Paper deep dive
Riemannian Deep Learning:Modules, Networks, and Geometries
Chen Ziheng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/25/2026, 1:39:03 AM
Summary
This thesis presents a unified framework for Riemannian deep learning, addressing limitations of Euclidean approximations and costly geometric operations. It introduces reusable neural modules such as Lie Group and Gyrogroup Batch Normalization, extends Multinomial Logistic Regression to SPD and general Riemannian manifolds, and develops specific networks for hyperbolic spaces (Proper Velocity, Busemann) and full-rank correlation matrices. Additionally, it proposes adaptive and stable Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and Cholesky-based product geometries, validated across vision, signal processing, and genomics.
Entities (13)
Relation Signals (9)
Ziheng Chen → authored → Riemannian Deep Learning: Modules, Networks, and Geometries
confidence 99% · Riemannian Deep Learning: Modules, Networks, and Geometries Ziheng Chen
Nicu Sebe → supervised → Ziheng Chen
confidence 99% · Advisor Prof. Nicu Sebe
Riemannian Deep Learning → includes → Lie Group Batch Normalization
confidence 92% · generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups
Riemannian Deep Learning → includes → Gyrogroup Batch Normalization
confidence 92% · introduce pseudo-reductive gyrogroups... and build a normalization framework on this structure
Adaptive Log-Euclidean Metrics → appliesto → SPD Manifolds
confidence 90% · study how the geometry itself can improve learning on SPD manifolds... make the metric learnable through parameterized matrix logarithms
Cholesky Product Geometry → appliesto → SPD Manifolds
confidence 90% · exploit the product structure of Cholesky factors to construct SPD metrics
Riemannian Deep Learning → extends → Multinomial Logistic Regression
confidence 90% · extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds
Proper Velocity Neural Networks → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroups, and extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds. It further develops neural networks for several important geometric representations, including an unconstrained model of hyperbolic space, Busemann-based hyperbolic learning, and full-rank correlation matrices. Finally, it introduces adaptive and computationally efficient Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and fast, stable Cholesky-based geometries. The proposed methods are supported by theoretical analysis and validated through numerical experiments and applications in vision, signal processing, graph learning, and genomics.
Tags
Links
- Source: https://arxiv.org/abs/2607.19305v1
- Canonical: https://arxiv.org/abs/2607.19305v1
Trouble viewing inline? Open PDF directly →
Full Text
747,883 characters extracted from source content.
Expand or collapse full text
Doctoral School in Information and Communication Technology Riemannian Deep Learning: Modules, Networks, and Geometries Ziheng Chen Advisor Prof. Nicu Sebe Università di Trento ELLIS Co-supervisor Prof. Bernhard Schölkopf Max Planck Institute for Intelligent Systems September 2026 arXiv:2607.19305v1 [cs.LG] 21 Jul 2026 Publications († corresponding author, ‡ equal contribution, ♣ equal supervision) The thesis is based on the following publications: • Chapter 3: [1] Ziheng Chen, Yue Song, Yunmei Liu, and Nicu Sebe. “A Lie Group Approach to Riemannian Batch Normalization.” ICLR 2024. [2] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. “Gyrogroup Batch Normalization.” ICLR 2025. • Chapter 4: [3] Ziheng Chen, Yue Song, Gaowen Liu, Ramana Rao Kompella, Xiao-Jun Wu, and Nicu Sebe. “Riemannian Multinomial Logistics Regression for SPD Neural Networks.” CVPR 2024. [4] Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, and Nicu Sebe. “RMLR: Extending Multinomial Logistic Regression into General Geometries.” NeurIPS 2024. • Chapter 5: [5] Ziheng Chen ‡ , Zihan Su ‡ , Bernhard Schölkopf, and Nicu Sebe. “Proper Ve- locity Neural Networks.” ICLR 2026. [6] Ziheng Chen, Bernhard Schölkopf, and Nicu Sebe. “Hyperbolic Busemann Neural Networks.” CVPR 2026. [7] Ziheng Chen, Xiao-Jun Wu, Bernhard Schölkopf, and Nicu Sebe. “Rieman- nian Networks over Full-Rank Correlation Matrices.” ICML 2026. • Chapter 6: [8] Ziheng Chen, Yue Song, Tianyang Xu, Zhiwu Huang, Xiao-Jun Wu, and Nicu Sebe. “Adaptive Log-Euclidean Metrics for SPD Matrix Learning.” IEEE TIP 2024. [9] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. “Fast and Stable Riemannian Metrics on SPD Manifolds via Cholesky Product Geometry.” ICLR 2026. The following papers are published but are not included in this thesis: (10) Ziheng Chen, Yue Song, Xiao-Jun Wu, Gaowen Liu, and Nicu Sebe. “Under- standing Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian Geometry.” ICLR 2025. (11) Shaocheng Jin, Tao Zhou, Rui Wang, Ziheng Chen, Xiaoqing Luo, Xiao-Jun Wu, and Josef Kittler. “Towards Robust EEG Decoding Based on Riemannian Self- Attention.” KDD 2026. (12) Rui Wang, Zihao Bi, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen † . “Riemannian Graph Convolutional Network for Skeleton-Based Two-Person Interaction Recognition.” IJCAI 2026. i (13) Xianglong Shi ‡ , Ziheng Chen †,‡ , Yunhan Jiang, and Nicu Sebe. “Intrinsic Lorentz Neural Network.” ICLR 2026. (14) Shanglin Li, Shiwen Chu, Okan Koç, Yi Ding, Qibin Zhao, Motoaki Kawanabe, and Ziheng Chen † . “HEEGNet: Hyperbolic Embeddings for EEG.” ICLR 2026. (15) Chen Hu ‡ , Ziheng Chen ‡ , Rui Wang, Yefeng Zheng, and Nicu Sebe. “Riemannian High-Order Pooling for Brain Foundation Models.” ICLR 2026. (16) Rui Wang, Yuting Jiang, Xiaoqing Luo, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen † . “Wasserstein-Aligned Hyperbolic Multi-View Clustering.” AAAI 2026 (Oral). (17) Rui Wang, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen † . “Towards a General Attention Framework on Gyrovector Spaces for Matrix Mani- folds.” NeurIPS 2025. (18) Rui Wang, Shaocheng Jin, Zhenyu Cai, Ziheng Chen † , Xiao-Jun Wu † , and Josef Kittler. “Learning a Better SPD Network for Signal Classification: A Riemannian Batch Normalization Method.” IEEE TNNLS 2025. (19) Chen Hu, Rui Wang ♣ , Xiaoning Song, Tao Zhou, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen ♣ . “A Correlation Manifold Self-Attention Network for EEG Decod- ing.” IJCAI 2025. (20) Rui Wang, Shaocheng Jin, Ziheng Chen † , Xiaoqing Luo, and Xiao-Jun Wu. “Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry.” CVPR 2025. (21) Rui Wang, Jiayao Jin, Ziheng Chen † , Cong Wu † , Xiao-Jun Wu, and Nicu Sebe. “Structural Topology Refinement Network for Skeleton-Based Action Recognition.” IEEE TIM 2025. (22) Rui Wang, Chen Hu, Ziheng Chen † , Xiao-Jun Wu † , and Xiaoning Song. “A Grassmannian Manifold Self-Attention Network for Signal Classification.” IJCAI 2024. (23) Rui Wang, Xiao-Jun Wu, Ziheng Chen, Cong Hu, and Josef Kittler. “SPD Man- ifold Deep Metric Learning for Image Set Classification.” IEEE TNNLS 2024. i Abstract Recently, deep neural networks operating on manifold-valued representations have gar- nered significant attention across various machine learning applications. However, many basic neural components remain tied to particular manifolds or rely on Euclidean approximations, while the underlying geometry can make repeated computations costly or numerically unstable. This thesis addresses these limitations by developing reusable modules, exploiting manifold-specific structures when general constructions are insuffi- cient, and introducing fast and stable geometries. We first generalize batch normaliza- tion and classification beyond individual manifolds. For normalization, we develop a framework on Lie groups, with theoretical control over Riemannian sample means and variances. To extend this principle beyond Lie groups, we introduce pseudo-reductive gyrogroups, which generalize classical gyrogroups and Lie groups, and build a normaliza- tion framework on this structure for a wider range of manifolds. For classification, we first extend Euclidean Multinomial Logistic Regression (MLR), which consists of a fully connected layer followed by softmax, to Symmetric Positive Definite (SPD) manifolds with flat metrics. We then use Riemannian trigonometry to extend MLR to general Riemannian manifolds. When a general formulation cannot exploit useful structure, we design networks for particular representations. For stable hyperbolic deep learning, we use Proper Velocity, an unconstrained representation of hyperbolic space, and develop its geometry and core neural layers. We also use Busemann functions to build intrinsic and efficient hyperbolic classification and fully connected layers. For full-rank correla- tion matrices, a normalized alternative to SPD matrices, we construct networks that operate directly on the manifold and derive accurate gradients for end-to-end training under two correlation geometries. Finally, we study how the geometry itself can im- prove learning on SPD manifolds. To move beyond fixed metrics, we make the metric learnable through parameterized matrix logarithms, allowing it to adapt to data and net- work dynamics with little additional computation. To improve efficiency and numerical stability, we exploit the product structure of Cholesky factors to construct SPD metrics with fast and stable closed-form operators. The proposed methods are supported by theo- retical analysis and validated through numerical experiments and empirical applications in vision, signal processing, graph learning, and genomics. i Keywords Riemannian deep learning, geometric deep learning, Riemannian manifolds, matrix manifolds, Lie groups, constant-curvature manifolds iv Acknowledgements My Ph.D. journey would not have been possible without the guidance, collaboration, and encouragement of many people. First and foremost, I would like to express my deepest gratitude to my advisor, Prof. Nicu Sebe. He gave me the freedom to pursue the questions that genuinely interested me and created an open research environment in which I could explore new directions with confidence. At the same time, he was always generous with his time and consistently offered encouragement and guidance in our discussions. His support extended well beyond research. He advised me on career decisions, helped me engage with the research community, and supported my scholarship applications. This balance between intellectual freedom and dependable support has profoundly shaped both the work presented in this thesis and the researcher I have become. I am also deeply grateful to Prof. Bernhard Schölkopf. During my research stay with him, his generosity and intellectual openness allowed me to explore distributional geometry in depth and to develop new perspectives on its role in machine learning. I value his insight, trust, and support. I have greatly enjoyed discussing research questions with him and feel fortunate that our collaboration will continue during my postdoctoral research. I would also like to thank Prof. Xiaojun Wu, who was my master’s advisor. We continued to collaborate throughout my Ph.D., and I deeply appreciate the generous help, guidance, and encouragement he provided along the way. My sincere thanks also go to my collaborators Yue Song, Rui Wang, and Shanglin Li, as well as to the students I mentored: Zihan Su, Xianglong Shi, Youxing Li, Ying Zhang, Chen Hu, Shaocheng Jin, Zihao Bi, and Yuting Jiang. Our discussions and joint efforts have been an essential part of my Ph.D. experience. I have learned a great deal from each of them, and I am grateful for the ideas, dedication, and enthusiasm they brought to our work together. Finally, I would like to thank my family. My deepest and most personal thanks go to my girlfriend, Yunmei Liu. Throughout this journey, she has supported me with patience and warmth and has always believed in me. She has made the difficult moments easier and the joyful moments more meaningful. She has also brought more good fortune into my life than I could ever have imagined. I am profoundly grateful to have her by my side. v vi Contents Notation and Conventionsxxi 1 Introduction1 1.1 Riemannian Deep Learning . . . . . . . . . . . . . . . . . . . . . . . . .1 1.2 Contributions and Outlines . . . . . . . . . . . . . . . . . . . . . . . .3 1.3 Summary of Papers Excluded from the Thesis . . . . . . . . . . . . . .5 2 Mathematical Background9 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 2.2 Topology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 2.3 Differential Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.4 Riemannian Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.5 Metric Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 2.6 Algebraic Structures on Manifolds . . . . . . . . . . . . . . . . . . . . . 29 2.7 Riemannian Optimization . . . . . . . . . . . . . . . . . . . . . . . . . 34 2.8 Matrix Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 2.8.1 Matrix Functions and Differentials . . . . . . . . . . . . . . . . 35 2.8.2 Backpropagation Through Matrix Functions . . . . . . . . . . . 36 2.9 Example Manifolds . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 2.9.1 Symmetric Positive Definite Manifolds . . . . . . . . . . . . . . 37 2.9.2 Full-Rank Correlation Manifolds . . . . . . . . . . . . . . . . . . 39 2.9.3 Grassmannian Manifolds . . . . . . . . . . . . . . . . . . . . . . 48 2.9.4 Special Orthogonal Groups . . . . . . . . . . . . . . . . . . . . . 50 2.9.5 Constant-Curvature Manifolds . . . . . . . . . . . . . . . . . . . 50 3 Riemannian Batch Normalization55 3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 3.2 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 56 3.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 3.2.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 3.2.3 Revisiting Normalization . . . . . . . . . . . . . . . . . . . . . . 59 3.2.4 LieBN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 3.2.5 Manifestations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 3.2.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 3.3 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 76 vii Contents 3.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 3.3.2 Pseudo-Reductive Gyrogroups . . . . . . . . . . . . . . . . . . . 79 3.3.3 GyroBN on Pseudo-Reductive Gyrogroups . . . . . . . . . . . . 84 3.3.4 Instantiations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 3.3.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 3.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 4 Riemannian Multinomial Logistic Regression109 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 4.2 Multinomial Logistic Regression on SPD Manifolds . . . . . . . . . . . 110 4.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 4.2.2 SPD Multinomial Logistic Regression . . . . . . . . . . . . . . . 111 4.2.3 SPD MLRs under Deformed LEM and LCM . . . . . . . . . . . 115 4.2.4 Rethinking the Existing LogEig Classifier . . . . . . . . . . . . . 116 4.3 Extension to General Riemannian Manifolds . . . . . . . . . . . . . . . 117 4.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 4.3.2 Riemannian Multinomial Logistic Regression . . . . . . . . . . . 119 4.3.3 SPD Multinomial Logistic Regressions . . . . . . . . . . . . . . 123 4.3.4 Lie Multinomial Logistic Regression . . . . . . . . . . . . . . . . 126 4.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 4.4.1 Experiments on the Proposed SPD MLRs . . . . . . . . . . . . 128 4.4.2 Experiments on the Proposed Lie MLR . . . . . . . . . . . . . . 131 4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 5 Riemannian Neural Networks133 5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133 5.2 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . . . . . 134 5.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134 5.2.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 5.2.3 Proper Velocity Geometry . . . . . . . . . . . . . . . . . . . . . 136 5.2.4 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . 139 5.2.5 Connections to the Hyperboloid . . . . . . . . . . . . . . . . . . 142 5.2.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 5.3 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . . . . . 149 5.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 5.3.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151 5.3.3 Busemann Multinomial Logistic Regression . . . . . . . . . . . . 152 5.3.4 Busemann Fully Connected Layer . . . . . . . . . . . . . . . . . 156 5.3.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 5.4 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . . . . . 163 5.4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163 5.4.2 Log-Euclidean Correlation Layers . . . . . . . . . . . . . . . . . 165 5.4.3 Poly-Hyperbolic-Cholesky Layers . . . . . . . . . . . . . . . . . 169 5.4.4 Backpropagation over Correlation Geometries . . . . . . . . . . 171 viii Contents 5.4.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 5.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183 6 Fast and Stable Geometries on SPD Manifolds185 6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 6.2 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . . . . . 186 6.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 186 6.2.2 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . 187 6.2.3 Parameter Learning . . . . . . . . . . . . . . . . . . . . . . . . . 195 6.2.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197 6.3 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . . . . . 204 6.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204 6.3.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 6.3.3 Product Geometries on the Cholesky . . . . . . . . . . . . . . . 206 6.3.4 Geometries on the SPD Manifold . . . . . . . . . . . . . . . . . 211 6.3.5 Applications to SPD Neural Networks . . . . . . . . . . . . . . 213 6.3.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213 6.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220 7 Conclusion and Future Work221 7.1 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 221 7.2 Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 222 Bibliography223 A Experimental Details and Additional Discussions245 A.1 Data Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245 A.1.1 Skeleton-Based Action Recognition and Gesture Data Sets . . . 245 A.1.2 Radar and EEG Signal Data Sets . . . . . . . . . . . . . . . . . 246 A.1.3 Image Classification Data Sets . . . . . . . . . . . . . . . . . . . 246 A.1.4 Graph Data Sets . . . . . . . . . . . . . . . . . . . . . . . . . . 246 A.1.5 Genomic Sequence Data Sets . . . . . . . . . . . . . . . . . . . 247 A.2 Backbone Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248 A.2.1 Backbone Networks on the SPD Manifold . . . . . . . . . . . . 248 A.2.2 Backbone Networks on the Grassmannian . . . . . . . . . . . . 249 A.2.3 Backbone Networks on Rotation Matrices . . . . . . . . . . . . 250 A.2.4 Backbone Networks on Hyperbolic Spaces . . . . . . . . . . . . 250 A.2.5 Riemannian Residual Network Backbones . . . . . . . . . . . . 252 A.3 Experimental Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253 A.3.1 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . 253 A.3.2 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . 257 A.3.3 Riemannian Multinomial Logistic Regression . . . . . . . . . . . 258 A.3.4 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . 261 A.3.5 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . 263 A.3.6 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . 266 ix Contents A.3.7 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . 269 A.3.8 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . 271 A.4 Additional Discussions . . . . . . . . . . . . . . . . . . . . . . . . . . . 273 A.4.1 Riemannian Multinomial Logistic Regression . . . . . . . . . . . 273 A.4.2 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . 282 A.4.3 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . 286 A.4.4 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . 289 A.4.5 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . 290 B Proofs291 B.1 Mathematical Background . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.1.1 Proof of Thm. 57 . . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.2 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 291 B.2.1 Proof of Thm. 61 . . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.2.2 Proof of Thm. 62 . . . . . . . . . . . . . . . . . . . . . . . . . . 292 B.2.3 Proof of Thm. 65 . . . . . . . . . . . . . . . . . . . . . . . . . . 292 B.2.4 Proof of Thm. 66 . . . . . . . . . . . . . . . . . . . . . . . . . . 293 B.2.5 Proof of Thm. 67 . . . . . . . . . . . . . . . . . . . . . . . . . . 293 B.2.6 Proof of Thm. 68 . . . . . . . . . . . . . . . . . . . . . . . . . . 294 B.2.7 Proof of Thm. 69 . . . . . . . . . . . . . . . . . . . . . . . . . . 294 B.2.8 Proof of Thm. 70 . . . . . . . . . . . . . . . . . . . . . . . . . . 296 B.2.9 Proof of Thm. 72 . . . . . . . . . . . . . . . . . . . . . . . . . . 297 B.3 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 298 B.3.1 Proof of Thm. 74 . . . . . . . . . . . . . . . . . . . . . . . . . . 298 B.3.2 Proof of Thm. 75 . . . . . . . . . . . . . . . . . . . . . . . . . . 299 B.3.3 Proof of Thm. 77 . . . . . . . . . . . . . . . . . . . . . . . . . . 300 B.3.4 Proof of Thm. 78 . . . . . . . . . . . . . . . . . . . . . . . . . . 300 B.3.5 Proof of Thm. 79 . . . . . . . . . . . . . . . . . . . . . . . . . . 302 B.3.6 Proof of Thm. 80 . . . . . . . . . . . . . . . . . . . . . . . . . . 303 B.3.7 Proof of Thm. 81 . . . . . . . . . . . . . . . . . . . . . . . . . . 304 B.3.8 Proof of Thm. 83 . . . . . . . . . . . . . . . . . . . . . . . . . . 305 B.3.9 Proof of Thm. 86 . . . . . . . . . . . . . . . . . . . . . . . . . . 306 B.3.10 Proof of Thm. 88 . . . . . . . . . . . . . . . . . . . . . . . . . . 309 B.3.11 Proof of Thm. 89 . . . . . . . . . . . . . . . . . . . . . . . . . . 309 B.3.12 Proof of Thm. 90 . . . . . . . . . . . . . . . . . . . . . . . . . . 312 B.3.13 Proof of Thm. 91 . . . . . . . . . . . . . . . . . . . . . . . . . . 312 B.3.14 Proof of Thm. 92 . . . . . . . . . . . . . . . . . . . . . . . . . . 314 B.3.15 Proof of Thm. 93 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.16 Proof of Thm. 95 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.17 Proof of Thm. 97 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.18 Proof of Thm. 98 . . . . . . . . . . . . . . . . . . . . . . . . . . 321 B.3.19 Proof of Thm. 99 . . . . . . . . . . . . . . . . . . . . . . . . . . 322 B.4 SPD Multinomial Logistic Regression . . . . . . . . . . . . . . . . . . . 326 B.4.1 Proof of Thm. 104 . . . . . . . . . . . . . . . . . . . . . . . . . 326 x Contents B.4.2 Proof of Thm. 105 . . . . . . . . . . . . . . . . . . . . . . . . . 327 B.4.3 Proof of Thm. 106 . . . . . . . . . . . . . . . . . . . . . . . . . 327 B.4.4 Proof of Thm. 107 . . . . . . . . . . . . . . . . . . . . . . . . . 327 B.4.5 Proof of Thm. 108 . . . . . . . . . . . . . . . . . . . . . . . . . 328 B.4.6 Proof of Thm. 109 . . . . . . . . . . . . . . . . . . . . . . . . . 328 B.4.7 Proof of Thm. 111 . . . . . . . . . . . . . . . . . . . . . . . . . 329 B.5 Riemannian Multinomial Logistic Regression . . . . . . . . . . . . . . . 332 B.5.1 Proof of Thm. 113 . . . . . . . . . . . . . . . . . . . . . . . . . 332 B.5.2 Proof of Thm. 114 . . . . . . . . . . . . . . . . . . . . . . . . . 333 B.5.3 Proof of Thm. 116 . . . . . . . . . . . . . . . . . . . . . . . . . 333 B.5.4 Proof of Thm. 117 . . . . . . . . . . . . . . . . . . . . . . . . . 333 B.5.5 Proof of Thm. 119 . . . . . . . . . . . . . . . . . . . . . . . . . 338 B.5.6 Proof of Thm. 120 . . . . . . . . . . . . . . . . . . . . . . . . . 338 B.6 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . . . . . 338 B.6.1 Derivation of the Proper Velocity Metric . . . . . . . . . . . . . 338 B.6.2 Proof of Thm. 121 . . . . . . . . . . . . . . . . . . . . . . . . . 339 B.6.3 Proof of Thm. 122 . . . . . . . . . . . . . . . . . . . . . . . . . 341 B.6.4 Proof of Thm. 123 . . . . . . . . . . . . . . . . . . . . . . . . . 342 B.6.5 Proof of Thm. 124 . . . . . . . . . . . . . . . . . . . . . . . . . 350 B.6.6 Proof of Thm. 125 . . . . . . . . . . . . . . . . . . . . . . . . . 350 B.6.7 Proof of Thm. 126 . . . . . . . . . . . . . . . . . . . . . . . . . 354 B.6.8 Proof of Thm. 127 . . . . . . . . . . . . . . . . . . . . . . . . . 359 B.6.9 Proof of Thm. 128 . . . . . . . . . . . . . . . . . . . . . . . . . 359 B.6.10 Proof of Thm. 129 . . . . . . . . . . . . . . . . . . . . . . . . . 360 B.7 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . . . . . 363 B.7.1 Proof of Thm. 130 . . . . . . . . . . . . . . . . . . . . . . . . . 363 B.7.2 Proof of Thm. 132 . . . . . . . . . . . . . . . . . . . . . . . . . 364 B.7.3 Proof of Thm. 135 . . . . . . . . . . . . . . . . . . . . . . . . . 366 B.7.4 Proof of Thm. 136 . . . . . . . . . . . . . . . . . . . . . . . . . 367 B.7.5 Proof of Thm. 137 . . . . . . . . . . . . . . . . . . . . . . . . . 368 B.8 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . . . . . 370 B.8.1 Proof of Thm. 138 . . . . . . . . . . . . . . . . . . . . . . . . . 370 B.8.2 Proof of Thm. 139 . . . . . . . . . . . . . . . . . . . . . . . . . 372 B.8.3 Proof of Thm. 142 . . . . . . . . . . . . . . . . . . . . . . . . . 373 B.8.4 Proof of Thm. 143 . . . . . . . . . . . . . . . . . . . . . . . . . 376 B.8.5 Proof of Thm. 145 . . . . . . . . . . . . . . . . . . . . . . . . . 377 B.8.6 Proof of Thm. 146 . . . . . . . . . . . . . . . . . . . . . . . . . 379 B.8.7 Proof of Thm. 144 . . . . . . . . . . . . . . . . . . . . . . . . . 380 B.9 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . . . . . 380 B.9.1 Proof of Thm. 147 . . . . . . . . . . . . . . . . . . . . . . . . . 380 B.9.2 Proof of Thm. 148 . . . . . . . . . . . . . . . . . . . . . . . . . 381 B.9.3 Proof of Thm. 149 . . . . . . . . . . . . . . . . . . . . . . . . . 381 B.9.4 Proof of Thm. 150 . . . . . . . . . . . . . . . . . . . . . . . . . 381 B.9.5 Proof of Thm. 152 . . . . . . . . . . . . . . . . . . . . . . . . . 382 xi Contents B.9.6 Proof of Thm. 154 . . . . . . . . . . . . . . . . . . . . . . . . . 382 B.9.7 Proof of Thm. 155 . . . . . . . . . . . . . . . . . . . . . . . . . 383 B.9.8 Proof of Thm. 156 . . . . . . . . . . . . . . . . . . . . . . . . . 384 B.9.9 Proof of Thm. 157 . . . . . . . . . . . . . . . . . . . . . . . . . 384 B.9.10 Proof of Thm. 158 . . . . . . . . . . . . . . . . . . . . . . . . . 385 B.9.11 Proof of Thm. 159 . . . . . . . . . . . . . . . . . . . . . . . . . 385 B.9.12 Proof of Thm. 160 . . . . . . . . . . . . . . . . . . . . . . . . . 385 B.9.13 Proof of Thm. 161 . . . . . . . . . . . . . . . . . . . . . . . . . 385 B.9.14 Proof of Thm. 163 . . . . . . . . . . . . . . . . . . . . . . . . . 386 B.9.15 Proof of Thm. 164 . . . . . . . . . . . . . . . . . . . . . . . . . 386 B.9.16 Proof of Thm. 167 . . . . . . . . . . . . . . . . . . . . . . . . . 386 B.9.17 Proof of Thm. 165 . . . . . . . . . . . . . . . . . . . . . . . . . 388 B.9.18 Proof of Thm. 166 . . . . . . . . . . . . . . . . . . . . . . . . . 389 B.10 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . . . . . 390 B.10.1 Proof of Thm. 169 . . . . . . . . . . . . . . . . . . . . . . . . . 390 B.10.2 Proof of Thm. 170 . . . . . . . . . . . . . . . . . . . . . . . . . 391 B.10.3 Proof of Thm. 172 . . . . . . . . . . . . . . . . . . . . . . . . . 393 B.10.4 Proof of Thm. 173 . . . . . . . . . . . . . . . . . . . . . . . . . 393 B.10.5 Proof of Thm. 174 . . . . . . . . . . . . . . . . . . . . . . . . . 395 B.10.6 Proof of Thm. 177 . . . . . . . . . . . . . . . . . . . . . . . . . 396 B.10.7 Proof of Thm. 179 . . . . . . . . . . . . . . . . . . . . . . . . . 396 xii List of Tables 2.1 Euclidean prototypes for topological concepts. . . . . . . . . . . . . . . 12 2.2 Euclidean prototypes for differential-geometric concepts. . . . . . . . . 18 2.3 Euclidean prototypes for Riemannian-geometric concepts. . . . . . . . . 24 2.4 Euclidean prototypes for metric-geometric concepts. . . . . . . . . . . . 29 2.5 Lie group structures and associated Riemannian operators on S n ++ . . . 39 2.6 Riemannian operators of (θ,α,β)-EM and BWM on S n ++ . . . . . . . . . 39 2.7 Isometric prototype spaces and diffeomorphisms on the correlation man- ifold. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 2.8 Vector operations and Riemannian operators under ECM and LECM. . 46 2.9 Vector operations and Riemannian operators under OLM and LSM. . . 46 2.10 Riemannian operators on the Grassmannian under ONB and P. . . . 49 2.11 Gyro operators on the Grassmannian under ONB and P. . . . . . . . 49 2.12 Lie group structures and Riemannian operators on rotation matrices. . 50 2.13 Gyro operators on stereographic and Beltrami–Klein models. . . . . . . 52 2.14 Riemannian operator templates for stereographic and radius constant- curvature coordinates. . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 3.1 Review of SPD Lie groups and invariant metrics. . . . . . . . . . . . . 59 3.2 Review of the rotation Lie group and invariant metric. . . . . . . . . . 59 3.3 Review of full-rank correlation Lie groups and invariant metrics. . . . . 60 3.4 Summary of some representative RBN methods. . . . . . . . . . . . . . 60 3.5 Summary of LieBN types. . . . . . . . . . . . . . . . . . . . . . . . . . 65 3.6 Key operators in calculating LieBN on SPD manifolds. . . . . . . . . . 68 3.7 Key operators in calculating LieBN on the rotation matrices. . . . . . . 68 3.8 Summary of LieBN on the correlation. . . . . . . . . . . . . . . . . . . 70 3.9 10-fold average results of SPDNet with and without SPDBN or LieBN. 72 3.10 Cross-validation results of TSMNet with SPDDSMBN and DSMLieBN. 73 3.11 Results of LieNet with or without rotation LieBN. . . . . . . . . . . . . 75 3.12 Results of SPDNet with or without correlation LieBN under different invariant metrics. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 3.13 Comparison of previous RBN methods with GyroBN, where M and V denote the sample mean and variance. . . . . . . . . . . . . . . . . . . 77 3.14 Summary of operators for GyroBN across representative manifolds. . . 99 xiii List of Tables 3.15 Efficiency (in μs) of gyroaddition on the radius manifold: closed form versus Riemannian definition. Values in parentheses indicate the runtime of the closed-form implementation as a percentage of the corresponding Riemannian implementation. The best results are bold. . . . . . . . . 100 3.16 Comparison of GyroBN against other Grassmannian BN methods un- der the GyroGr backbone. Here, accuracy is reported as a percent- age, fit time denotes the average training time per epoch (s/epoch), and #Params is reported in millions. Values in parentheses specify the di- mension of the Grassmannian input to the BN layer. The largest number of parameters is marked in red. . . . . . . . . . . . . . . . . . . . . . . 103 3.17 Comparison of GyroBN against LRBN across five CCSs. ROC is the test- ing AUC reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. When LRBN degenerates the backbone network, the results are highlighted with red. . . . . . . . . . . . . . . . 106 3.18 Comparison of CorNet with or without RBN layers. Accuracy is reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 4.1 Several MLRs on different geometries are special cases of our MLR. . . 123 4.2 Properties of deformed metrics on SPD manifolds (θ ̸= 0 and min(α,α + nβ) > 0). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124 4.3 Comparison of SPDNet with LogEig against SPD MLRs on the Radar data set. The best results are bold. . . . . . . . . . . . . . . . . . . . . 127 4.4 Comparison of SPDNet with LogEig against SPD MLRs on the HDM05 data set. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 4.5 Inter-session TSMNet results on Hinss2021. . . . . . . . . . . . . . . . . 127 4.6 Inter-subject TSMNet results on Hinss2021. . . . . . . . . . . . . . . . 128 4.7 Comparison of LogEig against SPD MLRs under the RResNet architecture.129 4.8 Comparison of LogEig against SPD MLRs under the SPDGCN architec- ture. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 4.9 Comparison of LogEig against SPD MLRs for direct classification. . . . 130 4.10 Results of LogEig MLR against Lie MLR under the LieNet architecture. 131 5.1 Failure and violation rates (%) of r⊗ H x in FP32. . . . . . . . . . . . . 144 5.2 ∥Log 0 (Exp 0 (v))− v∥. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 5.3 Gradient magnitude ∥∇ x f r (x)∥ across varying radii. . . . . . . . . . . . 145 5.4 Top-1 image classification accuracy (%) of hyperbolic MLRs on ResNet- 18. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic), which comes from Bdeir et al. [15, Tab. 1]. . . . . . . 146 5.5 Accuracies of hyperbolic networks on graph learning. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic). . 146 5.6 Results of Tangent FC (TFC) vs PV FC, and Tangent BN (TBN) vs GyroBN. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 5.7 Comparison of methods in calculating mean and variance in PV GyroBN. Time is measured in milliseconds per training epoch. . . . . . . . . . . 147 xiv List of Tables 5.8 Ablations on PVNN with or without exponential map for the input PV feature. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 5.9 Ablations on PV activations. . . . . . . . . . . . . . . . . . . . . . . . . 148 5.10 Comparison in MCC of hyperbolic and Euclidean convolutional networks, including PVCNN, on TEB data sets. . . . . . . . . . . . . . . . . . . . 149 5.11 Comparison of C-class MLR. In Dist, Real means the point-to-hyperplane distance is the real distance, obtained by inf y∈H d(x,y), where H is a hy- perplane and d is the geodesic distance; Pseudo denotes a surrogate that coincides with the real distance only in Euclidean geometry. Compact params indicate whether each logit avoids an additional manifold-valued parameter. Batch efficiency indicates whether the MLR can avoid ineffi- cient per-class loops in implementation (see Sec. A.4.2.1). In #Params, we highlight the heaviest in red. In FLOPs, we mark the slowest in red and the fastest in green. . . . . . . . . . . . . . . . . . . . . . . . . . . 155 5.12 Comparison of hyperbolic FC layers. For simplicity, BFC layers do not involve the gyroaddition and assume φ is the identity map, which is in line with the Möbius and Lorentz FC layers. . . . . . . . . . . . . . . . 158 5.13 Top-1 image classification accuracy (%) of MLR methods on the ResNet- 18 backbone. The best results within each hyperbolic model are bold. The slowest MLR and largest parameter count are shown in red. . . . 159 5.14 Genomic MCC of MLR methods under the CNN backbone. The best results within each hyperbolic model are bold. . . . . . . . . . . . . . . 160 5.15 Fit time (s/epoch) on genome sequence learning. The fastest times are bold and the slowest ones are red. . . . . . . . . . . . . . . . . . . . . 161 5.16 Comparison of hyperbolic FC layers on link prediction. The best results within each hyperbolic model are bold. . . . . . . . . . . . . . . . . . . 161 5.17 Node classification F1 scores of hyperbolic MLRs on the HGCN back- bone, where δ denotes graph hyperbolicity (lower is more hyperbolic). The best results within each hyperbolic model are bold. . . . . . . . . 162 5.18 Efficiency comparison: fit time (s/epoch) and parameter count. Slowest results and largest parameter counts are in red. . . . . . . . . . . . . . 163 5.19 Correspondence between Euclidean and correlation-based layers. For convolution, kernel-based FC refers to applying a convolution kernel to a receptive field, which is an FC transformation. . . . . . . . . . . . . . 165 5.20 Five-fold results and training time per epoch on four data sets. The top 3 results are highlighted with red, blue, and cyan. ∗ denotes reproduced results due to missing official code. . . . . . . . . . . . . . . . . . . . . 173 5.21 Ablations on mixed geometries. Each row shows the metric used for Con- volution (Conv), and each column is the metric for MLR. The diagonal entries indicate configurations where both layers use the same metric. The best result in each row is bold. . . . . . . . . . . . . . . . . . . . 174 5.22 SPDNet: SPD vs. correlation. . . . . . . . . . . . . . . . . . . . . . . . 175 xv List of Tables 5.23 Comparison of SPDMLR-Trivlz on raw covariances against CorMLR on raw correlations on all three data sets. The input matrix dimensions are 93× 93, 63× 63, and 20× 20, respectively. . . . . . . . . . . . . . . . . 175 5.24 SPD networks with or without normalized SPD inputs. . . . . . . . . . 180 5.25 Comparison of CorNet with or without activations. . . . . . . . . . . . 181 5.26 Average runtime (s) of a single forward pass in CorNet under different metrics and input dimensions. The best results are bold. . . . . . . . . 182 6.1 Parameter learning for the general matrix logarithm and exponential. . 196 6.2 Results of ALog on the HDM05 data set. The best results are bold. . . 198 6.3 Results of ALog on the FPHA data set. . . . . . . . . . . . . . . . . . . 199 6.4 Results of ALog on the AFEW data set. . . . . . . . . . . . . . . . . . 200 6.5 Results of fixed bases on the HDM05 and FPHA data sets. . . . . . . . 201 6.6 Comparison of RBN methods on the HDM05 data set. . . . . . . . . . 202 6.7 Experiments on RResNet under different geometries. . . . . . . . . . . 203 6.8 Comparison of Gyro MLRs on the NTU60 data set. . . . . . . . . . . . 203 6.9 Riemannian and gyro operators of different metrics on the Cholesky man- ifold. For the diagonal log metric, log(·) and exp(·) are diagonal loga- rithm and exponentiation. . . . . . . . . . . . . . . . . . . . . . . . . . 211 6.10 SPD MLRs under different metrics on the SPDNet backbone. The best results are bold. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214 6.11 SPD MLRs on the GyroSPD backbone. . . . . . . . . . . . . . . . . . . 214 6.12 Results on residual blocks. . . . . . . . . . . . . . . . . . . . . . . . . . 215 6.13 Failure probabilities (%) of geodesics under different metrics with small eigenvalues in L ∈ L n ++ . An output matrix containing any Inf or NaN is considered a failure. Here, DLM denotes the diagonal log metric, while DPM and DBWM denote θ-DPM and θ-DBWM, respectively. . . . . . 216 6.14 Swelling effects of geodesic SPD interpolations. Deeper greens indicate greater swelling. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 218 6.15 Number of matrix functions required per sample for a C-class SPD MLR. Spectral matrix functions include matrix logarithm, matrix power, and the Lyapunov operator. . . . . . . . . . . . . . . . . . . . . . . . . . . . 218 6.16 Asymptotic per-sample complexity of a C-class SPD MLR for an n× n input SPD matrix. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219 6.17 Average runtime (in seconds) of one SPD MLR training step across dif- ferent matrix dimensions. . . . . . . . . . . . . . . . . . . . . . . . . . . 220 A.1 Summary statistics for the graph data sets. . . . . . . . . . . . . . . . . 246 A.2 Summary statistics for TEB. . . . . . . . . . . . . . . . . . . . . . . . . 247 A.3 Summary statistics for the adopted GUE data sets. . . . . . . . . . . . 248 A.4 Matrix powers in LieBN-Cor under different metrics on each data set. . 257 A.5 (θ,α,β) of SPD MLRs on the SPDGCN backbone. . . . . . . . . . . . 259 A.6 Candidate values for hyperparameters in SPD MLRs. . . . . . . . . . . 260 A.7 Hyperparameters for PVNN that vary across graph data sets. . . . . . 262 A.8 Hyperparameters for PVNN that are shared across graph data sets. . . 262 xvi List of Tables A.9 Summary of the hyperbolic layers used in the graph node classification models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 263 A.10 Hyperparameters for TEB. . . . . . . . . . . . . . . . . . . . . . . . . . 263 A.11 Summary of hyperparameters used in the image classification task. . . . 264 A.12 Hyperparameters for genome sequence learning. . . . . . . . . . . . . . 265 A.13 Hyperparameters for node classification on Disease, Airport, PubMed, and Cora. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265 A.14 Hyperparameters in CorNets. . . . . . . . . . . . . . . . . . . . . . . . 267 A.15 Hyperparameter θ in θ-PCM and θ-BWCM. It is selected from the can- didate values used for SPDNet and GyroSPD. . . . . . . . . . . . . . . 272 A.17 Comparison of point-to-hyperplane distances. Real means the point-to- hyperplane distance is the real distance, obtained by inf y∈H d(x,y) with H as a hyperplane and d as the geodesic distance. Instead, Pseudo means the point-to-hyperplane distance is a surrogate, which only equals the real distance in Euclidean geometry. . . . . . . . . . . . . . . . . . 282 A.16 Comparison of hyperplanes. Compact params indicate whether the pa- rameterization requires an additional manifold-valued point. . . . . . . 282 xvii List of Tables xviii List of Figures 2.1 The black stars denote 2× 2 correlation matrices, while the red, green, and blue dots denote corresponding SPD matrices. The black dots denote the boundary of the SPD cone. . . . . . . . . . . . . . . . . . . . . . . 40 3.1 Illustration of LieBN on the SPD, rotation, and correlation Lie groups. 56 3.2 Minimal examples of applying LieBN. . . . . . . . . . . . . . . . . . . . 58 3.3 Visualization of input and output SPD matrices in LieBN. . . . . . . . 74 3.4 Test accuracy curves of LieNet with rotation LieBN. . . . . . . . . . . . 75 3.5 Illustration of GyroBN on manifold-valued data. Blue points, green points, and the red dashed curves indicate the input samples, normalized outputs, and data distributions, respectively. . . . . . . . . . . . . . . . 77 3.6 Minimal examples of applying GyroBN. . . . . . . . . . . . . . . . . . . 79 3.7 Comparison of derivation logic for gyrotranslation isometries. . . . . . . 80 3.8 Visualization of GyroBN across different geometries. Blue and green points represent input and normalized data, respectively. Red and cyan points denote the input and output batch means. Black points mark the manifold boundary, and the gray surface depicts the manifold. . . . . . 101 3.9 Training and testing curves of 1-block GyroGr on two NTU data sets. 104 4.1 Conceptual illustration of SPD hyperplanes induced by (α,β)-LEM and θ-LCM. In each subfigure, the black dots are SPSD matrices, denoting the boundary of S 2 ++ , while the blue, red, and yellow dots denote three SPD hyperplanes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 4.2 Illustration of the deformation (left) and Venn diagram (right) of met- rics on SPD manifolds, where IEM, SREM, and 1 4 PAM denote Inverse Euclidean Metric, Square Root Euclidean Metric, and Polar Affine Met- ric scaled by 1 /4, respectively. . . . . . . . . . . . . . . . . . . . . . . . 123 4.3 Conceptual illustration of SPD hyperplanes induced by five families of Riemannian metrics. The black dots denote the boundary of S 2 ++ . . . 124 4.4 Conceptual illustration of a Lie hyperplane. Each pair of antipodal black dots corresponds to a rotation matrix with an Euler angle of π, while the green dots denote a Lie hyperplane. . . . . . . . . . . . . . . . . . . . 127 5.1 Illustration: red curves are different horospheres of B v . . . . . . . . . . 152 5.2 Validation accuracy curves on ImageNet-1k. . . . . . . . . . . . . . . . 159 xix List of Figures 5.3 Illustration of the Log-Euclidean 1D convolution with two kernels. The 3-channel input is first split into two receptive fields along the channel dimension. In each receptive field, two kernels are applied to the product space. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168 5.4 Illustration of the PHCM convolution and MLR. The multi-channel input correlation matrices are denoted asC i c i=1 . For the convolutional layer, the illustration focuses on the transformation within a receptive field and assumes a single-channel output. . . . . . . . . . . . . . . . . . . . . . 170 5.5 Illustration of the decision hyperplanes in the correlation MLRs under five different geometries. The 3×3 correlation manifold can be embedded as an open elliptope in R 3 , by visualizing the strictly lower triangular part of each C ∈ Cor + (3). The black dots denote the boundary. The PHCM hyperplane is defined by the one in the β-concatenated Poincaré space. 174 5.6 Distribution of per-sample coefficients of variation of diagonal variances on FPHA. Higher values indicate stronger diagonal variability, which could cause nuisance noise. . . . . . . . . . . . . . . . . . . . . . . . . . 176 5.7 Distribution of per-sample coefficients of variation of diagonal variances on HDM05. Higher values indicate stronger diagonal variability, which could cause nuisance noise. . . . . . . . . . . . . . . . . . . . . . . . . . 177 5.8 Distribution of ratios of diagonal to off-diagonal entries on FPHA. . . . 178 5.9 Distribution of ratios of diagonal to off-diagonal entries on HDM05. . . 178 6.1 Accuracy curves on the FPHA data set. . . . . . . . . . . . . . . . . . . 199 6.2 Visualization of parameters in the ALog layer on the HDM05 and FPHA data sets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201 6.3 Geodesic interpolation of SPD matrices under different Riemannian met- rics. Each 3× 3 SPD matrix can be visualized as an ellipsoid [9]. The two endpoints are fixed across all metrics. . . . . . . . . . . . . . . . . 217 B.1 Illustration of the Euclidean spaces LT 0 (m), Hol(m) and Row 0 (m), where ⋆ can be obtained by symmetry. . . . . . . . . . . . . . . . . . . . . . 375 x Notation and Conventions This chapter collects the notation used throughout the thesis. General Sets and Linear Algebra R, R n The real numbers and the n-dimensional Euclidean space. R + , R ++ The non-negative and positive real numbers. I n , 0 n , 1 n , 1, 0 n×n The n× n identity matrix, the n-dimensional zero vector, the n-dimensional all-one vector (also written 1 when its dimension is clear), and the n × n zero matrix. Dimen- sion subscripts may be omitted when they are clear from context. ⟨x,y⟩, ∥x∥The Euclidean inner product and the induced norm. ∥A∥, tr(A), rank(A)The Frobenius norm, trace, and rank of a matrix. diag(x), diag(A)The diagonal matrix induced by a vector and the diagonal part of a matrix, with the intended meaning specified by context. argmin, argmaxThe minimizer and maximizer operators. Manifold Geometry X,YGeneric sets or topological spaces. XA generic metric space. M,NSmooth or Riemannian manifolds. C ∞ (M)The space of smooth real-valued functions on M. T p MThe tangent space of M at p. g p The Riemannian metric at p, viewed as an inner product on T p M. xxi Notation and Conventions g L , g R Left- and right-invariant Riemannian metrics on a Lie group. ⟨v,w⟩ p , ∥v∥ p The Riemannian inner product and norm in T p M. d M (x,y)The geodesic distance between x,y ∈M. α, L(α)A general curve on a space and its length. γA geodesic or geodesic ray. ηThe momentum parameter used in LieBN and GyroBN running-statistics updates. M n K , d K , D K The n-dimensional model space of constant curvature K, its distance, and the diameter of M 2 K . CAT(K)A metric space whose geodesic triangles satisfy the CAT(K) comparison inequality. π C The orthogonal projection onto a convex subset C of a CAT(0) space. ∂XThe boundary at infinity of a metric space, defined by equiv- alence classes of asymptotic geodesic rays when available. B γ (x), HB γ τ , H γ τ The Busemann function associated with a geodesic ray γ, its horoball, and its horosphere at level τ. Exp p , Log p The Riemannian exponential and logarithmic maps at p. PT p→q Parallel transport from T p M to T q M. T p→q A vector transport from T p M to T q M. H a,p , ̃ H ̃ A,P Euclidean and Riemannian hyperplanes used by multino- mial logistic regression. P k , A k , ̃ A k Class-wise Riemannian prototype, fixed tangent-space pa- rameter, and transported tangent parameter in RMLR. z k , r k , v k (x)The Euclidean direction parameter, scalar offset parameter, and class score used by the PV MLR and PV FC layers. M, B, v 2 , sThe Riemannian mean, biasing parameter, variance, and scaling factor used by Riemannian normalization layers. f ∗,p , d p f (v)The differential of a smooth map f at p, written in Chapter 2 as a linear map and, in later computations, equivalently as applied to a tangent vector v. xxii Notation and Conventions ∇ w f, grad w fThe Euclidean gradient of a scalar objective at a Euclidean parameter w and the Riemannian gradient at w ∈ M, re- spectively. X(M), X(α)Smooth vector fields on M and smooth vector fields along a curve α. [X,Y ]The Lie bracket of smooth vector fields X,Y ∈ X(M). D, D dt , V ′ The Levi-Civita connection and the induced derivative V ′ = D dt (V ) along a curve. FM(x i ), WFM(w i ,x i ) The Fréchet mean and weighted Fréchet mean. Algebraic and Gyrovector Operations G, e, EA group-like algebraic structure and its identity element. ⊕A binary operation for groups, Lie groups, and gyrogroups. L x , R x Left and right translations by x on a Lie group. L x Left gyrotranslation by x, defined by L x (y) = x⊕ y. ⊖xThe inverse of x under an addition-like operation. In gy- rogroups, this is called the gyro-inverse. ⊙Scalar multiplication in gyrovector spaces. gyr[x,y]The gyration generated by x and y. ⟨x,y⟩ gyr The gyro inner product. ∥x∥ gyr The gyronorm. d gyr (x,y)The gyrodistance. Bar η The binary barycenter operator with weight η, used for run- ning mean updates on gyrogroups. Matrix Manifolds S n The vector space of n× n symmetric matrices. S n ++ The manifold of n×n symmetric positive definite matrices. S + n The set of n× n symmetric positive semidefinite matrices. Chol(P )The Cholesky factor of an SPD matrix P. xxiii Notation and Conventions L n ++ The space of n× n lower triangular matrices with positive diagonal entries. log(P ), exp(A), log α (P ), log −1 α (A) The matrix logarithm and exponentiation, followed by the general matrix logarithm used by ALEM and its inverse. ⊕ ALE , ⊙ ALE The ALEM-induced element addition and scalar multipli- cation on S n ++ . f (S), f ∗,S , L f A symmetric matrix function induced by a scalar function f, its differential at S, and the corresponding Loewner ma- trix. A⊛ BThe Hadamard product of matrices A and B with the same size. log ∗,P , Chol ∗,P , Chol −1 ∗,L The differentials of the matrix logarithm, the Cholesky de- composition, and the inverse Cholesky map at the indicated points. P θ (P ), (P θ ) ∗,P The matrix power map P 7→ P θ on SPD matrices and its differential at P. φ :S n ++ →S n , ⊙ φ , g φ , d φ An isometry from SPD matrices to a Euclidean space and its induced abelian group operation, pullback Euclidean met- ric, and geodesic distance. Dlog(D), ψ LC (P )The diagonal element-wise logarithm and the Log-Cholesky map ψ LC (P ) =⌊L⌋ + Dlog(D(L)), where L = Chol(P ). L P [V ]The Lyapunov operator defined byL P [V ]P +PL P [V ] = V . ⟨A,B⟩ (α,β) , ∥A∥ (α,β) The two-parameter Euclidean inner product on S n and its induced norm. STThe admissible parameter set (α,β) ∈ R 2 | min(α,α + nβ) > 0. ⌊A⌋, D(A)The strictly lower triangular part and the diagonal matrix formed from the diagonal of A. A 1/2 The Cholesky half-diagonal operator A 1/2 =⌊A⌋ + 1 2 D(A). AIAffine-invariant geometry on SPD matrices. LELog-Euclidean geometry on SPD matrices. LCLog-Cholesky geometry on SPD matrices. PEPower-Euclidean geometry on SPD matrices. xxiv Notation and Conventions (α,β)-AIMThe two-parameter Affine-Invariant Metric (AIM) on SPD matrices. (α,β)-LEMThe two-parameter Log-Euclidean Metric (LEM) on SPD matrices. (θ,α,β)-AIMThe three-parameter Affine-Invariant Metric (AIM) on SPD matrices. θ-LCMThe parameterized Log-Cholesky Metric (LCM) on SPD matrices. (θ,α,β)-EMThe three-parameter Euclidean Metric (EM) on SPD ma- trices. (a,b)-ALEMThe Adaptive Log-Euclidean Metric on SPD matrices. θ-PCMThe deformed Power-Cholesky Metric on SPD matrices. (θ, M)-BWCMThe deformed Bures–Wasserstein–Cholesky Metric on SPD matrices. ⊕ AI , ⊙ AI AIM-induced SPD gyroaddition and scalar gyromultiplica- tion. ⊕ LieAI , ⊕ LE , ⊕ LC , ⊕ θ-AI , ⊕ θ-LC , ⊖ LieAI SPD Lie group operations under AIM, LEM, LCM, and their power-deformed AIM and LCM variants, with ⊖ LieAI denoting the inverse under ⊕ LieAI when operation-specific disambiguation is needed. ⊕ LE , ⊙ LE , ⊕ LC , ⊙ LC LEM- and LCM-induced SPD gyrovector operations, which coincide with linear operations in the corresponding global charts. BWBures–Wasserstein geometry on SPD matrices. g AI , g LE , g LC , g θ-E , g BW , g CRI , g θ-CRI Metric tensors associated with the corresponding SPD ge- ometries. GL(n), SL(n)The general linear group and the special linear group. O(n)The orthogonal group. St(p,n)The Stiefel manifold represented by orthonormal frames. Gr(p,n), f Gr(p,n)The Grassmannian manifold represented by orthonormal bases and projection matrices. xxv Notation and Conventions I p,n , e I p,n Identity points of the Grassmannian under the Orthonor- mal Basis (ONB) and Projector Perspective (P) represen- tations. ⊕ Gr , ⊙ Gr , ⊖ Gr Grassmannian gyroaddition, scalar multiplication, and in- verse under the ONB representation. e ⊕ Gr , e ⊙ Gr , e ⊖ Gr Grassmannian gyroaddition, scalar multiplication, and in- verse under the P representation. π : Gr(p,n)→ f Gr(p,n)The Grassmannian isometry π(U ) = U ⊤ from ONB to P. Matrix commutatorThe operation [A,B] = AB− BA. Cor + (n)The manifold of full-rank correlation matrices. Cor, I, invThe correlation normalization map, the cor-inversion oper- ator on correlation matrices, and the matrix inversion op- erator on SPD matrices. LT n The Euclidean space of n× n lower triangular matrices. LT 0 (n)The Euclidean space of n× n strictly lower triangular ma- trices. LT 1 (n)The affine space of n×n lower triangular matrices with unit diagonal. L n The manifold of Cholesky factors of full-rank correlation matrices, consisting of lower triangular matrices with posi- tive diagonal entries and unit row norm. Hol(n), Row 0 (n)The spaces of symmetric hollow matrices and symmetric matrices with null row sum, respectively. Θ, φ EC , Log ◦ , Exp ◦ , Log ⋆ , Exp ⋆ Diffeomorphisms and charts used by ECM, OLM, and LSM on correlation matrices. D(H), D ⋆ (C)The diagonal correction operators used in OLM and LSM, respectively. Diag(n), Diag + (n)The vector space of diagonal matrices and the positive di- agonal matrix manifold. S n , S ± (n)The permutation group and the signed permutation group acting on correlation matrices. (X) sym , (X) sym+ The averaged symmetrization (X + X ⊤ )/2 and the unnor- malized symmetrization X + X ⊤ , respectively. xxvi Notation and Conventions HS i , PHS n−1 , P n−1 The open hemisphere, product hemisphere space, and prod- uct of unit Poincaré balls used by PHCM and GyroBN on correlation manifolds. Φ : Cor + (n)→ P n−1 The row-wise Cholesky identification from full-rank corre- lation matrices to the product of unit Poincaré balls. SO(n), so(n)The special orthogonal group and its Lie algebra. Constant-Curvature Manifolds M n K The K-radius model with curvature parameter K. st n K , D n K The K-stereographic model and its positive-curvature pro- jected hypersphere case. S n K The spherical model space with curvature parameter K. S n , D n The unit sphere and unit projected hypersphere. Unit-space notation suppresses the fixed curvature subscript. P n K The Poincaré ball model of hyperbolic space with curvature K < 0. P n , K n , H n , L n The unit Poincaré ball, unit Beltrami–Klein ball, unit hy- perboloid, and unit Lorentz reference space. Unit-space notation suppresses the fixed curvature subscript. ⊕ K , ⊖ K , ⊙ K Gyroaddition, gyro-inverse, and scalar gyromultiplication on the K-stereographic model. λ K x The conformal factor of the K-stereographic model at x. ⊕ M , ⊖ M , ⊙ M Möbius gyroaddition, gyro-inverse, and scalar multiplica- tion on the Poincaré ball. ⊕ M K , ⊖ M K , ⊙ M K Gyroaddition, gyro-inverse, and scalar gyromultiplication on the K-radius model. L n K , H n K Equivalent notation for the Lorentz, or hyperboloid, model of hyperbolic space with curvature K < 0. ⊕ L , ⊖ L , ⊙ L Lorentz gyroaddition, gyro-inverse, and scalar gyromultipli- cation on the Lorentz model. ⟨x,y⟩ L , ∥x∥ L The Lorentzian inner product and its associated quantity. xxvii Notation and Conventions ⟨x,y⟩ K , ∥x∥ K The curvature-dependent ambient bilinear form on the K- radius model and its induced tangent-space norm. On the K < 0 branch, the ambient form is Lorentzian and becomes positive definite only after restriction to a tangent space. K n K The Beltrami–Klein model of hyperbolic space with curva- ture K < 0. ⊕ E , ⊖ E , ⊙ E Einstein gyroaddition, gyro-inverse, and scalar gyromulti- plication on the Beltrami–Klein model. γ x The Einstein gamma factor on the Beltrami–Klein model. tan K , sin K , cos K Curvature-aware trigonometric functions used for constant- curvature models. PV n K The Proper Velocity model of hyperbolic space with curva- ture K < 0. ⊕ U , ⊖ U , ⊗ U PV gyroaddition, gyro-inverse, and scalar gyromultiplica- tion. β x , γ y The relativistic beta factor on PV space and the gamma factor used by the PV–Poincaré isometry. π PV n K →P n K , π P n K →PV n K The mutually inverse isometries between the PV model and the Poincaré ball. π PV n K →H n K , π H n K →PV n K The mutually inverse isometries between the PV model and the hyperboloid model. xxviii Chapter 1 Introduction 1.1 Riemannian Deep Learning Over the past decade or so, Deep Neural Networks (DNNs) have achieved signifi- cant progress in machine learning [102, 127, 97, 202]. Traditionally, DNNs have been developed under the assumption that the latent geometry of the input data is Eu- clidean. However, many applications involve non-Euclidean structures, such as mani- folds [30, 88]. Therefore, deep learning over Riemannian spaces, referred to as Rieman- nian deep learning, has shown great success in diverse applications, such as computer vision [107, 106, 108, 47, 120], natural language processing [76, 181, 141], graph and knowledge-graph learning [163, 38, 42, 136], multimodal learning [67, 166], recommenda- tion systems [223], signal processing [31, 123], human neuroimaging [167, 134, 135, 233], medical imaging [37], astronomy [44], and genome sequence modeling [119]. Commonly encountered manifolds include special orthogonal groups [165], Symmetric Positive Def- inite (SPD) manifolds [9], Grassmannian manifolds [71, 19], spherical manifolds [165], and hyperbolic manifolds [173, 33]. Many of these manifolds admit computationally tractable Riemannian operators, including geodesics, exponential and logarithmic maps, and parallel transport. Building on these geometric tools, several fundamental Euclidean neural network components have been generalized to manifold-valued data, including normalization [31, 34, 142, 123], attention [167], residual blocks [201, 118], classification [76, 159, 161], and Fully Connected (FC) and convolutional layers [106, 107, 108, 76, 181, 45, 161]. However, most existing constructions either remain tied to particular geometries or carry out neural computations in intermediate flat spaces that approximate the intrinsic geometry. Existing normalization methods are either confined to particular SPD met- 1 1.1. Riemannian Deep Learning rics [31, 123], restricted to matrix Lie groups with a specific distance [34, Sec. 3.2], or applicable more generally but lack theoretical guarantees for controlling sample statis- tics, as in ManifoldNorm [34, Algs. 1–2] and RBN [142, Alg. 2]. Classification layers frequently map manifold-valued features to tangent Euclidean spaces [106, 31], ambient Euclidean spaces [107], or coordinate Euclidean spaces [36], while intrinsic alternatives may require the generalized law of sines [76] or gyrovector structures [159, 161]. FC and convolutional layers exhibit similar limitations. Early constructions are tailored to SPD, rotation, and Grassmannian manifolds [106, 107, 108], while hyperbolic coun- terparts rely on tangent-space mappings [76, 147], Lorentz spacetime [45], or Poincaré geometry [181]. The weighted-Fréchet-mean convolution [37] applies more broadly, but constrains the output manifold dimension to match the input dimension. These lim- itations motivate unified principles that formulate a module once at an appropriate geometric level and then instantiate it across manifolds carrying the required common structure. Beyond individual modules, deep architectures have been developed for matrix- valued manifolds, including SPD manifolds [106, 37], Grassmannian manifolds [108], and rotation manifolds [107], as well as vector-valued manifolds, including spherical manifolds [13, 182] and hyperbolic manifolds [76, 13, 182, 45, 15]. Nevertheless, ex- isting networks remain concentrated on a limited set of geometric representations and construction tools. Hyperbolic networks, for example, predominantly use the Poincaré ball [76, 181] and Lorentz (hyperboloid) models [45, 15], while alternative models remain less explored [33]. Beyond standard Riemannian and gyrovector operators, Busemann functions and horospheres [29, Ch. I.8] have been incorporated into hyperbolic SVMs [73], hyperbolic PCA [39], Sliced-Wasserstein distances [24], and prototype-learning methods [81]. Despite these developments, their use in neural network design remains limited. Likewise, correlation matrices have received far less attention than SPD co- variance representations despite being statistically compact alternatives to covariance matrices [7]. Their recently developed tractable Riemannian geometries [195, 191] in- dicate a broader design space that remains underexplored. Finally, both the module and network designs described above ultimately depend on the underlying Riemannian geometry. A Riemannian metric is not merely a means of measuring distance. It provides the theoretical foundation and concrete computa- tional primitives of Riemannian learning algorithms. By assigning inner products to tangent spaces, the metric determines geodesic distances, exponential and logarithmic maps, parallel transport, and Fréchet statistics [165]. These operators enter directly into the design of deep-network components. For example, hyperbolic geometry de- 2 Chapter 1. Introduction fines Riemannian classifiers [76, 181], weighted Fréchet means support convolutional layers [37] and normalization [31, 34, 123], and exponential and logarithmic maps un- derpin attention [167], residual blocks [118], and pooling layers [205]. Consequently, changing the Riemannian metric changes not only the geometry, but also the formulas, parameterizations, computational cost, numerical stability, and ultimately the practical behavior of the associated deep networks. This role is especially evident on the SPD manifold, where a broad range of metrics has been developed [171, 9, 70, 137, 22, 93]. However, most existing metric tensors are fixed, which may limit the expressivity of the induced geometry and its ability to adapt to data or the dynamics of deep networks. Moreover, numerical stability is particularly important in deep-network training, where metric-induced operators are repeatedly evaluated. Developing Riemannian metrics that balance flexibility, tractability, efficiency, and stability is therefore a significant problem for Riemannian deep learning. 1.2 Contributions and Outlines The contributions of this thesis are organized around three connected perspectives: uni- fied Riemannian module design across manifolds, manifold-specific Riemannian network design, and the design of the underlying Riemannian geometries. The first perspective formulates principled network modules which can be applied to different geometries, including normalization and classification layers. The second perspective addresses cases in which a fully general construction is either intractable or unable to exploit useful manifold-specific structure. It therefore uses the additional structures of partic- ular manifolds to develop manifold-specific modules and network architectures. The third perspective moves from network design under prescribed geometries to the de- sign of the geometry itself, developing flexible, efficient, and numerically stable metrics on SPD manifolds. The contributions and organization of the remaining chapters are summarized below. Chapter 2 establishes the mathematical foundations used throughout the thesis. The chapter is organized into two parts. The first part develops the general theory required by the subsequent methods, including topology, differential and Riemannian geome- try, metric geometry, algebraic structures on manifolds, Riemannian optimization, and matrix functions, and relates these constructions to their Euclidean counterparts. The second part turns to the concrete manifolds studied in later chapters: SPD, full-rank correlation, Grassmannian, and constant-curvature manifolds, together with special or- thogonal groups. 3 1.2. Contributions and Outlines Chapter 3 develops unified frameworks for Batch Normalization (BN) across struc- tured classes of manifolds, with the goal of controlling both the Riemannian mean and variance beyond a single manifold or metric. Sec. 3.2 first introduces Lie Group Batch Normalization (LieBN) on Lie groups under invariant metrics. LieBN uses group trans- lations for centering and biasing and tangent-space scaling at the identity for variance control, thereby extending Euclidean BN while preserving the manifold structure. Con- crete manifestations are developed on SPD, rotation, and full-rank correlation mani- folds, together with efficient implementations and experimental validation. Sec. 3.3 first introduces pseudo-reductive gyrogroups, a new algebraic structure that generalizes classical gyrogroups and Lie groups, thereby providing a more general foundation for principled normalization. Building on this structure, it develops Gyrogroup Batch Nor- malization (GyroBN), which recovers LieBN as a special case and extends the same nor- malization principle to manifolds that need not possess a Lie group structure. Finally, GyroBN is instantiated on the Grassmannian, constant-curvature, and full-rank corre- lation manifolds, and experiments evaluate the resulting layers across matrix-manifold, constant-curvature, and graph learning tasks. Chapter 4 develops a unified framework for intrinsic classification across Riemannian manifolds. Sec. 4.2 first extends Euclidean Multinomial Logistic Regression (MLR) to SPD manifolds with flat pullback metrics. By formulating classification through the geodesic margin between an SPD input and a decision hyperplane, it derives closed- form classifiers and provides an intrinsic explanation for the widely used LogEig MLR. Sec. 4.3 then replaces the potentially intractable point-to-hyperplane infimum with a Riemannian-trigonometric formulation, extending the classifier beyond flat SPD ge- ometries. The resulting Riemannian Multinomial Logistic Regression (RMLR) requires only an explicit Riemannian logarithmic map, incorporates several existing manifold classifiers as special cases, and yields new instantiations on SPD manifolds under mul- tiple metric families and on the special orthogonal group. Finally, Sec. 4.4 evaluates these classifiers within feedforward, residual, graph, and Lie group networks, as well as in direct manifold-valued classification. Chapter 5 turns to manifold-specific Riemannian network design, exploiting the ad- ditional structures of particular hyperbolic models and correlation manifolds. Sec. 5.2 first introduces the Proper Velocity (PV) model, an unconstrained model of hyper- bolic space, and establishes its Riemannian geometry. Based on this geometry, Proper Velocity Neural Networks (PVNNs) develop MLR, FC, convolutional, activation, and normalization layers for stable hyperbolic deep learning. Sec. 5.3 uses Busemann func- tions and horospheres to develop intrinsic and efficient Busemann Multinomial Logistic 4 Chapter 1. Introduction Regression (BMLR) and Busemann Fully Connected (BFC) layers on both the Poincaré and Lorentz models. BMLR interprets its logits through point-to-horosphere distances, while BFC uses the same Busemann logits to construct feature transformations. Sec. 5.4 then develops Correlation Networks (CorNets) for full-rank correlation matrices. It con- structs MLR, FC, and convolutional layers under five correlation geometries and derives accurate Riemannian backpropagation. Experiments across vision, graph, and genome learning tasks validate these manifold-specific designs. Chapter 6 develops flexible, fast, and numerically stable Riemannian metrics on SPD manifolds. Whereas the preceding chapters design modules and networks under prescribed geometries, this chapter designs the underlying metrics. Sec. 6.2 first devel- ops Adaptive Log-Euclidean Metrics (ALEMs). By parameterizing general matrix loga- rithms within a pullback framework, ALEMs adapt the geometry to data while retaining closed-form Riemannian operators and a compatible abelian Lie group structure; the learned metrics are instantiated in SPD networks and other Riemannian building blocks. Sec. 6.3 then develops product Cholesky geometries, including the Power-Cholesky Met- ric (PCM) and Bures–Wasserstein–Cholesky Metric (BWCM). These metrics avoid the scalar logarithms and exponentials used by the Log-Cholesky Metric (LCM) and admit fast and stable closed-form Riemannian and algebraic operators. Their effectiveness is evaluated in SPD classification, residual learning, and tensor interpolation, while sepa- rate experiments assess computational efficiency, scalability, and numerical stability. Chapter 7 summarizes the thesis and discusses future directions. 1.3 Summary of Papers Excluded from the Thesis The technical chapters focus on the core theoretical and methodological contributions. Fourteen additional publications on Riemannian deep learning belong to the same re- search theme and are summarized below. • Riemannian attention. We first developed self-attention on three specific man- ifolds: the Grassmannian [210], the full-rank correlation manifold [103], and the SPD manifold under the Bures–Wasserstein geometry [113]. Then, we general- ized geometry-specific constructions into a unified attention framework for general matrix manifolds [212]. • Geometry-aware electroencephalography (EEG) representation learn- ing. In Li et al. [135], we used hyperbolic embeddings to represent the hierarchical structure of EEG signals and improve cross-domain generalization. In Hu et al. 5 1.3. Summary of Papers Excluded from the Thesis [104], we introduced Riemannian high-order pooling on SPD manifolds to capture second-order correlations in EEG foundation models. • Geometry-specific Riemannian batch normalization. We developed SPD BN methods under specific Riemannian metrics [214, 215] for stable training. • Skeleton-based action recognition. We developed two SPD-based approaches to skeleton action recognition: a graph convolutional network that uses Gaussian embeddings of high-order skeletal statistics to model inter-subject interactions and global correlations for two-person interaction recognition [216], and an approach that partitions the skeleton into semantic body regions and models their long- range dependencies with SPD representations [213]. • Hyperbolic neural networks. In Shi et al. [179], we constructed a fully intrinsic hyperbolic Lorentz neural network whose FC, BN, concatenation, activation, and dropout modules operate within the Lorentz geometry. • Hyperbolic multi-view clustering. In Wang et al. [217], we developed hy- perbolic multi-view clustering that aligns view-specific distributions through a hyperbolic sliced-Wasserstein distance while preserving hierarchical semantics in the Lorentz manifold. • Riemannian interpretation of global covariance pooling. In Chen et al. [53], we provided a unified Riemannian interpretation of matrix functions in global covariance pooling, showing that their effectiveness is explained by the Rieman- nian classifiers they implicitly respect. • SPD deep metric learning. In Wang et al. [211], we combined an SPD network encoder with a Riemannian decoder, local covariance regularization, and deep metric learning to improve image-set classification. These papers are excluded from this thesis because they extend or apply the core works developed in this thesis. For example, CorAtt [103] applies OLM/LSM corre- lation geometry to EEG attention, while Sec. 5.4 systematizes how to build neural networks and differentiation over the correlation matrices. GyroAtt [212] adopts the gyro paradigm, following Sec. 3.3. HEEGNet [135] applies the Lorentz GyroBN de- veloped in Sec. 3.3 to domain-specific normalization. CBN [214] and GBWBN [215] specialize Sec. 3.2’s template to specific metrics. ILNN [179] accelerates the Lorentz 6 Chapter 1. Introduction GyroBN developed in Sec. 3.3 and follows Sec. 5.2’s point-to-hyperplane design. Fi- nally, RiemGCP [53] applies the RMLR framework developed in Sec. 4.3 to interpret matrix-function normalization as implicit SPD Riemannian classifiers. 7 1.3. Summary of Papers Excluded from the Thesis 8 Chapter 2 Mathematical Background 2.1 Introduction This chapter presents the mathematical foundations of the thesis in two parts. The first part develops the general theory and computational tools used throughout the subsequent chapters, beginning with topology and differential geometry [197], proceed- ing to Riemannian geometry [165], metric geometry [29], and algebraic structures on manifolds [128, 200], and then reviewing Riemannian optimization [1, 27], trivialization [132], and matrix functions and their differentials [21]. Although these concepts can be abstract, many generalize familiar Euclidean constructions, so Euclidean prototypes are provided whenever appropriate. The second part reviews the concrete spaces used later: Symmetric Positive Definite (SPD) and full-rank correlation manifolds, Grass- mannian manifolds, special orthogonal groups, and constant-curvature manifolds. Their Riemannian and algebraic operators expose both the common primitives that support unified module design across manifolds and the additional structures later exploited by manifold-specific methods. 2.2 Topology Topology provides the language for continuity and locality before coordinates are in- troduced [197, App. A]. This is essential because a manifold is usually not defined by a single global coordinate system as in Euclidean space, but by a family of local coordinate systems. Recall classical analysis on R n , where many basic notions are formulated in terms of open sets. In the standard Euclidean setting R n , openness is defined through open 9 2.2. Topology balls: a subset U ⊂ R n is open if, for every x ∈ U, there exists r > 0 such that the open ball B r (x) = y ∈ R n | ∥y− x∥ < r is contained in U. However, a general abstract space need not carry a norm or a distance, so it may not have a prior notion of open balls. Topology abstracts the essential closure properties of Euclidean open sets, leading to the following definition. Definition 1 (Topological space [197, Def. A.1]). A topology on a set X is a collection T of subsets of X such that (1) ∅∈T and X ∈T , (2) every union of elements of T belongs to T , i.e., for any U α α∈A ⊂T , [ α∈A U α ∈T .(2.1) (3) every finite intersection of elements of T belongs to T , i.e., for any integer m≥ 1 and any open sets U 1 ,...,U m ∈T , m \ i=1 U i ∈T .(2.2) The pair (X,T ) is called a topological space. Elements of T are called open sets. A subset F ⊂ X is called closed if X is open, where X =x∈ X | x /∈ F. A basis records enough open sets to reconstruct the full topology, and second count- ability controls the size of this local description. Definition 2 (Basis and second countability [197, Defs. A.6 and A.12]). Let (X,T ) be a topological space. A collectionB ⊂T is a basis forT if every open set U ∈T is a union of elements of B. Equivalently, for every x∈ U with U ∈T , there exists B ∈B such that x∈ B ⊂ U. A topological space is second countable if its topology has a countable basis. The next proposition gives the corresponding construction criterion when one starts from a candidate family of subsets rather than from an already specified topology. Proposition 3 (Criterion [197, Prop. A.8]). Let X be a set and letB be a collection of subsets of X. Then B is a basis for some topology T on X if and only if (1) X is the union of all sets in B, 10 Chapter 2. Mathematical Background (2) whenever x∈ B 1 ∩B 2 with B 1 ,B 2 ∈B, there exists B 3 ∈B such that x∈ B 3 ⊂ B 1 ∩ B 2 . In Euclidean space R n , the family of open balls generates the standard topology. Moreover, R n is second countable because the open balls with centers in Q n and positive rational radii form a countable basis. Subspace topology is used whenever a geometric object is described as a subset of an ambient space. Definition 4 (Subspace topology [197, Sec. A.2]). Let (X,T ) be a topological space and let Y ⊂ X. The subspace topology on Y is T Y =Y ∩ U | U ∈T.(2.3) With this topology, Y is called a subspace of X. For Euclidean space, the unit sphere S n−1 =x∈ R n |∥x∥ = 1 carries the subspace topology inherited from R n . Its open sets are exactly the sets S n−1 ∩ U, where U is open in R n . Definition 5 (Continuous map and homeomorphism [197, Sec. A.7]). Let (X,T X ) and (Y,T Y ) be topological spaces. A map f : X → Y is continuous if f −1 (V )∈T X for every V ∈T Y . A map f : X → Y is a homeomorphism if it satisfies: (1) f is bijective, (2) f and f −1 are continuous. In the Euclidean special case, this definition recovers the familiar notion of continuity from calculus. For a map f : R n → R m with the standard topologies, f is continuous in the sense of Thm. 5 if and only if, for every x ∈ R n and every ε > 0, there exists δ > 0 such that ∥f (y)− f (x)∥ < ε whenever ∥y− x∥ < δ. Homeomorphic spaces are topologically identical. A simple homeomorphism is the translation x 7→ x + a on R n , whose inverse is x7→ x− a. For manifolds, one also needs a separation condition that prevents distinct points from being topologically indistinguishable. Definition 6 (Hausdorff space [197, Def. A.16]). A topological space X is Haus- dorff if, for any two distinct points x,y ∈ X, there exist disjoint open sets U,V ⊂ X such that x∈ U and y ∈ V . 11 2.3. Differential Geometry Topological conceptEuclidean special case Open and closed setsB r (x) = y ∈ R n | ∥y− x∥ < r, B r (x) =y ∈ R n |∥y− x∥≤ r. Basis and second countability B Q =B r (q)| q ∈ Q n ,r ∈ Q >0 . Subspace topologyT S n−1 =S n−1 ∩ U | U ⊂ R n open. Continuous mapContinuous map in calculus. Hausdorff spacex ̸= y ⇒ B r (x)∩ B r (y) = ∅ for 0 < r < 1 2 ∥x− y∥. Table 2.1: Euclidean prototypes for topological concepts. The Hausdorff condition ensures that points can be separated by neighborhoods. For Euclidean space, if x ̸= y, then the open balls B r (x) and B r (y) are disjoint whenever 0 < r < 1 2 ∥x− y∥. Thus Euclidean space is Hausdorff. Tab. 2.1 summarizes the Euclidean special cases of the above topological concepts. 2.3 Differential Geometry In R n , a basis provides a global coordinate system: every point is represented by a unique coordinate tuple. By contrast, a manifold usually does not admit a single coordinate map covering the entire space. Instead, each point has a neighborhood homeomorphic to an open subset of R n , and this homeomorphism provides Euclidean coordinates on that neighborhood. This local Euclidean structure permits derivatives, tangent vectors, and smooth maps to be defined intrinsically. Definition 7 (Topological manifold [197, Sec. 5]). An n-dimensional topological manifold is a topological space M such that (1) M is Hausdorff, (2) M is second countable, (3) every point p∈M has a neighborhood U homeomorphic to an open subset of R n . The integer n is called the dimension of M. The Hausdorff and second-countability assumptions exclude pathological spaces and ensure that local coordinates behave like ordinary Euclidean neighborhoods. Charts are local coordinate systems, and an atlas is a collection of such coordinate systems covering the whole manifold. 12 Chapter 2. Mathematical Background Definition 8 (Chart and atlas [197, Sec. 5]). Let M be an n-dimensional topo- logical manifold. A chart on M is a pair (U,φ) satisfying: (1) U ⊂M is open, (2) φ : U → φ(U )⊂ R n is a homeomorphism onto an open subset of R n . Writing φ = (x 1 ,...,x n ), the functions x i : U → R are called the coordinate functions of the chart. Equivalently, if r i : R n → R denotes the i-th standard coordinate projection, then x i = r i ◦ φ. An atlas is a collection of charts A = (U α ,φ α ) α∈A satisfying M = [ α∈A U α .(2.4) To do calculus consistently across charts, changes of coordinates must preserve smoothness. Definition 9 (Smooth atlas and smooth manifold [197, Sec. 5]). Two charts (U,φ) and (V,ψ) on M are smoothly compatible if the following transition maps are smooth maps between open subsets of Euclidean spaces: (1) ψ◦ φ −1 : φ(U ∩ V )→ ψ(U ∩ V ), (2) φ◦ ψ −1 : ψ(U ∩ V )→ φ(U ∩ V ). A smooth atlas is an atlas whose charts are pairwise smoothly compatible. A smooth manifold is a topological manifold equipped with a maximal smooth atlas. Smooth compatibility means that changing coordinates does not destroy differentia- bility. This is the formal mechanism that allows a derivative computed in one coordinate chart to represent an intrinsic geometric object. Smooth maps between manifolds are defined by checking their coordinate represen- tations. Definition 10 (Smooth map and diffeomorphism [197, Sec. 6]). Let M and N be smooth manifolds. A map f :M→N is smooth if, for every p∈M, every chart (U,φ) around p, and every chart (V,ψ) around f (p) with f (U )⊂ V , the coordinate representation ψ◦ f ◦ φ −1 : φ(U )→ ψ(V )(2.5) is smooth. A smooth map f :M→N is a diffeomorphism if it satisfies: (1) f is bijective, 13 2.3. Differential Geometry (2) f −1 :N →M is smooth. Diffeomorphisms refine homeomorphisms by preserving smooth structure, so diffeo- morphic manifolds are identical from the viewpoint of smooth geometry. Before defining tangent spaces on manifolds, it is useful to recall what a tangent vector does in Euclidean calculus. Fix x∈ R n and a direction v ∈ R n . The directional derivative at x in the direction v can be viewed as an operator on smooth functions, D x,v : C ∞ (R n )→ R, D x,v (f ) = d dt t=0 f (x + tv).(2.6) This operator is linear in f and, by the ordinary product rule, satisfies D x,v (fh) = f (x)D x,v (h) + h(x)D x,v (f ), f,h∈ C ∞ (R n ).(2.7) Thus a Euclidean tangent vector can be recognized not only as an arrow v, but also as a first-order operator that differentiates smooth functions at x. The derivation viewpoint keeps precisely this algebraic behavior and extends it to manifolds, where no global vector structure is available. Definition 11 (Tangent space [197, Sec. 8]). LetM be a smooth manifold and let p∈M. A derivation at p is a map v : C ∞ (M)→ R satisfying: a (1) linearity, v(af + bh) = av(f ) + bv(h), a,b∈ R, f,h∈ C ∞ (M),(2.8) (2) the Leibniz rule, v(fh) = f (p)v(h) + h(p)v(f ), f,h∈ C ∞ (M).(2.9) The set of all derivations at p is a vector space, called the tangent space of M at p and denoted by T p M. a Strictly speaking, the domain should be the algebra of germs of smooth functions at p, namely equivalence classes of smooth functions that agree on some neighborhood of p. For simplicity, we write C ∞ (M). The tangent space T p M is a vector space whose operations are pointwise addition and scalar multiplication of derivations. Let (U,φ) = (U,x 1 ,...,x n ) be a chart contain- ing p. The chart induces the coordinate tangent vector ∂ ∂x i p , which is the derivation at 14 Chapter 2. Mathematical Background p defined by ∂ ∂x i p (f ) = ∂ (f ◦ φ −1 ) ∂r i (φ(p)), f ∈ C ∞ (M), i = 1,...,n.(2.10) Thus ∂ ∂x i p differentiates f in the i-th Euclidean coordinate direction after f is expressed in the chart. The coordinate tangent vectors ∂ ∂x 1 p ,..., ∂ ∂x n p form a basis of T p M [197, Prop. 8.9]. Consequently, every v ∈ T p M has a unique coordinate representation v = n X i=1 v i ∂ ∂x i p , v i ∈ R,(2.11) so this coordinate basis identifies T p M with R n . The differential of a smooth map generalizes the Jacobian matrix. Definition 12 (Differential [197, Sec. 8]). Let f :M→N be a smooth map and let p∈M. The differential of f at p is the linear map f ∗,p : T p M→ T f (p) N defined by (f ∗,p (v)) (h) = v(h◦ f ), v ∈ T p M, h∈ C ∞ (N ).(2.12) Throughout this chapter, we write the differential as f ∗,p and its action on v as f ∗,p (v). Later computations also use the notation d p f (v). The intrinsic differential recovers the ordinary Jacobian in local coordinates. Let (U,φ) = (U,x 1 ,...,x n ) be a chart around p∈M and let (V,ψ) = (V,y 1 ,...,y m ) be a chart around f (p)∈N . The differential f ∗,p is f ∗,p ∂ ∂x 1 p · ∂ ∂x n p = ∂ ∂y 1 f (p) · ∂ ∂y m f (p) ∂ (y 1 ◦ f ) ∂x 1 (p) · ∂ (y 1 ◦ f ) ∂x n (p) . . . . . . . . . ∂ (y m ◦ f ) ∂x 1 (p) · ∂ (y m ◦ f ) ∂x n (p) .(2.13) The (i,j)-th entry of this matrix is ∂ (y i ◦ f )/∂x j (p) for i = 1,...,m and j = 1,...,n. Hence, f ∗,p is represented by the Jacobian matrix of the local coordinate representation ψ◦ f ◦ φ −1 evaluated at φ(p) [197, Prop. 8.11]. The differential gives an intrinsic definition of the velocity of a curve. 15 2.3. Differential Geometry Definition 13 (Smooth curve and velocity [197, Sec. 8.6]). Let I ⊂ R be an open interval and let c : I →M be a smooth curve. The velocity vector of c at t 0 ∈ I is c ′ (t 0 ) := c ∗,t 0 d dt t=t 0 ! ∈ T c(t 0 ) M.(2.14) Equivalently, for every h∈ C ∞ (M), c ′ (t 0 )(h) = d dt t=t 0 h(c(t)).(2.15) Proposition 14 (Velocity in local coordinates [197, Prop. 8.15]). Let c : I → M be a smooth curve and let (U,φ) = (U,x 1 ,...,x n ) be a chart containing c(t). Then the velocity is obtained by differentiating the coordinate functions: c ′ (t) = n X i=1 d(x i ◦ c) dt (t) ∂ ∂x i c(t) .(2.16) Thus the coefficients of c ′ (t) in the chart-induced basis are precisely the ordinary derivatives of the coordinate representation φ◦ c. Every smooth curve through p yields a tangent vector in T p M, and conversely every tangent vector is the velocity of some smooth curve through p [197, Prop. 8.16]. Therefore, T p M =c ′ (0)| ε > 0, c : (−ε,ε)→M is smooth, c(0) = p.(2.17) For any curve c representing v ∈ T p M in this way, the derivation v acts as the directional derivative v(h) = d dt t=0 h(c(t)) [197, Prop. 8.17]. Curves also provide a classical method for computing differentials. Proposition 15 (Differential via curves [197, Prop. 8.18]). Let F : M → N be a smooth map, let p∈M, and let v ∈ T p M. If c : (−ε,ε)→M is any smooth curve satisfying c(0) = p and c ′ (0) = v, then the differential maps the velocity of c to the velocity of its image curve: F ∗,p (v) = (F ◦ c) ′ (0) = d dt t=0 F (c(t)).(2.18) 16 Chapter 2. Mathematical Background In particular, this expression is independent of the chosen representative curve. The following examples illustrate the above proposition in the vector and matrix cases. Example 16 (Sphere). Let S n = y ∈ R n+1 | ∥y∥ = 1 and consider the nor- malization map ν : R n+1 \0 → S n defined by ν(x) = x/∥x∥. For x ̸= 0 and v ∈ R n+1 , the curve c(t) = x + tv has initial velocity v. Applying Eq. (2.18) gives ν ∗,x (v) = d dt t=0 x + tv ∥x + tv∥ = 1 ∥x∥ v− ⟨x,v⟩ ∥x∥ 2 x ∈ T ν(x) S n .(2.19) In particular, when x∈ S n , the differential is ν ∗,x (v) = v−⟨x,v⟩x, the orthogonal projection of v onto T x S n . Example 17 (General linear group). Let GL(n) = A∈ R n×n | det(A)̸= 0 de- note the manifold of invertible real n × n matrices. For G ∈ GL(n), define L G : GL(n) → GL(n) by L G (B) = GB. Since GL(n) is an open submanifold of R n×n , its tangent spaces are identified with R n×n . For X ∈ T I n GL(n) ∼ = R n×n , the curve c(t) = I n + tX remains in GL(n) for sufficiently small t and satisfies c ′ (0) = X. Hence, (L G ) ∗,I n (X) = d dt t=0 G(I n + tX) = GX,(2.20) so the differential of left multiplication is again left multiplication [197, Ex. 8.19]. A vector field assigns a tangent vector to each point in a smooth way. Definition 18 (Vector field [197, Def. 12.7]). A vector field on a smooth manifold M is a map X :M→ TM such that X(p)∈ T p M for every p∈M. It is smooth if, for every smooth function f ∈ C ∞ (M), the function p7→ X(p)f is smooth. Tab. 2.2 summarizes the Euclidean special cases of the above differential-geometric concepts. 2.4 Riemannian Geometry Riemannian geometry equips each tangent space with an inner product. This turns local tangent vectors into measurable directions and induces global geometric objects 17 2.4. Riemannian Geometry Differential-geometric concept Euclidean special case Smooth manifoldsR n . Chartsφ = id R n . Smooth mapSmooth map in calculus. Tangent spaceT x R n ∼ = R n . Table 2.2: Euclidean prototypes for differential-geometric concepts. such as lengths and distances. Definition 19 (Riemannian manifold [165, Defs. 3.1 and 3.2]). LetM be a smooth manifold. A Riemannian metric on M is a smooth assignment p7→ g p : T p M× T p M→ R(2.21) such that, for every p∈M, g p satisfies: (1) g p is bilinear, (2) g p (v,w) = g p (w,v) for all v,w ∈ T p M, (3) g p (v,v) > 0 for all nonzero v ∈ T p M. The pair (M,g) is called a Riemannian manifold. We write the value of the metric as ⟨v,w⟩ p = g p (v,w) for v,w ∈ T p M. The induced norm is ∥v∥ p = q ⟨v,v⟩ p , v ∈ T p M.(2.22) The Riemannian metric extends the Euclidean inner product to curved spaces. Un- less explicitly stated otherwise, a manifold means a Riemannian manifold, and the pair (M,g) is abbreviated as M. We also use ⟨v,w⟩ p for the Riemannian metric. Once each tangent vector has a norm, one can measure the length of curves and define the induced shortest-path distance. Definition 20 (Length and geodesic distance [165, Defs. 5.11 and 5.15]). Let M be a manifold and let α : [a,b] → M be a piecewise smooth curve. The length of α is L(α) = Z b a ∥ ̇α(t)∥ α(t) dt.(2.23) 18 Chapter 2. Mathematical Background The geodesic distance between x,y ∈M is d M (x,y) = inf α L(α),(2.24) where the infimum is taken over all piecewise smooth curves α : [a,b] → M satis- fying α(a) = x and α(b) = y. The geodesic distance in R n is exactly the straight-line distance. On a connected manifold M, the distance function d M : M×M → [0,∞) is a metric, which makes (M, d M ) a metric space [165, Prop. 5.18]. The notion of a connection generalizes the directional derivative of tangent vectors to curved spaces. It is used to define geodesics, parallel transport, and curvature. Definition 21 (Connection and covariant derivative [165, Def. 3.9]). Let X(M) denote the space of smooth vector fields on M. A connection on M is a map D : X(M)× X(M)→ X(M),(X,Y )7→ D X Y,(2.25) that satisfies, for all X,Y,Z ∈ X(M), f 1 ,f 2 ,f ∈ C ∞ (M), and a,b∈ R: (1) C ∞ (M)-linearity in the first argument, D f 1 X+f 2 Y Z = f 1 D X Z + f 2 D Y Z,(2.26) (2) R-linearity in the second argument, D X (aY + bZ) = aD X Y + bD X Z,(2.27) (3) the Leibniz rule, D X (fY ) = X(f )Y + fD X Y.(2.28) The vector field D X Y is called the covariant derivative of Y in the direction X. A Riemannian metric determines a canonical connection by requiring zero torsion and compatibility with the metric. Theorem 22 (Levi-Civita connection [165, Thm. 3.11]). Every Riemannian man- ifold (M,g) admits a unique connection D satisfying, for all smooth vector fields X,Y,Z ∈ X(M): 19 2.4. Riemannian Geometry (1) torsion-freeness, D X Y − D Y X = [X,Y ],(2.29) (2) compatibility with the metric, X (⟨Y,Z⟩) =⟨D X Y,Z⟩ +⟨Y,D X Z⟩.(2.30) This connection is called the Levi-Civita connection. Example 23 (Euclidean Levi-Civita connection [165, Def. 3.8 and Lem. 3.14]). Let x 1 ,...,x n be the standard coordinates on R n , and let X = n X i=1 X i ∂ ∂x i , Y = n X j=1 Y j ∂ ∂x j (2.31) be smooth vector fields. The Levi-Civita connection of the Euclidean metric is D X Y = n X j=1 X(Y j ) ∂ ∂x j = n X i,j=1 X i ∂Y j ∂x i ∂ ∂x j .(2.32) Thus, D X Y is the ordinary directional derivative of the component functions of Y along X. In particular, the standard coordinate vector fields are parallel: D ∂ ∂x i ∂ ∂x j = 0,(2.33) for all i,j. The Levi-Civita connection is the canonical connection of Riemannian geometry. It determines geodesics, parallel transport, and curvature. It also induces a covariant derivative for vector fields along a curve. Proposition 24 (Induced covariant derivative [165, Prop. 3.18]). Let M be a manifold with Levi-Civita connection D, let α : I →M be a smooth curve, and let X(α) denote the space of smooth vector fields along α. There is a unique map X(α)→ X(α), V 7→ V ′ = D dt (V ),(2.34) called the induced covariant derivative along α, satisfying: 20 Chapter 2. Mathematical Background (1) (aV + bW ) ′ = aV ′ + bW ′ for V,W ∈ X(α) and a,b∈ R. (2) (fV ) ′ = df dt V + fV ′ for V ∈ X(α) and f ∈ C ∞ (I). (3) Let U ⊆ M be an open neighborhood of α(I), and let Y ∈ X(U ) be a smooth vector field on U. If V ∈ X(α) is the vector field along α obtained by restricting Y to the curve, namely V (t) = Y α(t) ∈ T α(t) M, then V ′ (t) = D ̇α(t) Y,(2.35) where the right-hand side denotes the covariant derivative of the field Y . (4) For V,W ∈ X(α), d dt ⟨V,W⟩ α(t) =⟨V ′ ,W⟩ α(t) +⟨V,W ′ ⟩ α(t) .(2.36) Geodesics are the Riemannian counterparts of straight lines. Definition 25 (Geodesic [165, Ch. 3]). Let M be a manifold with Levi-Civita connection D. A smooth curve γ : I →M is a geodesic if D ̇γ dt = D ̇γ ̇γ = 0.(2.37) Consequently, every geodesic has constant speed: ∥ ̇γ(t)∥ γ(t) is constant on I. Unlike straight lines in Euclidean space, geodesics on a general manifold need not be globally length-minimizing. They are locally length-minimizing curves [165, Lem. 5.14 and Prop. 5.16]. Geodesics also turn tangent vectors into manifold points through the exponential map, and locally turn nearby manifold points back into tangent vectors through the logarithmic map. Definition 26 (Exponential and logarithmic maps [165, Def. 3.29 and Prop. 3.30]). LetM be a manifold. For x∈M and v ∈ T x M, let γ x,v be the geodesic satisfying γ x,v (0) = x and ̇γ x,v (0) = v. The Riemannian exponential map at x is Exp x (v) = γ x,v (1),(2.38) where it is defined. As Exp x is locally invertible around 0∈ T x M, its local inverse 21 2.4. Riemannian Geometry is called the Riemannian logarithmic map at x and is denoted by Log x . This pair is used repeatedly in Riemannian neural networks to move between non- linear manifold-valued data and linear tangent-space computations. To compare or aggregate tangent vectors based at different manifold points, tangent vectors must be transported along curves. Definition 27 (Parallel transport [165, Prop. 3.19]). Let α : [a,b] → M be a smooth curve and let V (t) be a vector field along α. The field V is parallel along α if DV dt = 0(2.39) for all t ∈ [a,b]. Given v ∈ T α(a) M, the endpoint V (b) ∈ T α(b) M of the unique parallel vector field satisfying V (a) = v is called the parallel transport of v along α. When the connecting curve is the relevant geodesic from x to y, we denote parallel transport by PT x→y : T x M→ T y M. Proposition 28 (Parallel transport is an isometry [165, Lem. 3.20]). Let V (t) and W (t) be parallel vector fields along a smooth curve α : [a,b]→M. Then ⟨V (t),W (t)⟩ α(t) (2.40) is constant in t. Consequently, parallel transport along α defines a linear isometry between tangent spaces. The connection also measures the failure of second covariant derivatives to commute, which is encoded by the curvature tensor. Definition 29 (Curvature [165, Lem. 3.35]). LetM be a manifold with Levi-Civita connection D. The Riemannian curvature tensor is the map R : X(M)× X(M)× X(M)→ X(M)(2.41) defined by R(X,Y )Z = D X D Y Z− D Y D X Z− D [X,Y ] Z,(2.42) 22 Chapter 2. Mathematical Background where [X,Y ] denotes the Lie bracket of vector fields: [X,Y ](f ) = X(Y (f ))− Y (X(f )), f ∈ C ∞ (M).(2.43) Sectional curvature extracts from the curvature tensor the curvature of each two- dimensional tangent plane. Definition 30 (Sectional curvature [165, Lem. 3.39]). Let M be a manifold with curvature tensor R. For a two-dimensional subspace σ ⊂ T x M spanned by linearly independent vectors u,v ∈ T x M, the sectional curvature of σ is K x (σ) = ⟨R(u,v)v,u⟩ x ⟨u,u⟩ x ⟨v,v⟩ x −⟨u,v⟩ 2 x .(2.44) It is independent of the choice of basis u,v for σ. Positive, zero, and negative curvature correspond to spherical, Euclidean, and hy- perbolic behavior, which will be discussed later. A central consequence of non-positive curvature is that the exponential map can become globally invertible under suitable assumptions. Theorem 31 (Cartan–Hadamard theorem [165, Thm. 10.22]). Let M be a com- plete, connected, and simply connected manifold whose sectional curvature is every- where non-positive. Then, for every x ∈ M, the exponential map Exp x : T x M → M is a global diffeomorphism. This theorem explains why several geometries in machine learning are algorithmi- cally convenient, as their exponential maps allow tangent-space computations to be performed globally on the manifold. The distance also supports averaging manifold-valued data through Fréchet means. Definition 32 (Fréchet mean and variance [171, Sec. 2]). Let (X, d) be a metric space and let x 1 ,...,x N ∈ X with weights w i > 0 satisfying P N i=1 w i = 1. A weighted Fréchet mean is any minimizer WFM(w i ,x i )∈ argmin z∈X N X i=1 w i d(x i ,z) 2 .(2.45) When w i = 1 N for all i, it is called the Fréchet mean, and we write FM(x i ) for 23 2.4. Riemannian Geometry Geometric conceptEuclidean special case Riemannian metric ⟨u,v⟩ x =⟨u,v⟩. Geodesic distanced(x,y) =∥x− y∥. Geodesicγ(t) = (1− t)x + ty. Exponential mapExp x (v) = x + v. Logarithmic mapLog x (y) = y− x. Parallel transportPT x→y = id R n . Sectional curvatureK ≡ 0 on R n . Weighted Fréchet meanWFM(w i ,x i ) = P N i=1 w i x i . Fréchet varianceVar = min z P N i=1 w i ∥x i − z∥ 2 . Table 2.3: Euclidean prototypes for Riemannian-geometric concepts. the corresponding minimizer. The infimum of the objective is called the Fréchet variance. The objective above is defined on any metric space. On a manifold, the distance is usually the geodesic distance. If the data lie in a sufficiently small geodesic ball, the weighted Fréchet mean exists and is unique [2, Thm. 2.1]. Pullback metrics allow a complex manifold to inherit a computationally convenient geometry from a simpler prototype space. Definition 33 (Pullback and Riemannian isometry [165, Defs. 2.8 and 3.6]). Let M and N be smooth manifolds, let g be a Riemannian metric on N , and let f : M → N be smooth. The pullback of g by f is the symmetric tensor field f ∗ g on M defined by (f ∗ g) p (v,w) = g f (p) (f ∗,p (v),f ∗,p (w)), p∈M, v,w ∈ T p M.(2.46) If f ∗ g is positive definite at every point, it is a Riemannian metric on M. In particular, when f is a diffeomorphism andM is equipped with the pullback metric f ∗ g, the map f : (M,f ∗ g)→ (N,g) is a Riemannian isometry. A Riemannian isometry is the Riemannian version of a diffeomorphism: it preserves the metric, lengths, geodesic distances, geodesics, exponential maps, logarithmic maps, parallel transport, curvature, and distance-based quantities such as Fréchet means and variances [165, Ch. 3]. Tab. 2.3 summarizes the Euclidean special cases of the above geometric concepts. 24 Chapter 2. Mathematical Background 2.5 Metric Geometry Metric geometry extends Riemannian geometry to the more general setting of metric spaces, where no differentiable structure is assumed. 1 The theory develops geodesics and curvature without smooth structure, providing the metric tools used later for hyperbolic neural layers. We begin by recalling basic notions in metric spaces. Definition 34 (Metric space [29, Def. I.1.1]). A metric space is a pair (X,d) where X is a nonempty set and d :X ×X → R satisfies, for all x,y,z ∈X : (1) d(x,y)≥ 0 and d(x,y) = 0 if and only if x = y, (2) d(x,y) = d(y,x), (3) d(x,z)≤ d(x,y) + d(y,z). A connected Riemannian manifold becomes a metric space when equipped with its geodesic distance [165, Prop. 5.18]. The metric-space perspective abstracts away coordinates and focuses on distances, geodesics, and comparison geometry. Geodesics, rays, and lines generalize unit-speed minimizing geodesics to metric spaces. Definition 35 (Geodesic, geodesic ray, and geodesic line [29, Ch. I.1]). Let (X,d) be a metric space. A geodesic joining x to y is a continuous map γ : [0,l] → X with γ(0) = x and γ(l) = y such that d (γ(t),γ(t ′ )) =|t− t ′ |, t,t ′ ∈ [0,l].(2.47) A geodesic ray is a continuous map γ : [0,∞)→X such that d (γ(t),γ(t ′ )) =|t−t ′ | for all t,t ′ ≥ 0. A geodesic line is a continuous map γ : R → X such that d (γ(t),γ(t ′ )) =|t− t ′ | for all t,t ′ ∈ R. Geodesic metric spaces abstract the requirement that every pair of points be joined by a distance-realizing geodesic. Definition 36 (Geodesic metric space [29, Ch. I.1]). The metric space (X,d) is a geodesic metric space, or more briefly a geodesic space, if every pair of points in X is joined by a geodesic. It is uniquely geodesic if there is exactly one geodesic 1 Here, the missing differentiable structure refers specifically to a smooth manifold structure. Metric analogues of geodesics and curvature can still be defined without it. 25 2.5. Metric Geometry joining x to y for all x,y ∈X. Convexity is defined through geodesics, generalizing linear convexity in Euclidean space and geodesic convexity on manifolds. Definition 37 (Convex subset [29, Ch. I.1]). Let (X,d) be a metric space. A subset C ⊆ X is convex if every pair x,y ∈ C can be joined by a geodesic in X and the image of every such geodesic is contained in C. We next review concepts that extend curvature from manifolds to metric spaces. The reference spaces are the model spaces of constant curvature. Definition 38 (Model space [29, Ch. I.2]). For K ∈ R, the model space (M n K ,d K ) is given by (M n K ,d K ) = S n , 1 √ K d , K > 0, (R n ,d),K = 0, L n , 1 √ −K d , K < 0, (2.48) where d is the geodesic distance in the corresponding manifold. The unit sphere S n and unit Lorentz manifold L n are reviewed in Sec. 2.9.5. The diameter of M 2 K is denoted by D K = π/ √ K, K > 0, ∞,K ≤ 0. (2.49) Definition 39 (Comparison triangle [29, Lem. I.2.14 and Sec. I.1]). Let (X,d) be a geodesic metric space and let △(x,y,z) be a geodesic triangle in X with side lengths a = d(y,z), b = d(x,z), and c = d(x,y). A comparison triangle for △(x,y,z) in M 2 K is a triangle △( ̄x, ̄y, ̄z) such that d K ( ̄y, ̄z) = a, d K ( ̄x, ̄z) = b, and d K ( ̄x, ̄y) = c. When a + b + c < 2D K , the comparison triangle exists. If p lies on the side from x to y, then a comparison point for p is the point ̄p on the side from ̄x to ̄y such that d(x,p) = d K ( ̄x, ̄p) and d(y,p) = d K ( ̄y, ̄p). Comparison points on the other two sides are defined analogously. CAT(K) spaces encode curvature through triangle comparison with the model plane M 2 K . Intuitively, a CAT(K) space is a metric space whose triangles are thinner than the corresponding comparison triangles in M 2 K . 26 Chapter 2. Mathematical Background Definition 40 (CAT(K) space [29, Def. I.1.1]). Let (X,d) be a metric space and let K ∈ R. Let ∆ be a geodesic triangle in X with perimeter less than 2D K , and let ̄ ∆⊂ M 2 K be a comparison triangle for ∆. The triangle ∆ satisfies the CAT(K) inequality if, for all p,q ∈ ∆ and all corresponding comparison points ̄p, ̄q ∈ ̄ ∆, d(p,q)≤ d K ( ̄p, ̄q).(2.50) Then X is called a CAT(K) space as follows: (1) If K ≤ 0, then X is a CAT(K) space if X is a geodesic space all of whose geodesic triangles satisfy the CAT(K) inequality. (2) If K > 0, then X is a CAT(K) space if X is D K -geodesic and all geodesic triangles in X of perimeter less than 2D K satisfy the CAT(K) inequality. Here, D K -geodesic means that for every pair of points x,y ∈X with d(x,y) < D K , there is a geodesic joining x to y. Definition 41 (Hadamard space [29, p. 159]). A Hadamard space is a complete CAT(0) space. In particular, a Hadamard manifold, a complete, simply connected Riemannian manifold with non-positive sectional curvature, is a Hadamard space [29, Thm. I.4.1]. Proposition 42 (Orthogonal projection [29, Prop. I.2.4]). Let (X,d) be a CAT(0) space and let C ⊆X be a convex subset that is complete in the induced metric. For every x∈X, there exists a unique point π C (x)∈ C such that d (x,π C (x)) = inf y∈C d(x,y) = d(x,C).(2.51) If x ′ belongs to the geodesic segment [x,π C (x)], then π C (x ′ ) = π C (x).(2.52) The map π C :X → C is called the orthogonal projection, or simply the projection. With geodesics, one can define asymptotic rays, boundary points, Busemann func- tions, horoballs, and horospheres in metric spaces. 27 2.5. Metric Geometry Definition 43 (Asymptotic rays and boundary points [29, Def. I.8.1]). Let (X,d) be a metric space. Two geodesic rays γ,η : [0,∞)→X are asymptotic if there exists C ≥ 0 such that d (γ(t),η(t))≤ C for all t≥ 0.(2.53) The set ∂X of boundary points, also called points at infinity or ideal points, is the set of equivalence classes of geodesic rays, where two geodesic rays are equivalent if and only if they are asymptotic. Asymptotic rays generalize parallel lines. In hyperbolic geometry, they point to- ward the same ideal boundary point, which becomes a direction for defining Busemann functions. Definition 44 (Busemann function, horoball, and horosphere [29, Def. I.8.17]). Let (X,d) be a metric space and let γ : [0,∞)→X be a geodesic ray. If the limit exists, the Busemann function associated with γ is B γ (x) = lim t→∞ (d (x,γ(t))− t), x∈X.(2.54) For τ ∈ R, the sublevel set HB γ τ =x∈X | B γ (x)≤ τ(2.55) is a horoball, and the level set H γ τ =x∈X | B γ (x) = τ(2.56) is a horosphere. In Hadamard spaces, the Busemann limit exists [29, Lem. I.8.18]. Moreover, Buse- mann functions associated with asymptotic rays agree up to an additive constant. Corollary 45 (Busemann functions of asymptotic rays [29, Cor. I.8.20]). If X is a Hadamard space, then the Busemann functions associated with asymptotic rays in X are equal up to addition of a constant. In Euclidean space, if γ(t) = tv with ∥v∥ = 1, then B γ (x) = −⟨x,v⟩. Thus the Busemann function provides an intrinsic generalization of the Euclidean inner prod- uct, up to sign. Therefore, horospheres generalize Euclidean hyperplanes. Tab. 2.4 28 Chapter 2. Mathematical Background Metric-geometric conceptEuclidean special case Geodesic, geodesic ray, and geodesic lineStraight segment, ray, and line. Asymptotic geodesic raysParallel rays. Busemann function −B γ (x)Inner product ⟨x,v⟩. Horosphere H γ τ Hyperplane. Horospheres associated with asymptotic raysParallel hyperplanes. Table 2.4: Euclidean prototypes for metric-geometric concepts. summarizes the Euclidean special cases of the metric-geometric concepts. 2.6 Algebraic Structures on Manifolds Euclidean neural networks rely on vector addition and scalar multiplication to combine, translate, and rescale features. Manifold-valued representations generally lack a global linear structure, which motivates algebraic structures on manifolds that play analogous roles. This section reviews groups, Lie groups, gyrogroups and gyrovector spaces, which generalize addition, subtraction, and scalar multiplication to curved spaces. Definition 46 (Group [128, Ch. I.2]). A group is a nonempty set G equipped with a binary operation a ⊕ : G× G→ G satisfying, for all x,y,z ∈ G: (1) associativity, x⊕ (y⊕ z) = (x⊕ y)⊕ z, (2) identity, there exists e∈ G, called the identity element or neutral element, such that e⊕ x = x⊕ e = x, (3) inverse, for every x∈ G, there exists⊖x∈ G such that (⊖x)⊕x = x⊕(⊖x) = e. a The group operation is usually written multiplicatively and called multiplication, as it is usually noncommutative. For a commutative group, it is often written additively and called addition. In this thesis, we use ⊕ for simplicity. Definition 47 (Abelian group [128, Ch. I.2]). A group (G,⊕) is an abelian group, or commutative group, if x⊕ y = y⊕ x for all x,y ∈ G. Lie groups are simultaneously algebraic and geometric objects. They add smooth structure to groups and require the group operations to be smooth. 29 2.6. Algebraic Structures on Manifolds Definition 48 (Lie group [197, Def. 6.20]). A Lie group is a smooth manifold G equipped with a binary operation ⊕ : G× G→ G satisfying: (1) (G,⊕) is a group, (2) the multiplication map (x,y)7→ x⊕ y is smooth, (3) the inversion map x7→⊖x is smooth. Definition 49 (Abelian Lie group). A Lie group (G,⊕) is an abelian Lie group, or commutative Lie group, if its underlying group is abelian, namely x⊕y = y⊕x for all x,y ∈ G. Example 50 (Matrix Lie groups). The general linear group GL(n) introduced in Thm. 17 is a Lie group under matrix multiplication. Standard matrix Lie subgroups of GL(n) include the special linear group, the orthogonal group, and the special orthogonal group: SL(n) =A∈ GL(n)| det(A) = 1, O(n) = A∈ GL(n)| A ⊤ A = I n , SO(n) = O(n)∩ SL(n). (2.57) Gyrogroups relax the associativity axiom of groups. The deviation from associativity is controlled by gyrations through the gyroassociative law. Definition 51 (Gyrogroup [200, Def. 2.7]). Given a nonempty set G with a binary operation⊕ : G×G→ G, (G,⊕) forms a gyrogroup if its binary operation satisfies the following axioms for any x,y,z ∈ G: (1) Left identity: there exists at least one element e ∈ G, called a left identity or neutral element, such that e⊕ x = x. (2) Left inverse: there exists an element ⊖x ∈ G, called a left inverse of x, such that ⊖x⊕ x = e. (3) Left gyroassociative law: there exists an automorphism gyr[x,y] : G → G for each x,y ∈ G such that x⊕ (y⊕ z) = (x⊕ y)⊕ gyr[x,y]z.(2.58) The automorphism gyr[x,y] is called the gyroautomorphism, or the gyration 30 Chapter 2. Mathematical Background of G generated by x,y. (4) Left reduction law: gyr[x,y] = gyr[x⊕ y,y].(2.59) If every gyration is the identity map, the gyroassociative law reduces to ordinary associativity. Hence, gyrogroups naturally generalize groups. Definition 52 (Gyrocommutative gyrogroup [200, Def. 2.8]). A gyrogroup (G,⊕) is gyrocommutative if it satisfies x⊕ y = gyr[x,y](y⊕ x) (gyrocommutative law).(2.60) Gyrocommutativity replaces commutativity. It says that exchanging two operands is possible after applying the appropriate gyration. Similarly, a gyrovector space gen- eralizes a vector space. Definition 53 (Gyrovector space [54]). A gyrocommutative gyrogroup (G,⊕) equipped with a scalar gyromultiplication ⊙ : R × G → G is called a gyrovec- tor space if it satisfies the following axioms for s,t∈ R and x,y,z ∈ G: (1) Identity scalar multiplication: 1⊙ x = x.(2.61) (2) Scalar distributive law: (s + t)⊙ x = s⊙ x⊕ t⊙ x.(2.62) (3) Scalar associative law: (st)⊙ x = s⊙ (t⊙ x).(2.63) (4) Gyroautomorphism: gyr[x,y](t⊙ z) = t⊙ gyr[x,y]z.(2.64) (5) Identity gyroautomorphism: gyr[s⊙ x,t⊙ x] = id,(2.65) 31 2.6. Algebraic Structures on Manifolds where id is the identity map. Remark 54. Nguyen [157, Def. 2.3] presented a similar definition, except that the identity scalar multiplication axiom also includes 0⊙x = t⊙e = e and (−1)⊙x = ⊖x. As implied by Ungar [200, Thm. 6.4], these conditions are redundant. Just as an inner product space augments a vector space with a compatible inner product, a real inner product gyrovector space augments a gyrovector space with an ambient inner product and corresponding compatibility axioms. Definition 55 (Real inner product gyrovector space [200, Def. 6.2]). Let (G,⊕,⊙) be a gyrovector space and let ⟨·,·⟩ denote the Euclidean inner product on R n with associated norm ∥·∥. We call (G,⊕,⊙,⟨·,·⟩) a real inner product gyrovector space if the following conditions hold. (1) G⊆ R n and inherits the inner product ⟨·,·⟩ and norm ∥·∥. (2) Inner product gyroinvariance: ⟨gyr[x,y]u, gyr[x,y]v⟩ =⟨u,v⟩, ∀x,y,u,v ∈ G.(2.66) (3) Scaling property: |s|⊙ x ∥s⊙ x∥ = x ∥x∥ , ∀x∈ G\0, ∀s∈ R\0.(2.67) (4) Let ∥G∥ = ±∥x∥ | x ∈ G ⊂ R. The set ∥G∥ forms a one-dimensional real vector space with respect to the vector addition and scalar multiplication induced by ⊕ and ⊙ on G. (5) Homogeneity property: ∥s⊙ x∥ =|s|⊙∥x∥, ∀x∈ G, ∀s∈ R.(2.68) (6) Gyrotriangle inequality: ∥x⊕ y∥≤∥x∥⊕∥y∥, ∀x,y ∈ G.(2.69) 32 Chapter 2. Mathematical Background Definition 56 (Gyrovector space isomorphisms [200, Def. 6.89]). Let (G 1 ,⊕ 1 ,⊙ 1 ) and (G 2 ,⊕ 2 ,⊙ 2 )(2.70) be real inner product gyrovector spaces. A map φ : G 1 → G 2 is a gyrovector space isomorphism if it is bijective and satisfies φ(x⊕ 1 y) = φ(x)⊕ 2 φ(y), ∀x,y ∈ G 1 ,(2.71) φ(t⊙ 1 x) = t⊙ 2 φ(x), ∀x∈ G 1 ,∀t∈ R,(2.72) and preserves the inner product of unit gyrovectors, ⟨φ(x),φ(y)⟩ ∥φ(x)∥φ(y)∥ = ⟨x,y⟩ ∥x∥y∥ , ∀x,y ∈ G 1 with x̸= 0,y ̸= 0.(2.73) A useful property is that gyrovector space isomorphisms preserve the gyration, in- verse, and identity. Proposition 57. [↓] Let (G 1 ,⊕ 1 ,⊙ 1 ) and (G 2 ,⊕ 2 ,⊙ 2 ) be real inner product gy- rovector spaces with gyrations gyr 1 and gyr 2 , respectively. If φ : G 1 → G 2 is a gyrovector space isomorphism, then for all x,y,z ∈ G 1 , φ (gyr 1 [x,y]z) = gyr 2 [φ(x),φ(y)]φ(z),(2.74) φ(e 1 ) = e 2 ,(2.75) φ(⊖ 1 x) =⊖ 2 φ(x),(2.76) where e 1 and e 2 are the gyro identities in G 1 and G 2 , respectively. Intuitively, a gyrovector space generalizes a vector space to curved spaces. The gyro- operations associated with a manifoldM can be defined as follows. Given a predefined origin e ∈ M, we assume that the relevant exponential maps, logarithmic maps, and parallel transports are well-defined. Following Nguyen and Yang [159, Eqs. (1)–(3)], for x,y,z ∈M and t∈ R, these operations are defined as x⊕ y = Exp x (PT e→x (Log e (y))),(2.77) t⊙ x = Exp e (t Log e (x)),(2.78) ⊖x = Exp e (− Log e (x)),(2.79) 33 2.7. Riemannian Optimization gyr[x,y]z = (⊖(x⊕ y))⊕ (x⊕ (y⊕ z)).(2.80) The corresponding gyro inner product, gyronorm, and gyrodistance are ⟨x,y⟩ gyr =⟨Log e (x), Log e (y)⟩ e ,(2.81) ∥x∥ gyr = q ⟨x,x⟩ gyr ,(2.82) d gyr (x,y) =∥⊖x⊕ y∥ gyr .(2.83) If the above operations satisfy the axioms in Thm. 53, then (M,⊕,⊙) is a gyrovector space. In Euclidean space, the above definitions reduce to the ordinary vector-space structure. 2.7 Riemannian Optimization This section briefly reviews Riemannian optimization, which addresses the following manifold-constrained optimization problems: min w∈M f (w).(2.84) Here, M is a Riemannian manifold and f :M→ R is a smooth objective function. Riemannian gradient [165, Def. 3.47]. The Riemannian gradient of f at w ∈ M, denoted by grad w f ∈ T w M, is the unique tangent vector satisfying ⟨grad w f,v⟩ w = f ∗,w (v), ∀v ∈ T w M.(2.85) Eq. (2.85) characterizes the Riemannian gradient through the directional derivative f ∗,w (v). This is the same characterization as in Euclidean space, where the Euclidean gradient satisfies f ∗,w (v) = ⟨∇ w f,v⟩ for w,v ∈ R n . As in R n , the Cauchy–Schwarz inequality shows that, among all unit tangent directions, the directional derivative is minimized by − grad w f/∥grad w f∥ w whenever grad w f ̸= 0. Therefore, − grad w f is the direction of steepest descent. Riemannian optimizers. To solve the objective in Eq. (2.84), we take Stochastic Gradient Descent (SGD) [175] as an example to show how to generalize a Euclidean optimizer to the Riemannian setting [1]. At iteration t, let f t denote a stochastic or mini-batch objective associated with f. When M = R m , Euclidean SGD updates an 34 Chapter 2. Mathematical Background iterate w (t) ∈ R m as w (t+1) = w (t) − α t ∇ w (t) f t ,(2.86) where α t > 0 is the learning rate and ∇ w (t) f t is the Euclidean gradient of f t . On a general manifold, the descent direction must belong to the tangent space at the current iterate, and ordinary vector addition cannot return that direction to the manifold. The corresponding Riemannian Stochastic Gradient Descent (RSGD) [25, Sec. 2.3] update is w (t+1) = Exp w (t) (−α t grad w (t) f t ),(2.87) whenever the exponential map is defined for this step. In practical algorithms, a re- traction can replace the exact exponential map by a first-order approximation [27, Def. 3.47 and Sec. 4.3], but a systematic treatment of retractions is beyond the scope of this chapter. A widely used package is the Geoopt [125], which integrates manifold- valued parameters and Riemannian optimizers into PyTorch, supporting RSGD and adaptive optimizers such as Riemannian Adam [18]. In-depth discussions of Rieman- nian optimization can be found in Absil et al. [1], Boumal [27]. Trivialization. Apart from solving Eq. (2.84) directly on the manifold, another approach is to parameterize manifold-valued variables with Euclidean parameters and thereby use Euclidean optimization, which is called trivialization [132]. Let φ : R m → M be a smooth surjective map, let z ∈ R m be a trainable Euclidean parameter, and set w = φ(z)∈M. The constrained objective can then be written as min z∈R m f (φ(z)).(2.88) In practice, the map φ could be the exponential map or retraction. 2.8 Matrix Functions This section reviews some matrix functions widely used on matrix manifolds. 2.8.1 Matrix Functions and Differentials We denote the Euclidean space of n× n real symmetric matrices by S n and the SPD manifold of n × n SPD matrices by S n ++ . Let ̊ I be an open interval of R and let f : ̊ I → R be a smooth function. For any symmetric matrix S whose eigenvalues lie in 35 2.8. Matrix Functions ̊ I, the associated symmetric matrix function is defined by f : S 7−→ Uf (Σ)U ⊤ ∈S n , with S = U ΣU ⊤ as the eigendecomposition.(2.89) Its differential is known as the Daleckii–Krein formula: f ∗,S (V ) = U L f ⊛ U ⊤ V U U ⊤ , ∀V ∈S n ,(2.90) [L f ] i,j = f (σ i )−f (σ j ) σ i −σ j , if σ i ̸= σ j f ′ (σ i ),otherwise (2.91) where L f is called the Loewner matrix, its (i,j)-th entry is defined in Eq. (2.91), and ⊛ denotes the Hadamard product. Three special cases are the matrix logarithm log : S n ++ → S n , the matrix exponential exp : S n → S n ++ , which is its inverse, and the matrix power map P θ : S n ++ → S n ++ defined by P θ (P ) = P θ . See Bhatia [20, Eqs. (2.38)–(2.40)] or Bhatia [21, Thm. V.3.3] for more details. Let L n ++ be the Cholesky space of lower triangular matrices with positive diagonal entries. The Cholesky map is Chol : S n ++ → L n ++ with inverse Chol −1 (L) = L ⊤ . Let P ∈ S n ++ , L = Chol(P ), V ∈ T P S n ++ , and X ∈ T L L n ++ . For any square matrix A, set A 1/2 = ⌊A⌋ + 1 2 D(A), where ⌊A⌋ is the strictly lower triangular part and D(A) is the diagonal matrix formed from the diagonal of A. As shown by Lin [137, Prop. 4], the differentials of Chol and Chol −1 are Chol ∗,P (V ) = L L −1 V L −⊤ 1/2 ,Chol −1 ∗,L (X) = XL ⊤ + LX ⊤ .(2.92) 2.8.2 Backpropagation Through Matrix Functions Symmetric matrix functions. The above symmetric matrix functions can be back- propagated using the Daleckii–Krein formula. Although PyTorch supports automatic differentiation through eigendecomposition [169], its backward pass requires computing (σ i −σ j ) −1 [110, Prop. 1], which may trigger numerical instability when two eigenvalues approach each other. Following Brooks et al. [31, Eq. (13)], we instead use the Daleckii– Krein expression in Eqs. (2.90) and (2.91) for the backward pass [21, Thm. V.3.3]. As shown in Eq. (2.91), the divided difference converges to the derivative f ′ (σ i ) when two eigenvalues approach each other, making the expression numerically more stable. Cholesky decomposition. The backpropagation of the Cholesky decomposition has been studied by Murray [154]. In the experiments, we use torch.linalg.cholesky 36 Chapter 2. Mathematical Background and its automatic differentiation. 2.9 Example Manifolds This section collects the concrete model spaces that recur in later chapters. Each man- ifold is first described by its underlying set and representation, and then summarized through the Riemannian and algebraic operators. 2.9.1 Symmetric Positive Definite Manifolds The SPD manifold has shown great success in diverse applications [106, 37, 205, 141, 48, 123]. We denote the set of n×n SPD matrices byS n ++ and the vector space of n×n real symmetric matrices by S n . As shown by Arsigny et al. [9], S n ++ forms an open submanifold of the Euclidean space S n . We review five common Riemannian metrics on S n ++ : the affine-invariant metric (AIM) [171], the log-Euclidean metric (LEM) [9], the power-Euclidean metric (PEM) [70], the log-Cholesky metric (LCM) [137], and the Bures–Wasserstein metric (BWM) [22]. LEM, AIM, and PEM are represented by the parameterized families (α,β)-LEM, (α,β)-AIM, and (θ,α,β)-EM, respectively. Their common (α,β) parameters refer to the following O(n)-invariantinner product on S n [196]: ⟨A,B⟩ (α,β) = α⟨A,B⟩ + β tr(A) tr(B),(2.93) where A,B ∈ S n and (α,β) ∈ ST = (α,β) ∈ R 2 | min(α,α + nβ) > 0. The standard LEM and AIM are recovered from (α,β)-LEM and (α,β)-AIM, respectively, at (α,β) = (1, 0). Likewise, the standard θ-PEM is recovered from (θ,α,β)-EM at (α,β) = (1, 0). The (θ,α,β)-EM formulas below assume θ ̸= 0; their limit as θ → 0 is (α,β)-LEM. Tabs. 2.5 and 2.6 summarize the associated operators under these five geometries with the following notation. (1) General notation. Let P,Q ∈ S n ++ be SPD matrices, let V,W ∈ T P S n ++ be tangent vectors, and let P i N i=1 ⊂ S n ++ be a dataset. The norms induced by ⟨·,·⟩ (α,β) and the standard Frobenius inner product⟨·,·⟩ are denoted by∥·∥ (α,β) and ∥·∥, respectively. 2 2 We use ∥·∥ for both the Frobenius norm of matrices and the ℓ 2 norm of vectors, since both norms are induced by the standard Euclidean inner product on the ambient Euclidean space. 37 2.9. Example Manifolds (2) Symmetric matrix functions. Symmetric matrix functions and their differen- tials are reviewed in Sec. 2.8.1. In addition to the matrix logarithm and exponential recorded there, the SPD operator tables use the matrix power map P θ (P ) = P θ , whose differential is obtained by specializing Eq. (2.90). (3) LCM. The Cholesky map, its inverse, the half-diagonal operator, and their dif- ferentials are reviewed in Sec. 2.8.1. Let L = Chol(P ) and K = Chol(Q). In the LCM column, the corresponding tangent vectors are X = Chol ∗,P (V ) and Y = Chol ∗,P (W ). We define the Log-Cholesky map by ψ LC (P ) =⌊L⌋+Dlog(D(L)), where Dlog(·) is the diagonal element-wise logarithm. The diagonal matrices K, L, X, and Y contain the diagonal entries of K, L, X, and Y , respectively. (4) Gyrovector. The three geometries in Tab. 2.5 also admit gyrovector spaces. As shown by Nguyen [157, Sec. 3.1] and Nguyen and Yang [159, Sec. 2.4], AIM, LEM, and LCM each induce a gyro-structure via Eqs. (2.77) to (2.83). For LEM and LCM, the gyrovector spaces reduce to vector spaces: P ⊕ LE Q = exp (log(P ) + log(Q)), t⊙ LE P = exp (t log(P )), P ⊕ LC Q = ψ −1 LC (ψ LC (P ) + ψ LC (Q)), t⊙ LC P = ψ −1 LC (tψ LC (P )). (2.94) The above two vector additions are the Lie group operations in Tab. 2.5. For AIM, the gyroaddition and scalar gyromultiplication are P ⊕ AI Q = P 1/2 QP 1/2 , t⊙ AI P = P t .(2.95) In particular, the AIM gyroaddition above is not the AIM Lie group operation; the latter is listed in Tab. 2.5. (5) BWM. The Lyapunov operator L P [V ] is defined by L P [V ]P + PL P [V ] = V.(2.96) For BWM parallel transport, we only present the case where P and Q are com- muting matrices, namely P = U ΣU ⊤ and Q = U ∆U ⊤ , with Σ = diag(σ i ) and ∆ = diag(δ i ). (6) Completeness. BWM and (θ,α,β)-EM with θ ̸= 0 are geodesically incomplete, meaning that their exponential maps are not defined globally. For BWM, the exponential map is locally defined only for tangent vectors satisfying L P [V ] + I n ∈ 38 Chapter 2. Mathematical Background Metric(α,β)-LEM(α,β)-AIMLCM Q⊕ Pexp (log(P ) + log(Q))KPK ⊤ Chol −1 (⌊L + K⌋ + KL) g P (V,W ) log ∗,P (V ), log ∗,P (W ) (α,β) ⟨P −1 V,WP −1 ⟩ (α,β) ⟨⌊X⌋,⌊Y⌋⟩ +⟨XL −1 , YL −1 ⟩ d(P,Q)∥log(P )− log(Q)∥ (α,β) log(Q −1/2 PQ −1/2 ) (α,β) ∥ψ LC (P )− ψ LC (Q)∥ FM(P i )exp 1 N P N i=1 log(P i ) Karcher flowψ −1 LC 1 N P N i=1 ψ LC (P i ) Log P Q(log ∗,P ) −1 [log(Q)− log(P )]P 1/2 log(P −1/2 QP −1/2 )P 1/2 Chol −1 ∗,L [⌊K⌋−⌊L⌋ + L Dlog(L −1 K)] Exp P Vexp log(P ) + log ∗,P (V ) P 1/2 exp P −1/2 V P −1/2 P 1/2 Chol −1 (⌊L⌋ +⌊X⌋ + L exp (XL −1 )) γ(t;P,Q)exp [log(P ) + t(log(Q)− log(P ))]P 1/2 (P −1/2 QP −1/2 ) t P 1/2 Chol −1 n ⌊L⌋ + t(⌊K⌋−⌊L⌋) + K t L t−1 o References[9, 196]; Thm. 147[171, 196][137]; Thm. 147 Table 2.5: Lie group structures and associated Riemannian operators on S n ++ . Metric(θ,α,β)-EMBWM g P (V,W ) 1 θ 2 ⟨(P θ ) ∗,P (V ), (P θ ) ∗,P (W )⟩ (α,β) 1 2 ⟨L P [V ],W⟩ d(P,Q) 1 |θ| Q θ − P θ (α,β) tr(P ) + tr(Q)− 2 tr (PQ) 1/2 1/2 Log P Q[(P θ ) ∗,P ] −1 Q θ − P θ (PQ) 1/2 + (QP ) 1/2 − 2P PT P→Q (V )[(P θ ) ∗,Q ] −1 ◦ (P θ ) ∗,P (V )U h q δ i +δ j σ i +σ j U ⊤ V U ij i U ⊤ Exp P V P θ + (P θ ) ∗,P (V ) 1/θ P + V +L P [V ]PL P [V ] γ(t;P,Q) (1− t)P θ + tQ θ 1/θ A t A ⊤ t A t = (1− t)P 1/2 + tP −1/2 P 1/2 QP 1/2 1/2 References[70, 196][22, 196] Table 2.6: Riemannian operators of (θ,α,β)-EM and BWM on S n ++ . S n ++ . For (θ,α,β)-EM, the exponential map is locally defined only for tangent vectors satisfying P θ + (P θ ) ∗,P (V )∈S n ++ . (7) Fréchet mean. Karcher flow [116] refers to the standard iterative solver for the Fréchet mean objective in Thm. 32. 2.9.2 Full-Rank Correlation Manifolds Given a covariance matrix Σ, its correlation matrix is defined as C = Cor(Σ) = D(Σ) − 1 /2 ΣD(Σ) − 1 /2 ,(2.97) where D(·) extracts the diagonal part of Σ as a diagonal matrix. This diagonal nor- malization yields a scale-invariant representation: for any positive diagonal matrix D, Cor(DΣD) = Cor(Σ). Hence, correlation matrices remove marginal scales and empha- 39 2.9. Example Manifolds Figure 2.1: The black stars denote 2× 2 correlation matrices, while the red, green, and blue dots denote corresponding SPD matrices. The black dots denote the boundary of the SPD cone. size pairwise dependencies rather than raw variances. Only recently have Riemannian structures been developed for correlation matrices. The space of n×n full-rank correlation matrices, denoted by Cor + (n), forms a Rieman- nian manifold and can be identified as a quotient manifold of the SPD manifold [62, Thm. 1]. As illustrated in Fig. 2.1, each correlation matrix corresponds to a surface in the SPD manifold. However, this quotient geometry does not guarantee uniqueness or closed forms of the Riemannian logarithm and Fréchet mean [195, Sec. 1.1]. To ad- dress this limitation, recent advances introduced five convenient Riemannian metrics on Cor + (n): the Euclidean–Cholesky metric (ECM) [195], log-Euclidean–Cholesky metric (LECM) [195], poly-hyperbolic–Cholesky metric (PHCM) [195], off-log metric (OLM) [191], and log-scaled metric (LSM) [191]. These metrics are pullback metrics from simpler prototype spaces: ECM, LECM, OLM, and LSM are induced from Euclidean spaces, while PHCM is induced from a product of hyperbolic open hemispheres. We first review the associated prototype spaces. (1) LT 1 (n) is the affine space of n× n lower triangular matrices with unit diagonal. (2) LT 0 (n) is the Euclidean space of n×n lower triangular matrices with null diagonal. (3) L n is the manifold of n× n lower triangular matrices with positive diagonals and unit row ℓ 2 -norm. (4) Hol(n) is the Euclidean space of n×n symmetric matrices with null diagonals. The 40 Chapter 2. Mathematical Background tangent space T C Cor + (n) at C ∈ Cor + (n) can be identified with Hol(n). (5) Row 0 (n) is the Euclidean space of n× n symmetric matrices with null row sum. ECM. It is derived from LT 1 (n) by Cor + (n) Θ=D(Chol(·)) −1 Chol(·) −⇀ ↽− Θ −1 =Cor◦ Chol −1 LT 1 (n),(2.98) where Θ(C) = D(Chol(C)) −1 Chol(C) for any C ∈ Cor + (n). Here, Chol(C) is the Cholesky decomposition C = Chol(C) Chol(C) ⊤ and D(·) returns a diagonal matrix consisting of the input diagonals. As LT 1 (n) = I n + LT 0 (n), ECM is essentially induced from the Euclidean space of LT 0 (n). Proposition 58 (ECM). Let φ EC (C) =⌊Θ(C)⌋, where ⌊·⌋ returns a strictly lower triangular matrix. ECM over Cor + (n) is the pullback metric from the Euclidean space LT 0 (n) by φ EC . LECM. It is defined by further pulling back ECM: Cor + (n) log◦Θ −⇀ ↽− (log◦Θ) −1 =Cor◦ Chol −1 ◦ exp LT 0 (n),(2.99) where log(·) : LT 1 (n) −→ LT 0 (n) is the matrix logarithm with the matrix exponential exp(·) as its inverse. OLM. It is derived from a permutation-invariant inner product over Hol(n) by Cor + (n) Log ◦ =off◦log −⇀ ↽− Exp ◦ Hol(n).(2.100) For any symmetric hollow matrix H ∈ Hol(n), the operator D(H) returns a unique diagonal matrix such that Exp ◦ (·) : Hol(n) ∋ H 7−→ exp (D(H) + H) ∈ Cor + (n) is a diffeomorphism. As shown by Archakov and Hansen [6, Cor. 1 and Sec. 5], D(H) can be computed by the following exponentially converging algorithm: D k+1 = D k − log (D (exp (D k + H))), with D 0 = 0 n×n as the zero matrix. LSM. It is derived from a permutation-invariant inner product over Row 0 (n) by Cor + (n) Log ⋆ −⇀ ↽− Exp ⋆ =Cor◦ exp Row 0 (n).(2.101) For any correlation matrix C ∈ Cor + (n), there exists a unique positive diagonal matrix D ⋆ (C) such that Log ⋆ (·) : Cor + (n)∋ C 7−→ log(D ⋆ (C)CD ⋆ (C))∈ Row 0 (n) is a diffeo- 41 2.9. Example Manifolds morphism. As shown by Thanwerdas [191, Sec. 3.5], D ⋆ (C) corresponds to the unique zero of f : x ∈ R n ++ 7−→ Cx− 1 x , where R n ++ denotes the set of n-dimensional posi- tive vectors and 1 x = 1 x 1 ,..., 1 x n . This equation can be solved by damped Newton’s method. Tab. 2.7 summarizes the associated prototype spaces and diffeomorphisms. Tabs. 2.8 and 2.9 summarize the vector operations and Riemannian operators for ECM, LECM, OLM, and LSM. We use the following notation and record the following remarks. (1) General notation. Let C,C ′ ∈ Cor + (n) be correlation matrices, let C i N i=1 ⊂ Cor + (n), and let V,W ∈ T C Cor + (n) ∼ = Hol(n) be tangent vectors. Let L = Chol(C). (2) ECM and LECM. For any K ∈ LT 1 (n) and X,ξ ∈ LT 0 (n), the maps and differentials involved in ECM and LECM are Θ(C) = D(L) −1 L,(2.102) Θ −1 (K) = D(K ⊤ ) − 1 2 K ⊤ D(K ⊤ ) − 1 2 ,(2.103) log(K) = n−1 X k=1 (−1) k−1 k (K− I n ) k ,(2.104) exp(ξ) = n−1 X k=0 1 k! ξ k ,(2.105) Θ ∗,C (V ) = Θ(C) L −1 V L −⊤ 1 2 − 1 2 D L −1 V L −⊤ Θ(C),(2.106) (Θ ∗,C ) −1 (ξ) = Lξ ⊤ − CD Lξ ⊤ D(L) + D(L) ξL ⊤ − D Lξ ⊤ C , (2.107) log ∗,K (ξ) = n−1 X k=1 (−1) k−1 k h (K− I n ) k−1 ξ +· + ξ (K− I n ) k−1 i ,(2.108) exp ∗,X (ξ) = n−1 X k=1 1 k! X k−1 ξ + X k−2 ξX +· + ξX k−1 ,(2.109) (log◦Θ) ∗,C (V ) = log ∗,Θ(C) (Θ ∗,C (V )).(2.110) Due to the nilpotency of LT 0 (n), the matrix logarithm over LT 1 (n) and exponenti- ation over LT 0 (n) are free from eigendecomposition. Although the Euclidean inner product in the ECM and LECM columns can be any inner product, we use the canonical one in this thesis. (3) OLM and LSM. Let H,W ∈ Hol(n), S = H + D(H), Σ = D ⋆ (C)CD ⋆ (C), X = Log ⋆ (C) = log(Σ) ∈ Row 0 (n), and Y ∈ Row 0 (n). The involved maps and 42 Chapter 2. Mathematical Background differentials are [191, Thms. 2.4 and 4.1]: Log ◦ ∗,C (V ) = off log ∗,C (V ) ,(2.111) Exp ◦ ∗,H (W ) = exp ∗,S (W +D ∗,H (W )),(2.112) D ∗,H (W ) =− diag H 0 −1 D exp ∗,S (W ) 1 ,(2.113) H 0 = [H 0 il ]∈S n ++ , H 0 il = X j,k U ij U ik U lj U lk [L exp ] j,k ,(2.114) Log ⋆ ∗,C (V ) = log ∗,Σ ∆V ∆ + 1 2 V 0 Σ + ΣV 0 ,(2.115) Exp ⋆ ∗,X (Y ) = ∆ −1 exp ∗,X (Y )− 1 2 ∆ −2 D exp ∗,X (Y ) Σ(2.116) +ΣD exp ∗,X (Y ) ∆ −2 ∆ −1 ,(2.117) where S = U diag (λ 1 ,...,λ n )U ⊤ , L exp is the Loewner matrix of exp ∗,S , and 1 is the vector of all ones. Here, log ∗ and exp ∗ are given by the Daleckii–Krein formula in Eqs. (2.90) and (2.91), while diag(·) : R n → Diag(n) returns a diagonal matrix from an input vector. The remaining auxiliary quantities are ∆ = D(Σ) 1 2 , V 0 =−2 diag (I n + Σ) −1 ∆V ∆1 .(2.118) (4) Permutation. Let S n be the group of permutation matrices P σ = δ i,σ(j) 1≤i,j≤n associated with permutations σ, and let D ± (n) =diag (ε 1 ,...,ε n )| ε∈−1, 1 n be the group of diagonal matrices with entries in −1, 1. Thanwerdas [191, Thm. 1.1] showed that the largest congruence action on full-rank correlation ma- trices is the action of signed permutation matrices: ⋆ : (A,C)∈ S ± (n)× Cor + (n)7−→ ACA ⊤ ∈ Cor + (n), S ± (n) =D ± (n)S n . (2.119) As both Log ⋆ ∗ and Log ◦ ∗ are permutation-equivariant [191, Thms. 2.2(1) and 3.6(1)], permutation-invariant metrics over the correlation manifold can be induced by permutation-invariant inner products over Hol(n) and Row 0 (n), respectively. (5) Invariant inner products on Hol(n). For n ≥ 4, permutation-invariant inner 43 2.9. Example Manifolds products on Hol(n) are [190, Thm. 8.7] ⟨X 1 ,X 2 ⟩ (α,β,γ) = α tr(X 1 X 2 ) + β Sum (X 1 X 2 ) + γ Sum(X 1 ) Sum(X 2 ), ∀X 1 ,X 2 ∈ Hol(n), (2.120) with α > 0, 2α + (n − 2)β > 0, and α + (n − 1)(β + nγ) > 0. For n = 3, permutation-invariant inner products have the same form with α = 0: ⟨X 1 ,X 2 ⟩ (α,β,γ) = β Sum(X 1 X 2 ) + γ Sum(X 1 ) Sum(X 2 ), with β > 0 and β + 3γ > 0. (2.121) For n = 2, they have the same form with α = β = 0: ⟨X 1 ,X 2 ⟩ (α,β,γ) = γ Sum(X 1 ) Sum(X 2 ), with γ > 0.(2.122) (6) Invariant inner products on Row 0 (n). For n ≥ 4, permutation-invariant inner products on Row 0 (n) are [191, Thm. 4.2] ⟨Y 1 ,Y 2 ⟩ (α,δ,ζ) = α tr(Y 1 Y 2 ) + δ tr(D(Y 1 )D(Y 2 )) + ζ tr(Y 1 ) tr(Y 2 ), ∀Y 1 ,Y 2 ∈ Row 0 (n), (2.123) with α > 0, nα + (n− 2)δ > 0, and nα + (n− 1)(δ + nζ) > 0. For n = 3, the permutation-invariant inner products have the same form with α = 0. For n = 2, they have the same form with α = δ = 0. (7) OLM and LSM invariance. Combining the permutation-equivariant diffeomor- phisms with the above invariant inner products gives permutation-invariant OLM and LSM. As shown by Thanwerdas [191, Thm. 2.7], OLM is further invariant un- der signed permutations when β = γ = 0, in which case the associated ⟨·,·⟩ (α,0,0) reduces to the scaled canonical Euclidean inner product: ⟨V,W⟩ (α,0,0) = α⟨V,W⟩, ∀V,W ∈ Hol(n).(2.124) In this thesis, we assume that ⟨·,·⟩ (α,β,γ) and ⟨·,·⟩ (α,δ,ζ) are the canonical Euclidean inner products. For n≤ 3, these inner products remain permutation invariant; the dimension-specific forms above account for redundancies among the displayed trace and sum terms. 44 Chapter 2. Mathematical Background (8) Inverse consistency. We briefly review inverse-consistency, a property exclusive to LSM. The cor-inversion is defined as I : Cor + (n)∋ C 7−→ Cor (C −1 )∈ Cor + (n) [191, Def. 1.4]. It corresponds to the matrix inversion inv : S n ++ ∋ Σ 7−→ Σ −1 ∈ S n ++ , as represented by the following commuting diagram: S n ++ S n ++ Cor + (n)Cor + (n) inv CorCor I (2.125) As shown by Thanwerdas [191, Thm. 1.7], LSM enjoys inverse-consistency: Log ⋆ (I(C)) =− Log ⋆ (C), ∀C ∈ Cor + (n).(2.126) (9) Vector structure. ECM, LECM, OLM, and LSM are all pulled back from Eu- clidean vector spaces. It is therefore natural to inherit vector addition and scalar multiplication from their prototype spaces. For each corresponding isometry φ, these operations take the form C⊕ C ′ = φ −1 (φ(C) + φ(C ′ )), t⊙ C = φ −1 (tφ(C)),(2.127) with t∈ R. 45 2.9. Example Manifolds MetricPrototype spaceDiffeomorphismsProperties ECM [195] LT 1 (n) = LT 0 (n) + I n Θ : C ∈ Cor + (n)7−→ D(Chol(C)) −1 Chol(C)∈ LT 1 (n) Θ −1 = Cor◦ Chol −1 : LT 1 (n)−→ Cor + (n) Null curvature LECM [195] LT 0 (n) log◦Θ : Cor + (n)−→ LT 0 (n) (log◦Θ) −1 = Cor◦ Chol −1 ◦ exp : LT 0 (n)−→ Cor + (n) Null curvature OLM [191] Hol(n) Log ◦ : C ∈ Cor + (n)7−→ (off◦ log)(C)∈ Hol(n) (Log ◦ ) −1 = Exp ◦ : H ∈ Hol(n)7−→ exp (D(H) + H)∈ Cor + (n) Permutation-invariance Null curvature LSM [191] Row 0 (n) Log ⋆ : C ∈ Cor + (n)7−→ log(D ⋆ (C)CD ⋆ (C))∈ Row 0 (n) (Log ⋆ ) −1 = Exp ⋆ : R∈ Row 0 (n)7−→ Cor(exp(R))∈ Cor + (n) Permutation-invariance Inverse-consistency Null curvature PHCM [195] PHS n−1 Chol : Cor + (n)−→L n ∼ = PHS n−1 Chol −1 :L n ∼ = PHS n−1 −→ Cor + (n) Nonpositive sectional curvature Table 2.7: Isometric prototype spaces and diffeomorphisms on the correlation manifold. OperationECMLECM C⊕ C ′ (φ EC ) −1 φ EC (C) + φ EC (C ′ ) (log◦Θ) −1 (log◦Θ(C) + log◦Θ(C ′ )) t⊙ C(φ EC ) −1 tφ EC (C) (log◦Θ) −1 (t log◦Θ(C)) g C (V,W )⟨Θ ∗,C (V ), Θ ∗,C (W )⟩⟨(log◦Θ) ∗,C (V ), (log◦Θ) ∗,C (W )⟩ Exp C (V )Θ −1 (Θ(C) + Θ ∗,C (V ))(log◦Θ) −1 (log◦Θ(C) + (log◦Θ) ∗,C (V )) Log C (C ′ )Θ −1 ∗,Θ(C) (Θ(C ′ )− Θ(C))(log◦Θ) −1 ∗,log◦Θ(C) (log◦Θ(C ′ )− log◦Θ(C)) γ(t;C,C ′ )Θ −1 ((1− t)Θ(C) + tΘ(C ′ ))(log◦Θ) −1 ((1− t) log◦Θ(C) + t log◦Θ(C ′ )) d(C,C ′ )∥Θ(C)− Θ(C ′ )∥log◦Θ(C)− log◦Θ(C ′ )∥ Fréchet meanΘ −1 1 N P N i=1 Θ(C i ) (log◦Θ) −1 1 N P N i=1 (log◦Θ)(C i ) Curvature00 PT C→C ′ (V )(Θ ∗,C ′ ) −1 (Θ ∗,C (V ))((log◦Θ) ∗,C ′ ) −1 ((log◦Θ) ∗,C (V )) Table 2.8: Vector operations and Riemannian operators under ECM and LECM. OperationOLMLSM C⊕ C ′ Exp ◦ (Log ◦ (C) + Log ◦ (C ′ ))Exp ⋆ (Log ⋆ (C) + Log ⋆ (C ′ )) t⊙ CExp ◦ (t Log ◦ (C))Exp ⋆ (t Log ⋆ (C)) g C (V,W ) Log ◦ ∗,C (V ), Log ◦ ∗,C (W ) (α,β,γ) Log ⋆ ∗,C (V ), Log ⋆ ∗,C (W ) (α,δ,ζ) Exp C (V )Exp ◦ Log ◦ (C) + Log ◦ ∗,C (V ) Exp ⋆ Log ⋆ (C) + Log ⋆ ∗,C (V ) Log C (C ′ )Exp ◦ ∗,Log ◦ (C) (Log ◦ (C ′ )− Log ◦ (C))Exp ⋆ ∗,Log ⋆ (C) (Log ⋆ (C ′ )− Log ⋆ (C)) γ(t;C,C ′ )Exp ◦ ((1− t) Log ◦ (C) + t Log ◦ (C ′ ))Exp ⋆ ((1− t) Log ⋆ (C) + t Log ⋆ (C ′ )) d(C,C ′ )∥Log ◦ (C)− Log ◦ (C ′ )∥ (α,β,γ) ∥Log ⋆ (C)− Log ⋆ (C ′ )∥ (α,δ,ζ) Fréchet meanExp ◦ 1 N P N i=1 Log ◦ (C i ) Exp ⋆ 1 N P N i=1 Log ⋆ (C i ) Curvature00 PT C→C ′ (V )(Log ◦ ∗,C ′ ) −1 Log ◦ ∗,C (V ) (Log ⋆ ∗,C ′ ) −1 Log ⋆ ∗,C (V ) Table 2.9: Vector operations and Riemannian operators under OLM and LSM. 46 Chapter 2. Mathematical Background PHCM [195, Def. 4.3 and Thm. 4.4]. It is defined through the Cholesky decomposition and a product of open-hemisphere models of hyperbolic space. For a correlation matrix C ∈ Cor + (n), let L = Chol(C). For k = 2,...,n, define the nonzero part of the k-th row of L by ℓ k = (L k1 ,...,L k )∈ HS k−1 ,HS k−1 = x∈ R k |∥x∥ = 1, x k > 0 . (2.128) Thus,L n is identified with the product of n−1 open hemispheres, denoted by PHS n−1 = Q n−1 i=1 HS i . The first row, for which L 11 = 1 and HS 0 =1, is trivial and omitted from the product. PHCM is the pullback by the Cholesky decomposition of the product metric on Q n−1 i=1 (HS i ,α i g HS i ), where each α i is a positive weight and g HS i denotes the metric tensor on HS i . In particular, PHCM with all weights equal to 1 is called the canonical PHCM, on which we focus below. Given C ∈ Cor + (n) and L = Chol(C)∈L n , define Ψ = ψ 1 ×·× ψ n−1 :L n → n−1 Y i=1 HS i ,(2.129) where ψ i (L) = (L i+1,1 ,...,L i+1,i+1 )∈ HS i .(2.130) For Z ∈ T L L n , its differential is the corresponding row extraction, ψ i ∗,L (Z) = (Z i+1,1 ,...,Z i+1,i+1 )∈ T ψ i (L) HS i ,Ψ ∗,L = ψ 1 ∗,L ×·× ψ n−1 ∗,L . (2.131) The Riemannian operators under PHCM are obtained from the product geometry and the geometry of each HS i . Let C ′ ∈ Cor + (n), L ′ = Chol(C ′ ), and V,W ∈ T C Cor + (n). Write a i = ψ i (L), a ′ i = ψ i (L ′ ), and ξ i (V ) = ψ i ∗,L (Chol ∗,C (V )). Then g C (V,W ) = D D(L) −1 L L −1 V L −⊤ 1 2 , D(L) −1 L L −1 WL −⊤ 1 2 E ,(2.132) Exp C (V ) = Chol −1 Ψ −1 Exp HS 1 a 1 (ξ 1 (V )),..., Exp HS n−1 a n−1 (ξ n−1 (V )) !! ,(2.133) Log C (C ′ ) = Chol −1 ∗,L (Ψ ∗,L ) −1 Log HS 1 a 1 (a ′ 1 ),..., Log HS n−1 a n−1 a ′ n−1 !! ,(2.134) γ(t;C,C ′ ) = Chol −1 Ψ −1 γ HS 1 (t;a 1 ,a ′ 1 ),..., γ HS n−1 t;a n−1 ,a ′ n−1 !! ,(2.135) 47 2.9. Example Manifolds d(C,C ′ ) 2 = n−1 X i=1 arccosh 1 + 1−⟨a i ,a ′ i ⟩ L i+1,i+1 L ′ i+1,i+1 2 .(2.136) Here, Log HS i , Exp HS i , and γ HS i are the corresponding operators on HS i ; the distance and these closed forms follow from the hyperboloid–hemisphere isometry [195, Thm. 4.2]. 2.9.3 Grassmannian Manifolds The Grassmannian has been widely applied in machine learning, ranging from action recognition [108] to question answering [159], shape generation [224], image classification [204], and signal analysis [210]. The Grassmannian manifold is the set of p-dimensional subspaces of R n [19]. It has two matrix representations: the projector perspective (P) and the orthonormal-basis perspective (ONB): f Gr(p,n) = P ∈S n | P 2 = P,rank(P ) = p , Gr(p,n) = n [U ]| [U ] := n e U ∈ St(p,n)| e U = UR, R∈ O(p) o , (2.137) whereS n is the Euclidean space of symmetric matrices, St(p,n) is the Stiefel manifold, and O(p) is the orthogonal group. By abuse of notation, we use [U ] and U interchange- ably for elements of Gr(p,n), where the n× p column-wise orthonormal matrix U is taken as a representative of an equivalence class. Helmke and Moore [100] show that the ONB perspective is diffeomorphic to the P representation by π : Gr(p,n)∋ U 7→ U ⊤ ∈ f Gr(p,n).(2.138) As shown by Nguyen [157, Sec. 3.2] and Nguyen and Yang [159, Sec. 2.3.1], the Grass- mannian admits gyro-structures defined by Eqs. (2.77) to (2.83). Under the ONB perspective, tangent vectors are represented by horizontal lifts. Given an orthogonal complement U ⊥ ∈ St(n− p,n) of U, every tangent vector ∆ ∈ T U Gr(p,n) can be written as ∆ = U ⊥ B, B ∈ R (n−p)×p .(2.139) Under the P perspective, every point can be written as P = O e I p,n O ⊤ with O ∈ O(n), 48 Chapter 2. Mathematical Background OperatorONB: Gr(p,n)P: f Gr(p,n) g U (∆, Ξ) or eg P (A 1 ,A 2 ) g U (∆, Ξ) =⟨∆, Ξ⟩eg P (A 1 ,A 2 ) = 1 2 ⟨A 1 ,A 2 ⟩ d(U,V ) or e d(P,Q) ∥arccos(Σ)∥ U ⊤ V SVD := OΣR ⊤ 1 2 √ 2 ∥log ((I n − 2Q) (I n − 2P ))∥ Log U V or g Log P Q O arctan(Σ)R ⊤ I n − U ⊤ V U ⊤ V −1 SVD := OΣR ⊤ 1 2 [log ((I n − 2Q) (I n − 2P )),P ] Exp U ∆ or g Exp P A UR cos(Σ)R ⊤ + O sin(Σ)R ⊤ ∆ SVD := OΣR ⊤ exp ([A,P ])P exp (−[A,P ]) γ(t;U,V ) or eγ(t;P,Q) UR cos(tΣ)R ⊤ + O sin(tΣ)R ⊤ Log U V SVD := OΣR ⊤ exp t[ g Log P Q,P ] P exp −t[ g Log P Q,P ] PT U→V (∆) or PT P→Q (A) UR O − sin(Σ) cos(Σ) O ⊤ + I n − O ⊤ ∆ Log U V SVD := OΣR ⊤ exp [ g Log P Q,P ] A exp −[ g Log P Q,P ] FM(U i ) or g FM(P i ) Karcher flowKarcher flow References[71, 19][19] Table 2.10: Riemannian operators on the Grassmannian under ONB and P. OperatorONB: Gr(p,n)P: f Gr(p,n) GyroadditionU ⊕ Gr V = exp(Ω U )VP e ⊕ Gr Q = exp(Ω P )Q exp (−Ω P ) Gyro identityI p,n e I p,n Scalar gyromultiplication t⊙ Gr U = exp (tΩ U )I p,n t e ⊙ Gr P = exp (tΩ P ) e I p,n exp (−tΩ P ) Gyro-inverse⊖ Gr U = exp (−Ω U )I p,n e ⊖ Gr P = exp (−Ω P ) e I p,n exp(Ω P ) References[159][157] Table 2.11: Gyro operators on the Grassmannian under ONB and P. and every tangent vector A∈ T P f Gr(p,n) can be written as A = O " 0 B ⊤ B 0 # O ⊤ , B ∈ R (n−p)×p .(2.140) Tab. 2.10 and Tab. 2.11 summarize the Riemannian and gyro operators, respectively; their formulas use the singular value decomposition (SVD). In these tables, U,V ∈ Gr(p,n) and ∆, Ξ∈ T U Gr(p,n) refer to the ONB perspective, while P,Q∈ f Gr(p,n) and A,A 1 ,A 2 ∈ T P f Gr(p,n) refer to the P perspective. To distinguish the two perspectives, P Riemannian operators are marked by a tilde, such as g Log and g Exp. The differential of π is π ∗,U (∆) = U ∆ ⊤ + ∆U ⊤ . For gyro operators, define P = g Log e I p,n (P ). Let Ω U = [ U ⊤ , e I p,n ] and Ω P = [P, e I p,n ]. The identities are I p,n = [I p ,0] ⊤ and e I p,n = I p,n I ⊤ p,n . 49 2.9. Example Manifolds Operator R⊕ S g R (A 1 ,A 2 )d(R,S)Log R SExp R (A)γ(t;R,S)FM ExpressionRS ⟨A 1 ,A 2 ⟩ log(R ⊤ S) R log(R ⊤ S) R exp R ⊤ A R exp t log(R ⊤ S) Karcher flow References[146] Table 2.12: Lie group structures and Riemannian operators on rotation matrices. 2.9.4 Special Orthogonal Groups The set of n× n rotation matrices forms a Lie group, known as the special orthogonal group and denoted by SO(n) [197]: SO(n) = R∈ R n×n | R ⊤ R = I n ,det(R) = 1 .(2.141) Its group operation is the matrix product, with the identity matrix as the neutral element. Any tangent vector A ∈ T R SO(n) can be represented as A = RV , with V ∈ so(n). Here, so(n) is the Lie algebra of SO(n), which is the tangent space at the identity matrix, formed by the set of n× n skew-symmetric matrices: so(n) = Ω∈ R n×n | Ω ⊤ =−Ω .(2.142) The Fréchet mean can be obtained by Karcher flow [146]. Furthermore, if all rotations lie in a closed ball of radius r < π /2, then Karcher flow converges to the unique mean [146, Thm. 5]. Given R,S ∈ SO(n) and tangent vectors A,A 1 ,A 2 ∈ T R SO(n), Tab. 2.12 summarizes all the associated operators on SO(n). Across these matrix-manifold examples, the tables expose the Euclidean template behind the spaces used later: flat pullback metrics use ordinary addition and scaling in a chart, quotient metrics use horizontal representatives, product metrics act factor by factor, and Lie-group metrics use group translation. 2.9.5 Constant-Curvature Manifolds Definition 59 (Constant-curvature space [165, Ch. 8]). A constant-curvature space (CCS) is a complete, simply connected, n-dimensional Riemannian manifold of constant curvature K. By O’Neill [165, Cor. 8.25], any two CCSs with the same dimension and the same curvature K are isometric. Thus, a CCS is determined up to isometry by K: Euclidean space corresponds to K = 0, spherical space to K > 0, and hyperbolic space to K < 0. In the following, we introduce several concrete models of these spaces that are used 50 Chapter 2. Mathematical Background later: the K-stereographic model, the K-radius model, and the Beltrami–Klein model. K-stereographic model [13]. This model has shown success in different appli- cations, including computer vision [201], natural language processing [76, 180], graph learning [13, 85, 86], and astronomy [44]. It is defined as a model st n K with the conformal metric ⟨u,v⟩ st x = (λ K x ) 2 ⟨u,v⟩, λ K x = 2 1 + K∥x∥ 2 ,(2.143) where K ∈ R is the constant curvature and λ K x is a conformal factor. In particular, st n K is the scaled R n when K = 0. It unifies the spherical projected hypersphere D n K , Euclidean space R n , and the hyperbolic Poincaré ball P n K : st n K = D n K = R n ,For K > 0, spherical geometry, R n ,For K = 0, Euclidean geometry, P n K = x∈ R n |∥x∥ 2 <− 1 K , For K < 0, hyperbolic geometry. (2.144) Although D n K = R n for K > 0, its metric is conformal to the Euclidean one. We abbreviate the K-stereographic model as the stereographic model. Bachmann et al. [13, Eqs. 2–3] show that this model admits a gyro-structure. K-radius model [182]. This model has been effective in various applications [38, 45, 15, 166, 99, 119]. It provides an extrinsic representation of the space with constant curvature K ∈ R, encompassing the sphere S n K , Euclidean space R n , and the Lorentz, or hyperboloid, model L n K : M n K = S n K = x∈ R n+1 |∥x∥ 2 = 1 K ,For K > 0, spherical geometry, R n ,For K = 0, Euclidean geometry, L n K = x∈ R n+1 |⟨x,x⟩ L = 1 K ,x t > 0 , For K < 0, hyperbolic geometry, where⟨x,x⟩ L =∥x s ∥ 2 −x 2 t is the Lorentzian quadratic form. Following the conventions of the hyperboloid, we write x = (x t ,x ⊤ s ) ⊤ , where x t ∈ R is the time component and x s ∈ R n is the spatial component [173]. When K ̸= 0, the model can be written compactly as M n K = x∈ R n+1 |⟨x,x⟩ K = 1 K , ⟨·,·⟩ K = ⟨·,·⟩, K > 0, ⟨·,·⟩ L , K < 0, (2.145) with the hyperbolic branch restricted to x t > 0. We abbreviate the K-radius model 51 2.9. Example Manifolds OperatorStereographic model st n K Beltrami–Klein model K n K Gyroadditionx⊕ K y = 1− 2K⟨x,y⟩− K∥y∥ 2 x + 1 + K∥x∥ 2 y 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 x⊕ E y = 1 1− K⟨x,y⟩ x + 1 γ K x y− K γ K x 1 + γ K x ⟨x,y⟩x Gyro identity00 Scalar gyromultiplicationt⊙ K x = tan K t tan −1 K p |K|∥x∥ p |K| x ∥x∥ t⊙ E x = tanh t tanh −1 √ −K∥x∥ √ −K x ∥x∥ Gyro-inverse⊖ K x =−x⊖ E x =−x Gyrationgyr[x,y]z = z + 2 (A st x + B st y) D st gyr[x,y]z = z + A E x + B E y D E References[200][200] Table 2.13: Gyro operators on stereographic and Beltrami–Klein models. as the radius model. On its negative-curvature branch, we use M n K = L n K = H n K interchangeably; in hyperbolic-only contexts, the specializations of ⊕ M K , ⊖ M K , and ⊙ M K are denoted by ⊕ L , ⊖ L , and ⊙ L , respectively. Its gyro-structure is introduced later in Sec. 3.3. Beltrami–Klein model. There are five models of hyperbolic space [33]. Apart from the above Poincaré ball and hyperboloid models, we further study the Beltrami– Klein model: K n K = x∈ R n |∥x∥ 2 <− 1 K ,with g K x (v,w) = ⟨v,w⟩ 1 + K∥x∥ 2 − K⟨x,v⟩⟨x,w⟩ 1 + K∥x∥ 2 2 , where K < 0 is the constant curvature and g K is its Riemannian metric. Although the Poincaré ball and Beltrami–Klein models share the same underlying set, their Rieman- nian metrics differ. This model admits an Einstein gyrovector space [200, Sec. 6.18]. The gyro-structures for the stereographic and Beltrami–Klein models are summa- rized in Tab. 2.13. In the table, x,y,z ∈ st n K for the stereographic model, x,y,z ∈ K n K for the Beltrami–Klein model, t∈ R, and γ K x = 1 + K∥x∥ 2 −1/2 is the Einstein gamma factor. The stereographic convention is tan K = tanh for K < 0 and tan K = tan for K > 0. The stereographic gyration in Tab. 2.13, following Bachmann et al. [13, App. C.2.6], is gyr[x,y]z = z + 2 A st x + B st y D st ,(2.146) with A st =−K 2 ⟨x,z⟩∥y∥ 2 − K⟨y,z⟩ + 2K 2 ⟨x,y⟩⟨y,z⟩, B st =−K 2 ⟨y,z⟩∥x∥ 2 + K⟨x,z⟩, D st = 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 = (1− K⟨x,y⟩) 2 + K 2 ∥x∥ 2 ∥y∥ 2 −⟨x,y⟩ 2 ≥ 0. (2.147) 52 Chapter 2. Mathematical Background The Cauchy–Schwarz inequality gives the last relation. The gyration formula applies whenever D st > 0; singular positive-curvature configurations with D st = 0 are excluded. The Einstein gyration coefficients for the Beltrami–Klein model are A E = K γ K x 2 γ K x + 1 γ K y − 1 ⟨x,z⟩− Kγ K x γ K y ⟨y,z⟩ + 2K 2 γ K x 2 γ K y 2 (γ K x + 1) γ K y + 1 ⟨x,y⟩⟨y,z⟩, B E = K γ K y γ K y + 1 γ K x γ K y + 1 ⟨x,z⟩ + γ K x − 1 γ K y ⟨y,z⟩ , D E = 1 + γ K x γ K y (1− K⟨x,y⟩) = 1 + γ K x⊕ E y . (2.148) Tab. 2.14 summarizes the associated Riemannian operators for the stereographic and radius models when K ̸= 0; at K = 0, they reduce to the usual Euclidean operators. On the sphere, logarithmic maps and parallel transports are restricted away from antipodal pairs. The gyro operators of the radius model and the closed-form Riemannian operators of the Beltrami–Klein model are introduced later in Sec. 3.3. In Tab. 2.14, for K ̸= 0, the curvature-aware trigonometric functions are tan K (·) = tan(·), K > 0, tanh(·), K < 0, sin K (·) = sin(·), K > 0, sinh(·), K < 0, cos K (·) = cos(·), K > 0, cosh(·), K < 0, All ratios in these tables are understood by continuous extension at removable zero denominators: in particular, t⊙ K 0 = 0, Exp x (0) = x, and Log x (x) = 0. On the positive-curvature branch, logarithmic maps and parallel transports are restricted away from antipodal pairs and other stated singular configurations. For the radius model, if x ∈ M n K and v ∈ T x M n K , then ∥v∥ K = p ⟨v,v⟩ K ; for K < 0, the Lorentzian form is positive definite only after restriction to this tangent space. These formulas show that constant-curvature neural layers can be implemented by choosing a model, applying the matching exponential, logarithmic, and transport maps, and reducing to Euclidean vector operations when K = 0. 53 2.9. Example Manifolds OperatorStereographic: P n K , D n K Radius: L n K , S n K Metric⟨u,v⟩ x = (λ K x ) 2 ⟨u,v⟩⟨u,v⟩ x =⟨u,v⟩ K Geodesic distance 2 √ |K| tan −1 K p |K|∥−x⊕ K y∥ 1 √ |K| cos −1 K (K⟨x,y⟩ K ) Exponential map x⊕ K tan K p |K| λ K x ∥v∥ 2 v √ |K|∥v∥ cos K p |K|∥v∥ K x + sin K √ |K|∥v∥ K √ |K|∥v∥ K v Logarithmic map 2 tan −1 K √ |K|∥−x⊕ K y∥ √ |K|λ K x −x⊕ K y ∥−x⊕ K y∥ cos −1 K (β) √ sign(K)(1−β 2 ) (y− βx), β = K⟨x,y⟩ K Parallel transport λ K x λ K y gyr[y,−x]v− K⟨y,v⟩ K 1+K⟨x,y⟩ K (x + y) Fréchet mean K < 0: [142, Alg. 1] K > 0: Karcher flow K < 0: [142, Alg. 3] K > 0: Karcher flow References[182][182] Table 2.14: Riemannian operator templates for stereographic and radius constant- curvature coordinates. 54 Chapter 3 Riemannian Batch Normalization 3.1 Introduction Motivated by the great success of normalization techniques [109, 12, 198, 219], re- searchers have sought to devise normalization layers tailored for manifold-valued data. Brooks et al. [31] introduced Riemannian Batch Normalization (RBN) designed specifi- cally for the SPD manifold, with the ability to normalize the Riemannian mean. Kobler et al. [123] extended this approach to further control the Riemannian variance. How- ever, the above methods are constrained to AIM on the SPD manifold, limiting their applicability. On the other hand, Chakraborty [34] proposed two distinct Riemannian normalization frameworks: one for Riemannian homogeneous spaces [34, Algs. 1–2] and another for matrix Lie groups [34, Algs. 3–4]. Nonetheless, the normalization de- signed for Riemannian homogeneous spaces can normalize neither the mean nor the variance, while the one for matrix Lie groups is confined to a specific type of distance [34, Sec. 3.2]. Meanwhile, Lou et al. [142, Alg. 2] proposed an RBN layer for general geometries. However, similar to Chakraborty [34, Algs. 1–2], it lacks theoretical guar- antees for normalizing sample statistics. Therefore, a principled Riemannian normal- ization framework capable of controlling both Riemannian mean and variance remains unexplored. Given that Batch Normalization (BN) [109] serves as the foundational prototype for various types of normalization, this chapter focuses on RBN, with the potential to be extended to other normalization variants. We first present a general RBN framework for Lie groups, referred to as Lie Group Batch Normalization (LieBN), which can normalize both the Riemannian mean and variance under invariant metrics. To extend this nor- malization principle beyond Lie groups, we next introduce pseudo-reductive gyrogroups, 55 3.2. Lie Group Batch Normalization Figure 3.1: Illustration of LieBN on the SPD, rotation, and correlation Lie groups. The 2×2 SPD, 3×3 rotation, and 3×3 correlation manifolds can be embedded into R 3 as an open cone [220], a closed ball with antipodal points identified [95], and an open elliptope [195], respectively. LieBN is illustrated by (1) the left-invariant AIM and the proposed right-invariant CRIM geometry on the SPD manifold, (2) left or right translation under a bi-invariant metric on the rotation manifold, and (3) the bi-invariant ECM and LSM geometry on the correlation manifold. On the SPD and correlation manifolds, the batch mean and variance of the same input samples differ under different geometries. In all sub-figures, the black, blue, green, and red dots denote the boundary of the space, the input Lie group samples, the normalized samples, and the batch mean, respectively. As illustrated, our LieBN effectively normalizes the Lie group distribution. a relaxation of classical gyrogroups that also encompasses Lie groups as special cases. Building on this algebraic structure, we develop Gyrogroup Batch Normalization (Gy- roBN), which generalizes LieBN beyond group structures. We then instantiate LieBN and GyroBN on SPD manifolds, rotation matrices, correlation matrices, the Grassman- nian, and different CCSs. Extensive experiments across multiple tasks demonstrate the effectiveness of this unified design. 3.2 Lie Group Batch Normalization 3.2.1 Introduction Since several manifold-valued measurements form Lie groups, such as SPD manifolds [9, 137, 195], special orthogonal groups SO(n) [28], and full-rank correlation matrices 56 Chapter 3. Riemannian Batch Normalization [195, 191], we direct our attention to Lie groups. As each Lie group naturally admits left- and right-invariant metrics [69, Ch. 1.2], we propose a principled framework for RBN over Lie groups under invariant metrics, referred to as LieBN. Compared to previous work, our framework provides a theoretical guarantee for normalizing the Riemannian sample mean and variance. Empirically, we focus on the SPD, special orthogonal, and full-rank correlation man- ifolds. On SPD manifolds, we generalize three existing Lie group structures into pa- rameterized ones by matrix power deformation. Additionally, we propose a novel right- invariant metric, which, to the best of our knowledge, is the first non-trivial right- invariant SPD metric 1 , referred to as the Cholesky Right Invariant Metric (CRIM). We then instantiate our LieBN framework on SPD manifolds under these four Lie group structures. For rotation matrices, we adopt the popular bi-invariant metric [28], which will induce two types of LieBN: one w.r.t. left-invariance and another w.r.t. right- invariance. On the correlation manifold, we manifest our LieBN under four recently developed correlation geometries [195, 191]. To facilitate usage, we provide a LieBN toolbox compatible with PyTorch, which can be used as a drop-in module. Fig. 3.1 illustrates our LieBN on different geometries, while Fig. 3.2 illustrates a minimal demo. Extensive experiments on SPD, rotation, and correlation manifolds involving radar recognition, human action recognition, and electroencephalography (EEG) classifica- tion demonstrate the effectiveness of our methods. We emphasize that our work is fundamentally distinct from Brooks et al. [31], Kobler et al. [123], Lou et al. [142] in theory and more general than Chakraborty [34]. Pre- vious RBN methods are either designed for specific geometries [31, 123, 34] or fail to control both the mean and variance [142]. In contrast, our LieBN ensures the normal- ization of both the mean and variance across general Lie groups. In summary, our main contributions are: • A general LieBN framework with controllable first- and second-order moments; • A novel right-invariant metric on the SPD manifold, which is the first non-trivial right-invariant SPD metric; • Concrete instantiations of our LieBN framework on different geometries: four on SPD manifolds, one on rotation matrices, and four on correlation manifolds; 1 Although some metrics are bi-invariant, the associated group structures are commutative [9, 137]. Therefore, their bi-invariance is reduced to left-invariance. 57 3.2. Lie Group Batch Normalization from LieBN import LieBNSPD , LieBNRot , LieBNCor from LieBN.Geometry.SPD import SPDMatrices from LieBN.Geometry.Rotations import RotMatrices from LieBN.Geometry.Correlation import Correlation # ==== SPD matrices ==== P_spd = SPDMatrices(n=5).random(4, 2, 5, 5) # Implemented metrics: LEM ,ALEM ,LCM ,AIM ,CRIM liebn_spd = LieBNSPD ([2, 5, 5], metric="LEM", batchdim =[0]) output_spd = liebn_spd(P_spd) # ==== SO(3) matrices ==== P_so3 = RotMatrices ().random(4, 2, 3, 3, 3) # LieBN -Left if is_left else -Right liebn_so3 = LieBNRot ([3, 3, 3], batchdim =[0, 1], is_left=False) output_so3 = liebn_so3(P_so3) # ==== Correlation matrices ==== P_cor = Correlation(n=5).random(4, 2, 5, 5) # Implemented metrics: ECM ,LECM ,OLM ,LSM liebn_cor = LieBNCor ([2, 5, 5], metric="ECM", batchdim =[0]) output_cor = liebn_cor(P_cor) Figure 3.2: Minimal examples of applying LieBN. • Validation of the effectiveness of our LieBN framework by extensive experiments on different geometries. 2 Outline. Sec. 3.2.2 recalls the invariant metrics and Lie structures used by LieBN. Sec. 3.2.3 revisits Euclidean BN and RBN. Sec. 3.2.4 develops LieBN on Lie groups under left- and right-invariant metrics and establishes its statistical control. Sec. 3.2.5 instantiates LieBN on SPD, rotation, and full-rank correlation manifolds. Sec. 3.2.6 reports experiments that validate LieBN across these geometries. Proofs are deferred to Sec. B.2. 3.2.2 Preliminaries An invariant metric can be understood as the Lie-group analogue of the Euclidean inner product being unaffected by translations. In Euclidean space, adding the same vector on the left or on the right preserves inner products, while on a Lie group the corresponding requirement is that left or right group translations preserve the Riemannian metric. 2 The code is available at https://github.com/GitZH-Chen/LieBN.git. 58 Chapter 3. Riemannian Batch Normalization Operator(α,β)-AIM(α,β)-LEMLCM Q⊕ PKPK ⊤ exp (log(P ) + log(Q))Chol −1 (⌊L + K⌋ + KL) ⊖PChol −1 (L −1 )exp (− log(P ))Chol −1 (−⌊L⌋ + L −1 ) IdentityI n I n I n WFM Karcher Flowexp ( P i w i log(P i ))ψ −1 LC ( P i w i ψ LC (P i )) InvarianceLeft-invarianceBi-invarianceBi-invariance Table 3.1: Review of SPD Lie groups and invariant metrics. GroupQ⊕ P ⊖PIdentityWFMInvariance SO(n)QPP −1 = P ⊤ I n Karcher FlowBi-invariance Table 3.2: Review of the rotation Lie group and invariant metric. Definition 60 (Invariance [69]). A Riemannian metric g L over a Lie groupM,⊕ is left-invariant if, for any x,y ∈ M and V 1 ,V 2 ∈ T y M, it satisfies g L y (V 1 ,V 2 ) = g L L x (y) (L x∗,y (V 1 ), L x∗,y (V 2 )), with L x (y) = x⊕y as the left translation by x, and L x∗,y as the differential map of L x at y. Similarly, a right-invariant metric g R satisfies g R y (V 1 ,V 2 ) = g R R x (y) (R x∗,y (V 1 ), R x∗,y (V 2 )), with R x (y) = y⊕x as the right translation by x, and R x∗,y as the differential map of R x at y. Many popular matrix manifolds used in machine learning form Lie groups, includ- ing the SPD manifold, the full-rank correlation manifold, and the rotation group. Al- though the corresponding Riemannian geometries have been reviewed in Secs. 2.9.1, 2.9.2 and 2.9.4, Tabs. 3.1 to 3.3 provide a focused recap of their Lie structures. The notation follows Tabs. 2.5, 2.7 to 2.9 and 2.12. 3.2.3 Revisiting Normalization 3.2.3.1 Revisiting Euclidean Normalization In Euclidean DNNs, normalization is a significant technique for accelerating network training by mitigating the issue of internal covariate shift [109]. While various nor- malization methods have been introduced [109, 12, 198, 219], they all share a common purpose: the normalization of the first and second moments. We focus on BN, the prototype of other normalization variants. Given a batch of activations x i N i=1 , the core operations in the standard Euclidean 59 3.2. Lie Group Batch Normalization OperatorECM LECM OLM LSM C⊕ C ′ φ −1 (φ(C) + φ(C ′ )) ⊖Cφ −1 (−φ(C)) Identityφ −1 (0 n×n ) WFMφ −1 P N i=1 w i φ(C i ) InvarianceBi-invariance Table 3.3: Review of full-rank correlation Lie groups and invariant metrics. Here φ denotes Θ, log◦Θ, Log ◦ , and Log ⋆ for ECM, LECM, OLM, and LSM, respectively. Methods Involved Statistics Controllable Mean Controllable Variance Geometries SPDBN [31, Alg. 1]Mean✓N/ASPD manifolds under AIM SPDBN [124, Alg. 1]Mean+Variance✓SPD manifolds under AIM SPDDSMBN [123]Mean+Variance✓SPD manifolds under AIM ManifoldNorm [34, Algs. 1–2]Mean+Variance✗Riemannian homogeneous spaces ManifoldNorm [34, Algs. 3–4]Mean+Variance✓A specific Lie group structure and distance RBN [142, Alg. 2]Mean+Variance✗Geodesically complete manifolds LieBN (Ours)Mean+Variance✓Lie groups Table 3.4: Summary of some representative RBN methods. BN can be expressed as: ∀i≤ N,x i ← γ x i − μ b p v 2 b + ε + β(3.1) where μ b is the batch mean, v 2 b is the batch variance, γ is the scaling parameter, β is the biasing parameter, and ε is a small scalar for stability. 3.2.3.2 Revisiting RBN Although endeavors have been made to develop Riemannian normalization approaches tailored for manifolds, none of the existing methods effectively handle the first and second moments in a principled manner. Brooks et al. [31] introduced RBN over SPD manifolds under AIM. The core oper- ations are defined as follows: Centering from mean M ∈S n ++ : ̄ P i ← M − 1 2 P i M − 1 2 ,(3.2) Biasing towards parameter B ∈S n ++ : ˆ P i ← B 1 2 ̄ P i B 1 2 ,(3.3) where P i N i=1 are SPD matrices, and M is their Fréchet mean under AIM. Let Γ P→Q (S) = Exp Q [PT P→Q (Log P (S))],(3.4) 60 Chapter 3. Riemannian Batch Normalization where P,Q,S ∈S n ++ . Under AIM, Eqs. (3.2) and (3.3) can be more generally expressed as Γ I→B [Γ M→I (P i )].(3.5) However, Eqs. (3.2) and (3.3) only consider the Riemannian mean 3 and do not consider the Riemannian variance. To remedy this limitation, Kobler et al. [123] further extended the RBN to involve the second-order statistics. The key operation is formulated as ∀i≤ N, ̄ P i ← Γ I→B [(Γ M→I (P i )) s v ],(3.6) where v 2 is the Fréchet variance, and s ∈ R is a scaling factor. However, this method is still limited to SPD manifolds under AIM. In parallel, Chakraborty [34, Algs. 1–2] proposed a general framework for Riemannian homogeneous spaces based on Eq. (3.5), which involves both first and second moments. However, Eq. (3.5) does not gener- ally guarantee control over the Riemannian mean, resulting in agnostic Riemannian statistics [34, Sec. 3.1]. To mitigate this limitation, Chakraborty [34, Algs. 3–4] further proposed normalization over matrix Lie groups. However, the discussion is limited to a certain distance, limiting the applicability of their method. On the other hand, Lou et al. [142, Alg. 2] proposed an RBN based on a variant of Eq. (3.5). Similarly, their approach suffers from the same problem of agnostic Riemannian statistics on general manifolds. In summary, prevailing Riemannian normalization approaches lack a principled guarantee for controlling the first- and second-order statistics. In contrast, our method can normalize first- and second-order statistics over general Lie groups. We summarize the above RBN methods in Tab. 3.4. 3.2.4 LieBN Since every Lie group naturally admits invariant metrics, we propose BN over Lie groups based on invariant metrics, referred to as LieBN. We first introduce the core operations under left-invariant metrics and then extend them to right-invariant metrics. Finally, we present the theoretical LieBN framework. In the following, we denote the neutral element in the Lie group M as E 4 . 3 Although not discussed in Brooks et al. [31], the congruent actions in Eqs. (3.2) and (3.3) can transfer the batch mean to a desired value under AIM. 4 The neutral element E is not necessarily the identity matrix. 61 3.2. Lie Group Batch Normalization 3.2.4.1 Ingredients under Left-invariant Metrics In this subsection, we always assume that the Lie group M admits a left-invariant metric g L . Recalling the standard Euclidean BN [109] in Eq. (3.1), two key points are noteworthy: (a) the Euclidean BN implicitly assumes a Gaussian distribution and can effectively normalize the latent Gaussian distribution; (b) the centering and biasing operations control the mean, while the scaling controls the variance. Therefore, ex- tending BN to Lie groups requires Lie-group counterparts of the Gaussian distribution, centering, biasing, and scaling. There are several notions of Gaussian distribution over manifolds [170, 218, 35, 14]. We adopt the intrinsic definition from Chakraborty and Vemuri [35], which characterizes a Gaussian distribution on the Lie group M with a mean parameter M ∈ M and variance σ 2 . This distribution is denoted as N (M,σ 2 ), and its probability density function (PDF) is p X | M,σ 2 = k(σ) exp − d(X,M ) 2 2σ 2 ,(3.7) where k(σ) is the normalizing constant and d(·,·) is the geodesic distance. When M is R with the standard Euclidean metric, Eq. (3.7) reduces to the Euclidean Gaussian. On Lie groups, the natural counterparts of addition and subtraction in Eq. (3.1) are group operations. Therefore, centering and biasing on Lie groups can be defined by the left translation. Additionally, we define scaling via the tangent space. Specifically, for a batch of activations P i N i=1 ⊂M, we define the key operations of LieBN as follows: Centering from mean M ∈M : ̄ P i ← L ⊖M (P i ),(3.8) Scaling: ˆ P i ← Exp E s √ v 2 + ε Log E ( ̄ P i ) ,(3.9) Biasing towards parameter B ∈M : ̃ P i ← L B ˆ P i ,(3.10) where M is the Fréchet mean, v 2 is the Fréchet variance,⊖M ∈M is the group inverse of M, L ⊖M and L B are left translations (L B (P i ) = B ⊕ P i ), and s ∈ R \ 0 is a scaling parameter. The following two propositions demonstrate the above operations in normalizing mean and variance: one related to population statistics and the other related to sample statistics. Proposition 61 (Population). [↓] Given a random point X over M,⊕,g L , and the Gaussian distribution N (M,v 2 ) defined in Eq. (3.7), we have the following for 62 Chapter 3. Riemannian Batch Normalization the population statistics: (1) (MLE of M) Given P i N i=1 ⊂M i.i.d. sampled from N (M,v 2 ), the maximum likelihood estimator (MLE) of M is the sample Fréchet mean. (2) (Gaussian homogeneity) Given X ∼N (M,v 2 ) and B ∈M, we have L B (X)∼N (L B (M ),v 2 ).(3.11) Proposition 62 (Sample). [↓] Given N samples P i N i=1 over the Lie group M,⊕,g L , define φ s (P i ) = Exp E [s Log E (P i )].(3.12) We then have the following for the sample statistics. • Sample mean homogeneity: FML B (P i ) = L B (FMP i ),∀B ∈M.(3.13) • Controllable dispersion from E: X N i=1 w i d 2 (φ s (P i ),E) = s 2 X N i=1 w i d 2 (P i ,E),(3.14) where w i N i=1 are weights satisfying a convexity constraint, i.e., ∀i,w i > 0 and P i w i = 1. Thm. 61 and Eq. (3.13) imply that our centering and biasing in Eqs. (3.8) and (3.10) can transfer the sample and population mean. As the post-centering mean is E, Eq. (3.14) implies that Eq. (3.9) can control the sample variance. More interestingly, the latent Gaussian distribution can be transferred under some geometries, such as SPD manifolds under LEM and LCM. Remark 63. The MLE of the mean of the Gaussian distribution has been examined in several previous works [176, 35, 34]. However, these studies primarily focus on particular manifolds or specific metrics. In contrast, our contribution lies in presenting a general result for Lie groups. 63 3.2. Lie Group Batch Normalization Remark 64. While Eq. (3.7) appeared in Kobler et al. [124], the authors only focus on SPD manifolds under AIM. The transformation of the population under their proposed RBN remains unexplored as well. Besides, while Chakraborty [34] analyzed the population properties for their RBN over matrix Lie groups, their results were confined within a specific distance. In contrast, our work provides a more extensive examination, encompassing both population and sample properties of our LieBN in a general manner. 3.2.4.2 Ingredients under Right-invariant Metrics The key insight beneath Eqs. (3.8) and (3.10) and Thms. 61 and 62 is that left trans- lation is an isometry under left-invariant metrics. Similarly, right translation is an isometry under right-invariant metrics. Therefore, it can be used for centering and biasing under right-invariant metrics. Following the previous notations, we define the centering and biasing under a right-invariant metric g R as centering to E: ̄ P i ← R ⊖M (P i ),(3.15) biasing towards B: ̃ P i ← R B ( ˆ P i ).(3.16) Similar to the case under left-invariant metrics, Thms. 61 and 62 can be easily extended to right-invariant metrics. Notably, the proofs for the MLE of M in Thm. 61 and controllable dispersion in Thm. 62 can be directly applied to the right-invariant metric. Therefore, we only show the homogeneity in the following proposition. Proposition 65. [↓] Given a random point X ∼N (M,v 2 ) over M,⊕,g R , B ∈ M, and N samples P i N i=1 over M, we have: (1) Gaussian homogeneity: R B (X)∼N (R B (M ),v 2 ); (2) Sample homogeneity: FMR B (P i ) = R B (FMP i ). 3.2.4.3 LieBN under Invariant Metrics With the above ingredients, Alg. 1 presents our theoretical LieBN framework. Similar to Ioffe and Szegedy [109], we use the moving average to update the running statistics. For a bi-invariant metric, LieBN can be implemented using either left or right translation. If the Lie group is commutative, LieBN under left and right translations are equivalent. Tab. 3.5 summarizes the LieBN types under different conditions. 64 Chapter 3. Riemannian Batch Normalization CommutativityNon-commutativeCommutative InvarianceLeftRightBiLeft = Right = Bi LieBN TypesLeftRightLeft & RightLeft = Right Table 3.5: Summary of LieBN types. The centering and biasing in Euclidean BN correspond to the group action of R. From a geometric perspective, the standard Euclidean metric is invariant under this group operation. Consequently, it is not surprising that our LieBN algorithm naturally generalizes the standard Euclidean BN. Proposition 66. [↓] The LieBN algorithm presented in Alg. 1 is equivalent to the standard Euclidean BN when M = R n , both during the training and testing phases. 3.2.5 Manifestations This section instantiates our LieBN in Alg. 1 on nine different Lie groups, including four on the SPD manifold, one on rotation matrices, and four on the correlation manifold. 3.2.5.1 LieBN on SPD Manifolds We first extend the current Lie groups on SPD manifolds by the matrix power deforma- tion, resulting in three families of parameterized Lie groups. Then, we propose a novel right-invariant metric on the SPD manifold, the first non-trivial right-invariant metric on this manifold. Finally, we construct LieBN layers based on these Lie structures. Deformed Lie structures on SPD manifolds. As shown in Tab. 3.1, there are three Lie groups on SPD manifolds, each with a left-invariant metric. These metrics include (α,β)-AIM, (α,β)-LEM, and LCM. For clarity, we denote the group operations w.r.t. (α,β)-AIM, (α,β)-LEM and LCM as ⊕ LieAI , ⊕ LE and ⊕ LC , respectively. Recently, Thanwerdas and Pennec [193] further extended (α,β)-AIM into a three- parameter family of metrics via the pullback of the matrix power function P θ (·), scaled by 1 θ 2 and denoted by (θ,α,β)-AIM. The matrix power serves as a deformation, wherein (θ,α,β)-AIM encompasses (α,β)-AIM with θ = 1, and becomes (α,β)-LEM as θ ap- proaches 0 [192]. Inspired by the deforming utility of the power function, we define the power-deformed metrics of (α,β)-LEM and LCM as the pullback metrics by P θ and scaled by 1 θ 2 . We denote these two metrics as (θ,α,β)-LEM and θ-LCM, respectively. We have the following results with respect to the deformation. 65 3.2. Lie Group Batch Normalization Algorithm 1: Lie Group Batch Normalization (LieBN) Input: A batch of activations P i N i=1 over Lie groups M,⊕,g, a small positive constant ε, and momentum η ∈ [0, 1], running mean M r = E, running variance v 2 r = 1, biasing parameter B ∈M, and scaling parameter s∈ R\0. Output : Normalized activations ̃ P i N i=1 . if training then Compute batch mean M b and variance v 2 b Update running statistics: M r ← WFM(1− η,η,M r ,M b ) v 2 r ← (1− η)v 2 r + ηv 2 b Use the batch statistics, M ← M b ,v 2 ← v 2 b else Use the running statistics, M ← M r ,v 2 ← v 2 r for i← 1 to N do Centering to the neutral element E: if g is left-invariant then ̄ P i ← L ⊖M (P i ) else ̄ P i ← R ⊖M (P i ) Scaling the variance: ˆ P i ← Exp E h s √ v 2 +ε Log E ( ̄ P i ) i Biasing towards parameter B: if g is left-invariant then ̃ P i ← L B ( ˆ P i ) else ̃ P i ← R B ( ˆ P i ) Proposition 67 (Deformation). [↓] (θ,α,β)-LEM is equal to (α,β)-LEM. θ-LCM interpolates between ̃g-LEM (as θ → 0) and LCM (θ = 1). Here, given any P ∈ S n ++ and tangent vectors V,W ∈ T P S n ++ , ̃g-LEM is defined as ⟨V,W⟩ P = ̃g(log ∗,P (V ), log ∗,P (W )),(3.17) where ̃g(V 1 ,V 2 ) = 1 2 ⟨V 1 ,V 2 ⟩− 1 4 ⟨D(V 1 ), D(V 2 )⟩, D(V i ) is a diagonal matrix consisting of the diagonal elements of V i , and log ∗,P is the differential map at P. As (θ,α,β)-LEM is equal to (α,β)-LEM, we focus on (α,β)-LEM, (θ,α,β)-AIM, and θ-LCM in the following. As a diffeomorphism, P θ can also pull back the group operations ⊕ LieAI and ⊕ LC , denoted by ⊕ θ-AI and ⊕ θ-LC , respectively. We have the 66 Chapter 3. Riemannian Batch Normalization following proposition on the invariance. Proposition 68 (Invariance). [↓] (θ,α,β)-AIM is left-invariant w.r.t. ⊕ θ-AI , while θ-LCM is bi-invariant w.r.t. ⊕ θ-LC . SPD right-invariant metrics. AIM is left-invariant w.r.t. ⊕ LieAI . We can also define a right-invariant metric w.r.t. ⊕ LieAI by definition [69, Ch. 1.2]: g CRI P (V,W ) =⟨(R ⊖ LieAI P ) ∗,P (V ), (R ⊖ LieAI P ) ∗,P (W )⟩ I (3.18) where R (·) denotes Lie group right translation,⊖ LieAI P is the inverse of P under⊕ LieAI , and ⟨·,·⟩ I denotes an arbitrary inner product on T I S n ++ . We set ⟨·,·⟩ I to be the same as the AIM at I, i.e., ⟨·,·⟩ (α,β) . We call this metric CRIM, as the group operation is defined by the matrix product of Cholesky factors [195, Sec. 3.2]. Theorem 69. [↓] Given any SPD matrices P,Q and tangent vector V ∈ T P S n ++ , the Riemannian operators on S n ++ ,g CRI are g CRI P (V,V ) = L(L −1 V L −⊤ ) 1 2 L −1 sym+ (α,β) ! 2 (3.19) d(P,Q) = log e Q − 1 2 e P e Q − 1 2 (α,β) ,(3.20) Exp P (V ) =⊖ LieAI Exp AI e P − ̄ V ,(3.21) Log P (Q) =− L ⊤ L e V L ⊤ ⊤ 1 2 sym+ ,(3.22) where L is the Cholesky factor of P = L ⊤ , ⊖ LieAI (·) is the group inverse, e P and e Q are the group inverses of P and Q, ̄ V = L −1 V L −⊤ 1 2 L −1 L −⊤ sym+ , and e V = Log AI e P e Q . Here, (X) sym+ = X + X ⊤ ,∀X ∈ R n×n denotes unnormalized symmetrization, and (X) 1 2 =⌊X⌋ + 1 2 X. Corollary 70. [↓] CRIM is geodesically complete, and the associated geodesic con- necting SPD matrices P and Q is γ (P,Q) (t) =⊖ LieAI n γ AI (t; e P, e Q) o =⊖ LieAI e P 1 2 e P − 1 2 e Q e P − 1 2 t e P 1 2 , (3.23) 67 3.2. Lie Group Batch Normalization Metric(θ,α,β)-AIM(α,β)-LEMθ-LCMθ-CRIM InvarianceLeft-invarianceBi-invarianceRight-invariance LieBN TypeLieBN-LeftLieBN-Left = LieBN-RightLieBN-Right Pullback MapP θ logP θ ◦ψ LC P θ CodomainS n ++ ,⊕ LieAI , 1 θ 2 g (α,β)-AI S n ,⟨·,·⟩ (α,β) LT n , 1 θ 2 ⟨·,·⟩S n ++ ,⊕ LieAI , 1 θ 2 g CRI Riemannian and Lie group operators in the codomain L Q (P ) or R Q (P )KPK ⊤ P + QP + QLQL ⊤ L ⊖Q (P ) or R ⊖Q (P )K −1 PK −⊤ P − QP − QL −1 QL −⊤ Exp E [s Log E (P )]P s sPsP⊖ LieAI ⊖ LieAI P s FMKarcher Flow Arithmetic average Arithmetic average Karcher Flow WFM(1− η,η,P 1 ,P 2 )P 1 2 1 P − 1 2 1 P 2 P − 1 2 1 η P 1 2 1 Arithmetic weighted average Arithmetic weighted average ⊖ LieAI e P 1 2 1 e P − 1 2 1 e P 2 e P − 1 2 1 η e P 1 2 1 Table 3.6: Key operators in calculating LieBN on SPD manifolds. InvarianceLieBN Type⊖RL R (S)R R (S)Exp I [s Log I (R)]FMWFM(1− η,η,R,S) Bi-invarianceLieBN-Left & LieBN-Right R −1 RSSRexp (s log (R))[146, Alg. 1]R exp(η log(R ⊤ S)) Table 3.7: Key operators in calculating LieBN on the rotation matrices. where e P = ⊖ LieAI P and e Q = ⊖ LieAI Q are group inverses, with γ AI as the geodesic under AIM. Similar to the discussion in Sec. 3.2.5.1, we define θ-CRIM as the deformed metric of CRIM by the pullback of matrix power function P θ (·) and scaled by 1 θ 2 . As the pullback of CRIM, θ-CRIM is right-invariant w.r.t. ⊕ θ-AI by definition. Proposition 71. θ-CRIM is right-invariant w.r.t. ⊕ θ-AI . Manifestations on SPD manifolds. So far, there are four families of invariant metrics on the SPD Lie groups: (1) left-invariant (θ,α,β)-AIM w.r.t. ⊕ LieAI ; (2) bi- invariant (α,β)-LEM w.r.t. ⊕ LE and θ-LCM w.r.t. ⊕ LC ; (3) right-invariant θ-CRIM w.r.t. ⊕ LieAI . Since all the above metrics are pullback metrics, the LieBN based on these metrics can be simplified and calculated in the codomain. We first show a general result on LieBN under the pullback metric. We denote Alg. 1 on the Lie group M as LieBN(P i ;B,s,ε,η), P i ∈P j N j=1 ⊂M.(3.24) Then we can obtain the following theorem. Theorem 72. [↓] Given a Lie groupM 1 , a Lie groupM 2 with an invariant metric g 2 , and a map f : M 1 → M 2 that is both a diffeomorphism and a Lie-group isomorphism, the map f induces an invariant metric g 1 on M 1 , denoted as g 1 = f ∗ g 2 . For a batch of activations P i N i=1 in M 1 , LieBN 1 (P i ;B,s,ε,η) in M 1 can 68 Chapter 3. Riemannian Batch Normalization be calculated in M 2 by the following process: Mapping data into M 2 : ̄ P i = f (P i ), ̄ B = f (B),(3.25) Performing LieBN in M 2 : ˆ P i = LieBN 2 ( ̄ P i ; ̄ B,s,ε,η),(3.26) Mapping the resulting data back to M 1 : ̃ P i = f −1 ( ˆ P i ),(3.27) where LieBN 2 is the LieBN on M 2 . Given a metric g onS n ++ , the power-deformed metric ̃g = 1 θ 2 P ∗ θ g is equal to P ∗ θ ( 1 θ 2 g). Thm. 72 indicates that the LieBN under ̃g can be calculated by the LieBN under 1 θ 2 g. Besides, as the Christoffel symbols remain the same under constant scaling, the LieBNs under 1 θ 2 g and g only differ in the variance. We denote g (α,β)-AI and g (θ,α,β)-AI as the metric tensors of (α,β)-AIM and (θ,α,β)-AIM, respectively. Based on the above discussions, the computations of the LieBN under g (θ,α,β)-AI are reduced to the LieBN under 1 θ 2 g (α,β)-AI . Similarly, denoting g CRI and g θ-CRI as the metric tensors of CRIM and θ-CRIM, then the LieBN under θ-CRIM can be calculated by the one under 1 θ 2 g CRI . Furthermore, as shown in Sec. 4.2.3, (α,β)-LEM is a pullback metric from the Euclidean space S n of symmetric matrices, while θ-LCM is a pullback metric from the Euclidean space LT n of lower triangular matrices. As shown in Thm. 66, the LieBN in the Euclidean space S n or LT n is simplified to the standard Euclidean BN. Therefore, the LieBNs under (α,β)-LEM and θ-LCM can be calculated by the Euclidean BN over S n and LT n , respectively. We denote the LieBN under left and right translations as LieBN-Left and LieBN- Right, respectively. Then, the LieBNs under (θ,α,β)-AIM and θ-CRIM correspond to LieBN-Left and LieBN-Right, respectively. As ⊕ LC and ⊕ LE are commutative, the LieBN-Left and LieBN-Right under (α,β)-LEM and θ-LCM are equivalent. We denote P,Q,P 1 and P 2 as points in the codomain, i.e., S n ++ with scaled CRIM for θ-CRIM, S n ++ with scaled (α,β)-AIM for (θ,α,β)-AIM, S n for (α,β)-LEM, and LT n for θ-LCM, respectively. For CRIM, we denote e P i = ⊖ LieAI P i for i = 1, 2. We summarize all the necessary ingredients in Tab. 3.6 for calculating SPD LieBN. Note that for (θ,α,β)-AIM, our scaling operation defined in Eq. (3.9) encompasses the scaling operation in Kobler et al. [124, Eq. (9)] as a special case, when (θ,α,β) = (1, 1, 0). 3.2.5.2 LieBN on Rotation Matrices As the Riemannian metric on the rotation matrices is bi-invariant, there are two in- stantiations of LieBN on this manifold, i.e., LieBN-Left based on the left translation 69 3.2. Lie Group Batch Normalization MetricECMLECMOLMLSM InvarianceBi-invariance LieBN TypeLieBN-Left = LieBN-Right Pullback MapΘlog◦ΘLog ◦ Log ⋆ CodomainLT 1 (n),⟨·,·⟩ LT 0 (n),⟨·,·⟩ Hol(n),⟨·,·⟩ (α,β,γ) Row 0 (n),⟨·,·⟩ (α,δ,ζ) Table 3.8: Summary of LieBN on the correlation. ⟨·,·⟩ (α,β,γ) and ⟨·,·⟩ (α,δ,ζ) are permutation-invariant inner products [191]. and LieBN-Right based on the right translation. In particular, the scaling can be fur- ther simplified: Exp I (s Log I (R)) = exp (s log (R)). For the specific SO(3), the matrix exp and log can be efficiently calculated without matrix decomposition [95, Sec. 3.2]. Tab. 3.7 presents the expressions of the required operators in Alg. 1. 3.2.5.3 LieBN on Full-Rank Correlation Matrices As summarized in Tab. 3.3, all four correlation metrics are bi-invariant, and their as- sociated Lie groups are commutative. Consequently, LieBN-Left is identical to LieBN- Right. Moreover, all four correlation metrics are pullback metrics from simpler Eu- clidean spaces. Therefore, LieBN over the correlation can be implemented according to Thm. 72: (1) map the correlation into the prototype Euclidean space, (2) apply Euclidean BN, and (3) map back to the correlation. Optimization. Finally, we discuss the optimization of the correlation-valued bi- asing parameter B ∈ Cor + (n). As reviewed in Sec. 2.9.2, the correlation matrix can be identified by the product of hyperbolic spaces via the Cholesky decomposi- tion. Given C ∈ Cor + (n), the k-th row of the Cholesky factor L = Chol(C) is (L k1 ,...,L k,k−1 ,L k , 0,..., 0) with L k > 0, which belongs to the hyperbolic space of an open hemisphere HS k−1 = x∈ R k |∥x∥ = 1,x k > 0 .(3.28) Besides, the open hemisphere HS n is isometric to the Poincaré ball P n =x∈ R n |∥x∥ < 1(3.29) by π HS n →P n ((x ⊤ ,x n+1 ) ⊤ ) = x 1 + x n+1 .(3.30) Therefore, each correlation can be parameterized with n− 1 Poincaré vectors. Each 70 Chapter 3. Riemannian Batch Normalization Poincaré vector can be optimized directly on its manifold using the Riemannian opti- mization strategy reviewed in Sec. 2.7. The above process can be expressed as C7→ 10 ·0 L 21 L 22 ·0 . . . . . . . . . . . . L n1 L n2 · L n 7→ x 1 ∈ P 1 . . . x n−1 ∈ P n−1 .(3.31) 3.2.6 Experiments This section validates our LieBN on nine invariant metrics across the SPD, rotation, and correlation matrices. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.1. 3.2.6.1 Experiments of LieBN on the SPD Manifold Note that our LieBN layers are architecture-agnostic and can be applied to any ex- isting SPD neural network. Following the previous work [106, 31, 123], we focus on two network architectures: (1) SPDNet [106] for drone recognition on the Radar data set [31], and human action recognition on the HDM05 [153] and FPHA [80] data sets; (2) TSMNet [123] for EEG classification on the Hinss2021 data set [101]. In the EEG application, TSMNet is endowed with SPD Domain-Specific Momentum Batch Nor- malization (SPDDSMBN) (denoted TSMNet+SPDDSMBN) [123], which is a domain adaptation version of Kobler et al. [124]. For a fair comparison, we also implement a domain-specific momentum LieBN, referred to as DSMLieBN. The backbone network architectures are represented as d 0 ,d 1 ,...,d L , where the dimension of the parameter in the i-th BiMap layer is d i ×d i−1 . As (α,β) only affects variance calculation through- out LieBN, we simply set (α,β) = (1, 0) and only tune the deformation factor θ. For each family of LieBN or DSMLieBN, we report two representatives: the standard one induced from the standard metric (θ = 1), and the one induced from the deformed metric with proper θ. If the standard one is already saturated, we only report the result of that standard variant. Application to SPDNet. As SPDNet is a classical SPD network, we apply our LieBN to SPDNet on the Radar, HDM05, and FPHA data sets. Additionally, we com- pare our method with SPDNetBN, which applies the SPDBN in Eqs. (3.2) and (3.3) to SPDNet. Following Brooks et al. [31], we use the architectures of 20, 16, 8, 93, 30, and63, 33 for the Radar, HDM05, and FPHA data sets, respectively. The 10-fold av- 71 3.2. Lie Group Batch Normalization (a) Radar data set. AccSPDNetSPDNetBN SPDNetLieBN θ = 1Best θ AIM-(1)LEM-(1)LCM-(1)CRIM-(1)LCM-(-0.5) Fit time (s)0.981.561.621.281.111.821.43 Mean±STD93.25±1.1094.85±0.9995.47±0.9094.89±1.0493.52±1.0794.35±0.6894.80±0.71 Max94.496.1396.2796.895.295.695.73 (b) HDM05 data set. AccSPDNetSPDNetBN SPDNetLieBN θ = 1Best θ AIM-(1)LEM-(1)LCM-(1)CRIM-(1)AIM-(1.5)LCM-(0.5)CRIM-(0.5) Fit time (s)0.570.971.140.870.661.371.461.011.74 Mean±STD 59.13±0.6766.72±0.5267.79±0.6565.05±0.6366.68±0.7163.25±0.8868.16±0.6870.84±0.9265.76±0.54 Max60.3467.6668.7566.0568.5264.9469.2572.2766.96 (c) FPHA data set. AccSPDNetSPDNetBN SPDNetLieBN θ = 1Best θ AIM-(1)LEM-(1)LCM-(1)CRIM-(1)AIM-(1.5)LCM-(0.5)CRIM-(-0.5) Fit time (s)0.320.620.80.550.390.921.030.651.21 Mean±STD85.59±0.7289.33±0.4989.70±0.5186.56±0.7977.64±1.0084.65±1.2090.39±0.6686.33±0.4386.40±0.57 Max8690.1790.587.837986.6792.178787.17 Table 3.9: 10-fold average results of SPDNet with and without SPDBN or LieBN on the Radar, HDM05, and FPHA data sets. If the LieBN under the standard metric (θ = 1) is not saturated, the rightmost columns report the deformed LieBN. The best results are bold. erage results, including the average training time (s/epoch), are summarized in Tab. 3.9. We have three key observations regarding the choice of metrics, deformation, and train- ing efficiency. • The choice of metrics. The metric that yields the most effective LieBN layer differs for each data set. Specifically, the optimal LieBN layers on these three data sets are the ones induced by AIM-(1), LCM-(0.5), and AIM-(1.5), respectively, which improve the performance of SPDNet by 2.22, 11.71, and 4.80 percentage points. Additionally, although the LCM-based LieBN performs worse than other LieBN variants on the Radar and FPHA data sets, it exhibits the best performance on the HDM05 data set. These observations demonstrate the value of a unified normalization framework that can be instantiated under multiple admissible SPD geometries. • The effect of deformation. Deformation patterns also vary across data sets. Firstly, the standard AIM and CRIM are already saturated on the Radar data set. Secondly, the appropriate deformation θ can further enhance the performance of LieBN. Notably, even though the LieBNs induced by LCM-(1) and CRIM-(1) 72 Chapter 3. Riemannian Batch Normalization (a) Inter-session classification MethodFit Time Mean±STD SPDDSMBN0.1654.12±9.87 DSMLieBN AIM-(1)0.1655.10±7.61 LEM-(1)0.1354.95±10.09 LCM-(1)0.1051.54±6.88 CRIM-(1)0.2951.86±9.21 LCM-(0.5)0.1553.11±5.65 (b) Inter-subject classification MethodFit Time Mean±STD SPDDSMBN7.7450.10±8.08 DSMLieBN AIM-(1)6.9450.04±8.01 LEM-(1)4.7150.95±6.40 LCM-(1)3.5951.86±4.53 CRIM-(1)16.3550.71±8.1 CRIM-(1.5)19.5151.34±5.82 AIM-(-0.5)8.7153.97±8.78 Table 3.10: Cross-validation results of TSMNet with SPDDSMBN and DSMLieBN on the Hinss data set. If the DSMLieBN under the standard metric (θ = 1) is not saturated, the bottom rows report deformed DSMLieBN. impede the learning of SPDNet on the FPHA data set, they can improve the performance under an appropriate deformation θ. These findings highlight the efficacy of the deforming geometry on the SPD manifold. • Efficiency. Although our LieBN involves additional computations on variance compared with SPDNetBN, our LieBN achieves comparable or even better ef- ficiency than SPDNetBN. In particular, the LieBN induced by standard LEM or LCM exhibits better efficiency than SPDNetBN. Even with deformation, the LCM-based LieBN is still comparable with SPDNetBN in terms of efficiency. This phenomenon could be attributed to the fast and simple computation of LCM and LEM. Application to EEG classification. We apply our method to TSMNet under two scenarios, inter-session and inter-subject. Following Kobler et al. [123], we adopt the architecture of 40, 20. Compared to SPDDSMBN, DSMLieBN-AIM obtains the highest average scores of 55.10% and 53.97% in these two scenarios, outperforming SPDDSMBN by 0.98 and 3.87 percentage points, respectively. In the inter- subject scenario, the efficiency advantage of our LieBN over SPDDSMBN is more evi- dent. Specifically, both the LEM- and LCM-based DSMLieBN achieve similar or better performance compared to SPDDSMBN, while requiring considerably less training time. For example, DSMLieBN-LCM-(1) achieves better results with only half the training time of SPDDSMBN on inter-subject tasks. Interestingly, under the standard AIM, the sole difference between SPDDSMBN and our DSMLieBN is the way of centering and biasing. SPDDSMBN applies the matrix inverse square root and matrix square root to fulfill centering and biasing, while AIM-induced LieBN uses more efficient Cholesky decomposition. As such, the DSMLieBN induced by the standard AIM is more effi- cient than SPDDSMBN, particularly on the inter-subject task. On the other hand, 73 3.2. Lie Group Batch Normalization Figure 3.3: Visualization of input and output 30× 30 SPD matrices in LieBN using 2× 2 Riemannian t-SNE embeddings. The first row shows the input and output under different metrics. Due to the significant difference in magnitude between the t-SNE embeddings of LieBN’s input and output, the second row separately visualizes the LieBN output (at a smaller scale). the CRIM-based LieBN shows less efficiency, due to the relatively complex Riemannian computation of this metric. Visualization. We randomly select 50 samples and visualize the input and output of LieBN on the HDM05 data set. Using Riemannian t-SNE [65], we map the 30× 30 SPD matrices to 2× 2 low-dimensional representations. As shown in Fig. 3.3, LieBN effectively normalizes the data distribution. Specifically, the input t-SNE embeddings are largely scattered and their elements can reach up to 400K, whereas those of the output embeddings are mostly constrained within 20. 3.2.6.2 Experiments of LieBN on Rotation Matrices This subsection implements our LieBN on the special orthogonal groups, i.e., SO(n), also known as rotation matrices. As the Riemannian metric on SO(n) is bi-invariant, there are two instantiations of our LieBN on this group: LieBN-Left based on the left translation and LieBN-Right based on the right translation. We apply our LieBN to the classic LieNet backbone [107], where the latent space is the special orthogonal group. Following Huang et al. [107], we use three action recognition data sets, the G3D [23], HDM05 [153], and NTU60 [178] data sets. We denote the LieNet models 74 Chapter 3. Riemannian Batch Normalization 050100150 Epochs 60 70 80 90 Acc. G3D LieNet LieNetLieBN-Left LieNetLieBN-Right 050100150 Epochs 60 70 80 Acc. HDM05 LieNet LieNetLieBN-Left LieNetLieBN-Right 02550 Epochs 45 50 55 60 65 Acc. NTU60-2Blocks LieNet LieNetLieBN-Left LieNetLieBN-Right 02550 Epochs 45 50 55 60 65 Acc. NTU60-3Blocks LieNet LieNetLieBN-Left LieNetLieBN-Right Figure 3.4: Test accuracy curves corresponding to Tab. 3.11. Method G3DHDM05NTU60 Mean±STD MaxMean±STD Max2Blocks 3Blocks LieNet87.91±0.9089.7376.92±1.2779.1162.460.91 LieNetLieBN-Left88.88±1.6290.6778.89±1.0780.8863.5162.62 LieNetLieBN-Right88.12±1.1290.379.39±1.1380.6763.662.72 Table 3.11: Results of LieNet with or without rotation LieBN. with our LieBN-Left and LieBN-Right as LieNetLieBN-Left and LieNetLieBN-Right, respectively. Results. We conduct 10-fold experiments on the G3D and HDM05 data sets under the suggested 3Blocks 5 and 2Blocks architectures, respectively. On the NTU60 data set, we validate LieBN under the 2Blocks and 3Blocks settings. The results are presented in Tab. 3.11. Due to differences in software, our reimplemented LieNet (in PyTorch) per- forms slightly differently from the results reported by Huang et al. [107] (in MATLAB). However, we still observe a clear improvement when applying our LieBN to the vanilla LieNet backbone. Additionally, LieBN-Right performs slightly better than LieBN-Left. Although the effects of left and right translations on the sample statistics under the bi- invariant metric are identical, their transformations on each sample differ, as illustrated in Fig. 3.1. This difference could slightly affect the network performance. The specific optimal choice of left or right translations depends on the data set’s characteristics. Training dynamics. Fig. 3.4 presents the test accuracy curves. We have the following additional observations, which can be attributed to the mitigated covariate shift by our LieBN, as our LieBN can effectively normalize the sample statistics. • Accelerated convergence. LieBN significantly accelerates the convergence of LieNet. Specifically, on the NTU60 data set, the largest data set involved, LieNet with LieBN converges by the 5th epoch, whereas the vanilla LieNet does not 5 Each block consists of a RotMap layer followed by a RotPooling layer. For more details, please refer to Huang et al. [107]. 75 3.3. Gyrogroup Batch Normalization Data SetSPDNet SPDNetLieBN-Cor ECMLECMOLMLSM HDM0559.13±0.6765.37± 1.0761.35± 0.3460.33± 0.1260.00± 0.27 FPHA85.59±0.7287.20± 0.1287.03± 0.3286.80± 0.1286.77± 0.29 Table 3.12: Results of SPDNet with or without correlation LieBN under different in- variant metrics. converge until the 25th epoch. A similar phenomenon can also be observed on the HDM05 data set. • More stable performance. LieBN enhances the stability of network training. In particular, on the HDM05 and G3D data sets, the initial training fluctuations are greatly mitigated by our LieBN. 3.2.6.3 Experiments of LieBN on Correlation Matrices We apply our correlation LieBN (LieBN-Cor) to SPD networks. Our experiments focus on the SPDNet backbone using the FPHA and HDM05 data sets. LieBN-Cor is applied before the final classification layer. Specifically, SPD features are first activated by the power function, then mapped into correlation matrices via Cor(·), and finally processed by LieBN-Cor. Results. The 5-fold average results are presented in Tab. 3.12. Although LieBN- Cor is not specifically designed for SPD networks, it still improves SPDNet’s per- formance, demonstrating its effectiveness. Among the four invariant metrics, ECM achieves the best performance. LieBN-SPD outperforms LieBN-Cor when applied to SPDNet, as expected because SPDNet is tailored to SPD matrices. However, this does not undermine the validity of LieBN-Cor. The consistent improvement over vanilla SPDNet highlights the potential of applying LieBN-Cor to correlation manifolds. 3.3 Gyrogroup Batch Normalization 3.3.1 Introduction Although LieBN can normalize sample statistics, many important geometries in ma- chine learning do not admit a Lie group structure. As a result, existing methods still lack a principled solution for Riemannian normalization. Recently, gyro-structures have emerged as effective tools for building Riemannian networks across various geometries, 76 Chapter 3. Riemannian Batch Normalization M GyroBN Figure 3.5: Illustration of GyroBN on manifold-valued data. Blue points, green points, and the red dashed curves indicate the input samples, normalized outputs, and data distributions, respectively. Method Controllable Statistics Applied GeometriesIncorporated by GyroBN SPDBN [31]MSPD manifolds under AIM✓ SPDBN [124]M+VSPD manifolds under AIM✓ SPDDSMBN [123]M+VSPD manifolds under AIM✓ ManifoldNorm [34, Algs. 1–2]N/ARiemannian homogeneous spaces✗ ManifoldNorm [34, Algs. 3–4]M+V Matrix Lie groups under the distance d(X,Y ) =∥log (X −1 Y )∥ ✓ RBN [142, Alg. 2]N/AGeodesically complete manifolds✗ LieBN (Sec. 3.2)M+V Lie groups under invariant metrics ✓ GyroBNM+V Pseudo-reductive gyrogroups with gyroisometric gyrationsN/A Table 3.13: Comparison of previous RBN methods with GyroBN, where M and V denote the sample mean and variance. including SPD [157], Grassmannian [157], hyperbolic [76], and spherical manifolds [182]. They naturally extend Euclidean vector structures while encompassing Lie groups and non-group geometries. For instance, the Grassmannian, hyperbolic, and spherical man- ifolds do not form Lie groups but instead form gyrogroups. Based on the analysis above, this part of the thesis first introduces the pseudo- reductive gyrogroup, a relaxation of the classical gyrogroup that provides a broader algebraic foundation for Riemannian normalization. Building on this structure, we de- velop GyroBN, a general RBN framework on pseudo-reductive gyrogroups, as illustrated in Fig. 3.5. We employ gyrosubtraction, gyroaddition, and scalar gyromultiplication to generalize the centering (vector subtraction), biasing (vector addition), and scaling (scalar multiplication) in Euclidean BN to curved manifolds in a principled manner. We clarify why centering and biasing in GyroBN rely on left gyroaddition, rather than other candidates, such as right gyroaddition or gyrocoaddition [200, Def. 2.9]. We show that when gyrations are gyroisometries, GyroBN enjoys theoretical control over sample statistics. These conditions are satisfied by all known gyrogroups in machine learn- 77 3.3. Gyrogroup Batch Normalization ing, providing a principled and unified normalization mechanism. Moreover, several existing RBN methods arise as special cases of GyroBN, including LieBN and various SPD-based variants, as summarized in Tab. 3.13. Beyond the LieBN instantiations in Sec. 3.2.5, we instantiate GyroBN on seven rep- resentative geometries: the Grassmannian [19], five CCSs [76, 131, 13], and the full-rank correlation manifold [195]. For the Grassmannian, we propose an efficient implemen- tation. For CCSs, we cover five models: Poincaré ball, hyperboloid, Beltrami–Klein, sphere, and projected hypersphere. To enable these instantiations, we refine the pro- jected hypersphere structure [13], derive closed-form gyro-structures for the hyperboloid and sphere, and develop the Riemannian structure of the Beltrami–Klein model. For the correlation manifold, we demonstrate that its gyro-structure can be defined row- wise on its Cholesky factor. We also provide a PyTorch-compatible toolbox [169] with drop-in GyroBN layers, illustrated in Fig. 3.6. Experiments on networks over these seven geometries validate the effectiveness of our framework. In summary, the main contributions of this part are • Theoretical foundation: pseudo-reductive gyrogroups as a relaxation of classical gyrogroups; • General framework: GyroBN as a plug-and-play normalization mechanism, with pseudo-reduction and gyroisometric gyrations ensuring theoretical control of batch statistics; • Geometric insights: refined projected hypersphere gyro-structure, closed-form hy- perboloid and sphere gyro-structures, Riemannian structure of the Beltrami–Klein model, and row-wise correlation manifold gyro-structure; • Practical instantiations: implementations on the Grassmannian, five CCSs, and the correlation manifold with extensive experiments. 6 Outline. Sec. 3.3.2 introduces pseudo-reductive gyrogroups and analyzes their the- oretical properties, while Sec. 3.3.3 develops the GyroBN framework. Sec. 3.3.4 shows that prior RBN methods are special cases and instantiates GyroBN on seven represen- tative geometries. Sec. 3.3.5 reports experiments that validate GyroBN across these geometries. Proofs are deferred to Sec. B.3. 78 Chapter 3. Riemannian Batch Normalization from GyroBN import * from GyroBN.Geometry import * # ==== Grassmannian ==== manifold = GrassmannianGyro(n=50, p=10) X_gr = manifold.random_normal (30, 50, 10) gybn_gr = GyroBNGr(shape =[50, 10]) out_gr = gybn_gr(X_gr) # ==== Five CCSs ==== models = [ ("Poincare", Stereographic(K=-1.0)), ("Hyperboloid", Hyperboloid(K=-1.0)), ("Klein", Klein(K=-1.0)), ("Sphere", Sphere(K= 1.0)), ("ProjSphere", Stereographic(K= 1.0)), ] for name , manifold in models: X_ccs = manifold.random_normal (30, 16) gybn_ccs = GyroBNCCS(shape =[16], model=name , K=manifold.K) out_ccs = gybn_ccs(X_ccs) # ==== Full -Rank Correlation ==== manifold = CorPolyHyperbolicCholeskyMetric(n=10) X_cor = manifold.random (30, 10, 10) gybn_cor = GyroBNCor(shape =[10, 10]) out_cor = gybn_cor(X_cor) Figure 3.6: Minimal examples of applying GyroBN. 3.3.2 Pseudo-Reductive Gyrogroups Given a gyrogroup (G,⊕), the left gyrotranslation by x∈ G is defined as L x : G→ G, L x (y) = x⊕ y, ∀y ∈ G.(3.32) If any gyrotranslation is a gyroisometry, we can use gyrotranslation to center manifold- valued samples for the normalization layer. Nguyen and Yang [159] shows that any left gyrotranslation on the SPD and Grassmannian manifolds is a gyroisometry. However, the proof relies on the left cancellation law of gyrogroups, which does not hold for non-reductive gyrogroups, such as the Grassmannian. Here, non-reductive gyrogroups refer to gyro-like groupoids that satisfy the first three gyrogroup axioms but fail the 6 The code is available at https://github.com/GitZH-Chen/GyroBN.git. 79 3.3. Gyrogroup Batch Normalization Invariance of gyronorm under any gyration Left cancellation law Left gyrotranslation law Axiom (G1-3) Pseudo-reduction Gyroisometries of any gyration and left gyrotranslation Gyroisometry of the gyroinverse Gyrocommutativity Axiom (G1-3) Leftreduction (G4) Ours: Previous: Or Figure 3.7: A conceptual comparison of the derivation logic in our work with that in previous work [159], where the left gyrotranslation law is presented in Thm. 190. The previous work proves the results on the SPD and Grassmannian manifolds in a case- by-case manner. In contrast, we relax the left reduction into pseudo-reduction and give a general analysis. Our framework also corrects the proof for the Grassmannian cases. last axiom in Thm. 51, namely the left reduction law (G4), gyr[x,y] = gyr[x⊕ y,y]. Therefore, the proof is not generally valid for the Grassmannian. We propose an inter- mediate structure, referred to as pseudo-reductive gyrogroups, which supports the left cancellation law and, therefore, the gyroisometry of gyrotranslation. This structure forms the algebraic foundation for building the normalization layer. As illustrated in Fig. 3.7, our derivation extends the case-by-case approach by Nguyen and Yang [159] to general pseudo-reductive gyrogroups. 3.3.2.1 From Gyrogroups to Pseudo-Reductive Gyrogroups Definition 73 (Pseudo-reductive gyrogroups). A groupoid (G,⊕) is a pseudo- reductive gyrogroup if it satisfies the axioms (G1), (G2), (G3) and the following pseudo-reductive law: gyr[a,x] = id, for any left inverse a of x in G,(3.33) where id is the identity map. Eq. (3.33) can be intuitively viewed as an intermediate between reduction and non- reduction. For gyrogroups, Eq. (3.33) can be directly obtained from left gyroassociativ- ity (G3) and reduction (G4) [200, Thm. 2.10, item 3]. However, there is no theoretical guarantee that Eq. (3.33) holds for non-reductive gyrogroups. Therefore, we name Eq. (3.33) pseudo-reduction. Nevertheless, for the specific non-reductive Grassman- nian, it is indeed pseudo-reductive. 80 Chapter 3. Riemannian Batch Normalization Proposition 74. [↓] Gr(p,n) and f Gr(p,n) are pseudo-reductive gyrocommutative gyrogroups. Our pseudo-reductive gyrogroup naturally generalizes the vanilla gyrogroup, as it shares most of the basic properties of gyrogroups [200, Thms. 2.10–2.11]. Theorem 75 (First pseudo-reductive gyrogroup properties). [↓] Let (G,⊕) be a pseudo-reductive gyrogroup. For any elements x,y,z,a∈ G, we have: (1) If x⊕ y = x⊕ z, then y = z (General Left Cancellation law; see Case (8) below). (2) gyr[e,x] = id for any left identity e in G. (3) gyr[a,x] = id for any left inverse a of x in G. (4) There is a left identity that is a right identity. (5) There is only one left identity. (6) Every left inverse is a right inverse. (7) There is only one left inverse, ⊖x, of x, and ⊖(⊖x) = x. (8) The left cancellation law: ⊖x⊕ (x⊕ y) = y. (9) The gyrator identity: gyr[x,y]a =⊖(x⊕ y)⊕x⊕ (y⊕ a). (10) gyr[x,y]e = e. (11) gyr[x,y](⊖a) =⊖ gyr[x,y]a. (12) gyr[x,e] = id. (13) The gyrosum inversion law: ⊖(x⊕ y) = gyr[x,y](⊖y⊕⊖x). Remark 76. In non-reductive gyrogroups, the identities in Cases (2) and (3) are not guaranteed to hold. Consequently, any property relying on them, such as Case (4) and those from Case (6) to Case (10), is also not guaranteed to hold. The absence of these basic properties undermines the rationality of non-reductive gyrogroups. In contrast, our pseudo-reductive gyrogroups preserve most of the fundamental properties of gyrogroups. 81 3.3. Gyrogroup Batch Normalization 3.3.2.2 Isometries over Pseudo-Reductive Gyrogroups The gyro-structure in the following is assumed to be defined as Eq. (2.77)–Eq. (2.83). We first clarify that the Riemannian distance agrees with the gyrodistance, and the Riemannian isometry agrees with the gyroisometry. These justify gyrodistance and gyroisometry for gyrospaces over manifolds. Lemma 77 (Distances). [↓] Given a pseudo-reductive gyrogroup (M,⊕), we have d(x,y) =∥Log x (y)∥ x =∥⊖x⊕ y∥ gyr = d gyr (x,y), ∀x,y ∈M,(3.34) where d denotes the geodesic distance. 7 Lemma 78 (Isometries). [↓] Let (M,⊕) and ( f M, e ⊕) be two pseudo-reductive gy- rogroups. Their gyro identity elements are e ∈ M and e ∈ f M, respectively. If φ :M→ f M is a Riemannian isometry with e = φ(e), then the following hold. (1) The Riemannian isometry is a gyroisometry: d gyr (x,y) = g d gyr (φ(x),φ(y)),(3.35) where d gyr and g d gyr are the gyrodistances over M and f M, respectively. (2) If the gyroinverse, gyration, or left gyrotranslation over M is a gyroisometry, its counterpart over f M is also a gyroisometry. Thm. 77 implies that, as long as the gyro-structure is defined by Eq. (2.77)–Eq. (2.83), the gyrodistance coincides with the geodesic distance. Unless otherwise specified, we shall not distinguish between the two and uniformly denote them by d(·,·). Besides, the second result in Thm. 78 is particularly useful, as several geometries are isometric, such as the ONB and P Grassmannian, as well as different models in hyperbolic geometry. Now, we analyze gyroisometries over pseudo-reductive gyrogroups. The most related property in Thm. 75 is the left cancellation law, one of the key prerequisites for a gyrotranslation to be a gyroisometry. Note that the left cancellation comes from left gyroassociativity and Eq. (3.33) [200, Thm. 2.10, item 9]. Therefore, left cancellation does not generally hold for non-reductive gyrogroups but exists in pseudo-reductive 7 On Cartan–Hadamard manifolds, the statement holds for all x,y ∈ M. More generally, the equality requires x,y to lie within a geodesic ball of convexity radius to ensure the well-definedness of the minimizing geodesic and logarithm. In this section, we implicitly assume these conditions are satisfied. 82 Chapter 3. Riemannian Batch Normalization gyrogroups. We first present an if-and-only-if statement about gyroisometry, which will be useful in the following. Theorem 79. [↓] Given a pseudo-reductive gyrogroup (G,⊕), gyr[x,y] preserves the gyronorm for any x,y ∈ G if and only if gyr[x,y] is a gyroisometry for any x,y ∈ G. The gyroisometry of any gyration is a prerequisite for other operators to be gyroi- sometries. Theorem 80 (Gyroisometries). [↓] Given a pseudo-reductive gyrogroup (G,⊕) with any gyr[·,·] as a gyroisometry, we have the following. (1) The left gyrotranslation is a gyroisometry. (2) If (G,⊕) is gyrocommutative, then any gyroinverse is a gyroisometry. Now, we discuss the gyroisometries for the gyro-structures reviewed in Tabs. 2.5, 2.11 and 2.13. Theorem 81. [↓] For the pseudo-reductive gyrogroups corresponding to the SPD manifold (under AIM, LEM, and LCM), the ONB and P Grassmannian, and the stereographic model with K ≤ 0 (the Poincaré ball for K < 0 and Euclidean space for K = 0), the gyrodistance coincides with the geodesic distance. Moreover, the gyroinverse, any gyration, and any left gyrotranslation are gyroisometries. Credit and sketch of the proof. As Thm. 77 already shows that the gyrodistance coin- cides with the geodesic distance, it remains to establish the isometries. For the Grass- mannian and SPD manifolds, these were proved by Nguyen and Yang [159, Thms. 2.12– 2.14 and 2.16–2.18], although their arguments implicitly treated the non-reductive Grassmannian as a gyrogroup by using left cancellation. Our Thms. 74 and 75 con- firms that the Grassmannian is pseudo-reductive and does satisfy left cancellation, thereby validating their results. However, we can directly establish these properties from Thms. 79 and 80. The complete proof is given in Sec. B.3.7. Remark 82. The remaining constant-curvature manifolds reviewed in Sec. 2.9.5, including the stereographic model with K > 0, radius model, and Beltrami– Klein model, also satisfy these properties. Their verifications will be presented in Sec. 3.3.4. 83 3.3. Gyrogroup Batch Normalization 3.3.3 GyroBN on Pseudo-Reductive Gyrogroups Building on Thm. 81, which establishes that several geometries admit isometric gy- rotranslations, we develop RBN in a principled way for general pseudo-reductive gy- rogroups, referred to as GyroBN. Throughout, we assume (M,⊕) is a pseudo-reductive gyrogroup with the gyro-structure defined by Eq. (2.77)–Eq. (2.83). 8 As Thm. 77 es- tablishes the equivalence between gyrodistance and geodesic distance, we use the terms “gyromean” and “gyrovariance” interchangeably with their Riemannian counterparts. 3.3.3.1 GyroBN To generalize the Euclidean BN in Eq. (3.1) to gyrogroups, we first define sample mean, sample variance, centering, biasing, and scaling over gyrogroups. Then, we introduce the GyroBN framework with a theoretical analysis of the ability to normalize sample statistics. We define the gyromean as the Fréchet mean [74] under gyrodistance: μ = FM(x i ∈M N i=1 ) = argmin y∈M 1 N X N i=1 d 2 (x i ,y).(3.36) The gyrovariance is the corresponding Fréchet variance. By Thm. 77, the gyromean and gyrovariance coincide with the Riemannian mean and variance, i.e., the Fréchet mean and variance under geodesic distance. The existence and local uniqueness of the Fréchet mean are reviewed in Thm. 32. Easy computation shows that the Euclidean BN operations in Eq. (3.1) have direct gyrogroup counterparts in R n . Centering corresponds to gyrosubtraction (Eqs. (2.77) and (2.79)), biasing to gyroaddition (Eq. (2.77)), and scaling to scalar gyromultiplica- tion (Eq. (2.78)). Motivated by this, we define a normalization layer over gyrogroups via gyro operations. Given a batch of activations x i N i=1 ⊂M, the core operations of GyroBN are ∀i≤ N, ̃x i ← Biasing z| β⊕ Scaling z | s √ v 2 + ε ⊙ Centering z| ⊖μ⊕ x i ,(3.37) where μ ∈ M and v 2 are gyromean and gyrovariance, β ∈ M is the bias parameter, s ∈ R is the scaling parameter, and ε is a small value for numerical stability. The following theorem shows that Eq. (3.37) can normalize manifold-valued data. 8 In GyroBN, ⊙ is not required to satisfy the axioms of a gyrovector space (Thm. 53). 84 Chapter 3. Riemannian Batch Normalization Algorithm 2: Gyrogroup Batch Normalization (GyroBN) Require : batch of activations x i N i=1 ⊂M, small positive constant ε, and momentum η ∈ [0, 1], running mean μ r , running variance v 2 r , bias parameter β ∈M, scaling parameter s∈ R. Return : normalized batch ̃x i N i=1 ⊂M if training then Compute batch mean μ b and variance v 2 b of x i N i=1 ; Update running statistics μ r = Bar η (μ b ,μ r ), and v 2 r = ηv 2 b + (1− η)v 2 r ; (μ,v 2 ) = (μ b ,v 2 b ) if training else (μ r ,v 2 r ) ∀i≤ N, ̃x i = β⊕ s √ v 2 +ε ⊙ (⊖μ⊕ x i ) Theorem 83 (Homogeneity). [↓] Let (M,⊕) be a pseudo-reductive gyrogroup in which every gyration gyr[·,·] is a gyroisometry. For N samples x i N i=1 ⊂ M and any t∈ R, we have Homogeneity of gyromean: FM(β⊕ x i N i=1 ) = β⊕ FM(x i N i=1 ), ∀β ∈M, (3.38) Homogeneity of dispersion from e: 1 N X N i=1 d 2 (t⊙ x i ,e) = t 2 N X N i=1 d 2 (x i ,e), (3.39) The most important property of the Euclidean BN [109] lies in its ability to normalize the sample mean and variance. Thm. 83 shows that the formulation in Eq. (3.37) enjoys the same property: homogeneity of the gyromean guarantees that centering and biasing shift the gyromean, while homogeneity of the dispersion from e ensures that scaling controls the sample variance. As a result, GyroBN provides a theoretical guarantee of normalization on any pseudo-reductive gyrogroup with isometric gyrations. Moreover, since the gyromean and gyrovariance coincide with their Riemannian counterparts, GyroBN also normalizes Riemannian statistics. To finalize GyroBN, we define the running mean updates over gyrogroups as the binary barycenter based on gyrodistance: Bar η (x 1 ,x 2 ) = argmin y∈M η d 2 (x 1 ,y) + (1− η) d 2 (x 2 ,y) , η ∈ [0, 1],(3.40) which can be calculated by the geodesic. With these ingredients, the general framework for GyroBN is presented in Alg. 2. In particular, it recovers the classic Euclidean 85 3.3. Gyrogroup Batch Normalization BN [109] when M = R n . Remark 84. We make the following two remarks with respect to left gyrotranslation. • Other candidates. There are three alternatives to left gyrotranslation. However, they are not necessarily gyroisometries and therefore cannot sup- port the general GyroBN construction without additional isometry assump- tions. Specifically, analogous to the left gyrotranslation, the right gyrotrans- lation is R x : G→ G, R x (y) = y⊕ x, ∀y ∈ G.(3.41) Along with the gyroaddition, there is the gyrogroup coaddition [200, Def. 2.9]: x⊞ y = x⊕ gyr[x,⊖y]y, ∀x,y ∈ G.(3.42) Coaddition is symmetric to gyroaddition in many ways; for instance, when gyroaddition is gyrocommutative, coaddition is commutative [200, Thm. 3.3]. However, the right gyrotranslation, as well as the left and right translations by coaddition, are not guaranteed to be gyroisometries. Numerical experiments confirm that these three translations on the hyperbolic Poincaré ball fail to preserve gyrodistance (see gyrocoadd.py). A theoretical reason is that they all lack a counterpart of the left gyrotranslation law, which is important for the translation to be a gyroisometry (see Fig. 3.7). • Special cases. For Lie groups with a right-invariant Riemannian metric, right translations are isometries, and GyroBN can then be formulated us- ing them, with Thm. 83 extending directly. In general, however, only left gyrotranslations are guaranteed to be isometries. 3.3.4 Instantiations As indicated by Thms. 81 and 83, GyroBN can be applied to different geometries with guaranteed normalization of the sample statistics. Once the required operators are specified, Alg. 2 can be used in a plug-and-play manner. We first show that existing RBN methods with control over sample statistics, such as LieBN on Lie groups and AIM-based SPDBNs, are special cases of the framework. We then instantiate GyroBN on seven representative geometries: the Grassmannian, five CCS models (Poincaré ball, projected hypersphere, hyperboloid, sphere, and Beltrami–Klein), and the full-rank 86 Chapter 3. Riemannian Batch Normalization correlation manifold. To support these instantiations, we simplify the Grassmannian operators for efficient computation, refine the gyro-structures on the Poincaré ball and projected hypersphere, establish new gyro-structures for the hyperboloid and sphere, characterize the Beltrami–Klein geometry, and formulate a row-wise realization for the correlation manifold. 3.3.4.1 LieBN as a Special Case Chakraborty [34, Algs. 3–4] introduced the Riemannian normalization on matrix Lie groups under a specific distance. The extension to general Lie groups, yielding LieBN with theoretical normalization over the Riemannian mean and variance, is presented in Sec. 3.2. This subsection shows that LieBN is a special case of GyroBN. LieBN is formulated with an invariant metric on a Lie group. Centering and bias- ing are performed through group translations, while scaling is defined in the tangent space at the identity element. Since every Lie group is automatically a gyrogroup, gyrotranslation reduces exactly to the group translation. Consequently, the centering, biasing, and scaling in LieBN are identical to those in GyroBN. Moreover, as shown in Thm. 77, the mean, variance, and running mean update defined via the geodesic dis- tance in LieBN are equivalent to their counterparts based on gyrodistance in GyroBN. Therefore, LieBN is a special case of GyroBN. The LieBN instantiations presented in Sec. 3.2.5 are therefore incorporated by GyroBN. 3.3.4.2 AIM-Based SPDBNs as Special Cases As reviewed in Sec. 3.2.3.2, several RBNs on the SPD manifold were developed based on AIM [31, 124, 123]. These SPD normalization methods can be expressed in the following unified form: Normalization: ∀i≤ N, e P i ← B 1 2 M − 1 2 P i M − 1 2 s √ v 2 +ε B 1 2 ,(3.43) where M and v 2 are the Riemannian mean and variance. The running mean is updated by the binary barycenter under the geodesic distance. As gyrodistance is identical to the geodesic distance, the gyromean, gyrovariance, and running mean updates are identical to the Riemannian ones. Using the AIM gyro- operations in Eq. (2.95), Eq. (3.43) is exactly the specific implementation of Eq. (3.37) under the AIM-based gyrogroup on the SPD manifold. Therefore, the SPDBNs devel- oped by Brooks et al. [31], Kobler et al. [124, 123] are also special cases of GyroBN. 87 3.3. Gyrogroup Batch Normalization Remark 85. Brooks et al. [31] only considers centering and biasing. Kobler et al. [124] uses a running mean for centering during training. Kobler et al. [123] uses dif- ferent momentum parameters to update running statistics for training and testing, along with multi-channel mechanisms for domain adaptation. Nevertheless, all of them are based on Eq. (3.43). Therefore, tricks such as multi-channel and separate momentum can also be applied to GyroBN. This is what we mean by claiming that GyroBN incorporates their approaches. 3.3.4.3 Grassmannian Manifold We focus on the ONB perspective. Let U,V ∈ Gr(p,n), t ∈ R, and ∆ ∈ T U Gr(p,n). Using the Grassmannian operators reviewed in Tabs. 2.10 and 2.11, we instantiate GyroBN on the ONB Grassmannian. Instantiation. Building on these operators, we now implement the ONB Grass- mannian GyroBN. Given a batch of activationsU 1·N , the three core steps of GyroBN are Centering to the identity I p,n : U 1 i = exp −[ M ⊤ , e I p,n ] U i ,(3.44) Scaling the dispersion from I p,n : U 2 i = exp s √ v 2 + ε [ U 1 i (U 1 i ) ⊤ , e I p,n ] I p,n , (3.45) Biasing towards B ∈M: U 3 i = exp [ B, e I p,n ] U 2 i .(3.46) Here(·) = g Log e I p,n (·) is the Riemannian logarithm under the P Grassmannian, M (resp. v 2 ) is the Riemannian batch mean (resp. variance), and e I p,n = I p,n I ⊤ p,n is the P identity. The mean M can be obtained by the Karcher flow [116], with the Riemannian logarithm computed by Bendokat et al. [19, Alg. 5.3]. Efficient computation. The commutators [M ⊤ , e I p,n ] and [U 1 i (U 1 i ) ⊤ , e I p,n ] can be efficiently computed by the following result. Proposition 86. [↓] Given U = (U ⊤ 1 ,U ⊤ 2 ) ⊤ ∈ Gr(p,n) with U 1 ∈ R p×p and U 2 ∈ R (n−p)×p , then [ U ⊤ , e I p,n ] = 0 p×p − e U ⊤ 2 e U 2 0 (n−p)×(n−p) ! ,(3.47) where e U 2 = U 2 Q arcsin( ˆ S) ˆ S R ⊤ and U ⊤ 1 SVD := QSR ⊤ . Here S is in ascending order, Q and R are flipped column-wise, and ˆ S = p I p − S 2 . 88 Chapter 3. Riemannian Batch Normalization Remark 87. Two technical issues are worth noting. • Cut locus. The logarithm Log U (V ) exists only when U and V are not in each other’s cut locus [19]. Similarly, gyroaddition and gyromultiplication are not globally defined due to the cut locus [157, Sec. 3.2]. However, Bendokat et al. [19, Alg. 5.3] provides a numerical remedy. • P Grassmannian. Although our derivation is based on the ONB Grass- mannian, GyroBN under the P Grassmannian can be obtained by mapping data via π −1 : f Gr(p,n) → Gr(p,n), normalizing, and mapping back via π. This follows from the isometry π : Gr(p,n)→ f Gr(p,n). 3.3.4.4 Stereographic Model We begin by analyzing its gyro-structure and then instantiate GyroBN. Stereographic gyrovector space. As reviewed in Sec. 2.9.5, the stereographic model st n K unifies CCS geometries: the hyperbolic Poincaré ball P n K for K < 0, Eu- clidean space R n for K = 0, and the spherical projected hypersphere D n K for K > 0. For x,y,z ∈ st n K , t∈ R, and v ∈ T x st n K , its Riemannian and gyro operators are reviewed in Tabs. 2.13 and 2.14. For K < 0, the Poincaré ball (P n K ,⊕ K ,⊙ K ) forms a Möbius gyrovector space, as reviewed in Tab. 2.13. For K > 0, however, the situation is subtler [13]: • gyroaddition is well-defined except when x = y K∥y∥ 2 ; • gyromultiplication is well-defined except when r tan −1 ( √ K∥x∥) = π /2 + kπ for some k ∈ Z. Even if assumed well-defined, it remains unclear whether (st n K ,⊕ K ,⊙ K ) with K > 0 satisfies the axioms of a gyrovector space, in contrast to the Poincaré ball. Closing this gap is part of our contribution: under the assumption of well-definedness, we show that the stereographic gyro operations coincide with those in Eqs. (2.77) and (2.78) and prove that the stereographic model satisfies all axioms of a gyrovector space. Proposition 88. [↓] The stereographic gyroaddition and gyromultiplication coin- cide with the Riemannian definitions: x⊕ K y = Exp x (PT 0→x (Log 0 (y))), ∀x,y ∈ st n K ,(3.48) t⊙ K x = Exp 0 (t Log 0 (x)), ∀t∈ R,x∈ st n K .(3.49) 89 3.3. Gyrogroup Batch Normalization Theorem 89. [↓] For any K ∈ R, the stereographic model (st n K ,⊕ K ) satisfies all axioms of a gyrocommutative gyrogroup. When further endowed with gyromultipli- cation ⊙ K , it satisfies all axioms of a gyrovector space. Stereographic GyroBN. Next, we extend Thm. 81 to the stereographic model with arbitrary curvature. Theorem 90. [↓] For the stereographic model, the gyrodistance coincides with the geodesic distance. Moreover, the gyroinverse, any gyration, and any left gyrotrans- lation are gyroisometries. Combining Thm. 83 and Thm. 90, GyroBN in the stereographic model is theo- retically guaranteed to normalize sample statistics. Practically, implementation only requires substituting the stereographic operators reviewed in Tabs. 2.13 and 2.14 into Alg. 2. For efficient computation, the Poincaré Fréchet mean can be obtained using the algorithm of Lou et al. [142, Alg. 1], while the mean on the sphere is computed via the Karcher flow [116]. 3.3.4.5 Radius Model As reviewed in Sec. 2.9.5, the radius model M n K is another representation of constant- curvature spaces, unifying the hyperboloid H n K for K < 0, the sphere S n K for K > 0, and the Euclidean space R n for K = 0. The case K = 0 reduces trivially to Euclidean space, so the following analysis of the radius-model gyro-structure considers K ̸= 0. Although this model has been effective in various applications [38, 45, 15, 166, 99, 119], its gyro-structure has not been formalized. We first analyze the gyro-structure on M n K and then instantiate GyroBN. Radius gyrovector space. The radius modelM n K is isometric to the stereographic model st n K via stereographic projection fixing the south pole [182]: π M n K →st n K :M n K ∋ " ξ ∈ R x∈ R n # 7−→ x 1 + p |K|ξ ∈ st n K ,(3.50) π st n K →M n K : st n K ∋ y 7−→ 1 √ |K| 1−K∥y∥ 2 1+K∥y∥ 2 2y 1+K∥y∥ 2 ∈M n K .(3.51) The origin in M n K is defined as 0 = [ p 1/|K|, 0,..., 0] ⊤ , corresponding to 0 ∈ st n K . For K > 0, we have M n K = S n K and st n K = D n K . Since Eq. (3.50) is undefined at the 90 Chapter 3. Riemannian Batch Normalization south pole −0 when K > 0, we use the one-point compactification D n K ∪∞ with the identification π M n K →st n K (− 0) = ∞ [182, Rmk. A.9]. For simplicity, we use D n K and D n K ∪∞ interchangeably. Tab. 2.14 reviews the Riemannian operators. We adopt the following curvature- aware functions: sin K = sin if K > 0, sinh if K < 0, cos K = cos if K > 0, cosh if K < 0, ⟨·,·⟩ K = ⟨·,·⟩ if K > 0, ⟨·,·⟩ L if K < 0. (3.52) For v ∈ T x M n K , ∥v∥ K = p ⟨v,v⟩ K is the induced Riemannian norm. On the K < 0 branch, the ambient Lorentzian form is positive definite only after this tangent-space restriction. Moreover, (·) s denotes the space vector, and (·) t denotes the time scalar. Using the radius-model Riemannian operators reviewed in Tab. 2.14, we define the gyroaddition and gyromultiplication as Eqs. (2.77) and (2.78): x⊕ M K y = Exp x (PT 0→x (Log 0 (y))), ∀x,y ∈M n K ,(3.53) t⊙ M K x = Exp 0 (t Log 0 (x)), ∀t∈ R,∀x∈M n K .(3.54) We give the following clarifications regarding the above two gyro operations. • For hyperbolic geometry (K < 0), Eq. (3.53) has been employed in prior work [38, 99]. However, there exists no closed-form expression, which could be more efficient than the composition of Riemannian operators. Besides, the underlying gyro- structure has not been formally discussed. • For spherical geometry (K > 0), the geodesic between antipodal points (x and−x) is not unique, making the logarithm and parallel transport along such a geodesic ill-defined. For the sphere, Eq. (3.53) assumes that x ̸= − 0 and y ̸= −0, while Eq. (3.54) assumes that x̸=−0. Like before, we always make these assumptions implicitly. In the following, we first give the closed-form expressions of Eqs. (3.53) and (3.54). Then, we show that Eqs. (3.53) and (3.54) conform to all the axioms of a gyrovector space. 91 3.3. Gyrogroup Batch Normalization Proposition 91 (Gyromultiplication and gyroinverse). [↓] Let x = [x t ,x ⊤ s ] ⊤ be a point in M n K , where x t ∈ R is the time scalar, and x s ∈ R n is the spatial part. The gyromultiplication and inverse have closed-form expressions: t⊙ M K x = 0,t = 0∨ x =0, 1 √ |K| cos K t cos −1 K ( p |K|x t ) sin K t cos −1 K ( p |K|x t ) ∥x s ∥ x s , t̸= 0, (3.55) ⊖ M K x =−1⊙ M K x = " x t −x s # .(3.56) In particular, the gyro identity is 0. Besides, the following shows that the sphere gyromultiplication is still valid for the singular cases in the projected hypersphere. Assume K > 0 and x ̸= ±0, and let st n K ∋ u = π M n K →st n K (x) be its stereographic image. For t ∈ R, set θ = cos −1 √ Kx t ∈ (0,π). Then the following are equiva- lent: (i) t tan −1 √ K∥u∥ = π 2 + kπ ⇐⇒ (i) tθ = (2k + 1)π, k ∈ Z.(3.57) In these singular cases, Eq. (3.55) is still valid, while the stereographic gyromulti- plication returns infinity: t⊙ M K x =−0∈M n K , t⊙ K u =∞∈ st n K ,(3.58) where∞ denotes the added point in the one-point compactification D n K ∪∞ (st n K = D n K for K > 0) and π M n K →st n K (−0) =∞. Proposition 92 (Gyroaddition). [↓] Let x = [x t ,x ⊤ s ] ⊤ and y = [y t ,y ⊤ s ] ⊤ be points in M n K , where x t ,y t ∈ R are the time scalars, and x s ,y s ∈ R n are the spatial parts. 92 Chapter 3. Riemannian Batch Normalization Then, the gyroaddition on M n K admits the closed form: x⊕ M K y = x,y =0, y,x =0, 1 √ |K| D−KN D+KN 2(A s x s +A y y s ) D+KN , otherwise. (3.59) Here, A s = ab 2 − 2Kbs xy − Kan y and A y = b(a 2 + Kn x ), where the following quantities are defined by a = 1 + p |K|x t ,b = 1 + p |K|y t ,n x =∥x s ∥ 2 ,n y =∥y s ∥ 2 ,s xy =⟨x s ,y s ⟩. (3.60) D = a 2 b 2 − 2Kabs xy + K 2 n x n y ,N = a 2 n y + 2abs xy + b 2 n x .(3.61) Besides, the following shows that the sphere gyroaddition is still valid under the singular cases in the projected hypersphere. Assume K > 0 and x,y ̸=± 0, and let u = π M n K →st n K (x) and v = π M n K →st n K (y) be the stereographic images. The following statements are equivalent: (1) u = v K∥v∥ 2 (v ̸= 0); (2) x s = y s and x t =−y t (same meridian, mirrored across the equator); (3) D = 0. In such singular cases, we have N > 0 and Eq. (3.59) is still valid, while the stereographic gyroaddition returns infinity: x⊕ M K y =− 0, u⊕ K v =∞,(3.62) where ∞ is the point added in the one-point compactification of D n K . The above two propositions immediately imply that π M n K →st n K preserves the gyro operations. Corollary 93 (Isomorphism). [↓] For the hyperbolic geometry (K < 0), the isom- 93 3.3. Gyrogroup Batch Normalization etry π M n K →st n K : H n K → P n K preserves the gyro operations: x⊕ M K y = π st n K →M n K π M n K →st n K (x)⊕ K π M n K →st n K (y) , ∀x,y ∈M n K , r⊙ M K x = π st n K →M n K r⊙ K π M n K →st n K (x) , ∀r ∈ R,∀x∈M n K . (3.63) For the spherical geometry (K > 0), the isometry π M n K →st n K : S n K → D n K ∪∞ preserves the gyro operations: x⊕ M K y = π st n K →M n K π M n K →st n K (x)⊕ K π M n K →st n K (y) , ∀x,y ∈M n K /−0, r⊙ M K x = π st n K →M n K r⊙ K π M n K →st n K (x) , ∀r ∈ R,∀x∈M n K /− 0. (3.64) Remark 94. We provide two clarifications regarding Thm. 93. • Compactified projected hypersphere. Although stereographic opera- tions for K > 0 may be undefined in certain cases, they become well-defined on the one-point compactification D n K ∪∞, where undefined cases corre- spond to ∞. • Sphere vs. projected hypersphere. Thm. 93 suggests a numerical advan- tage of the sphere: gyro-operations are well-defined on S n K at all points except the single south pole− 0, including cases corresponding to singularities on the projected hypersphere. This broader domain can make computations on S n K more stable. From the above corollary, it is expected that the operations⊕ M K and⊙ M K also satisfy the axioms of a gyrovector space for both negative and positive curvature K. Theorem 95 (Radius gyrovector spaces). [↓] (M n K ,⊕ M K ) forms a gyrocommutative gyrogroup, and (M n K ,⊕ M K ,⊙ M K ) forms a gyrovector space. a a For K > 0, we implicitly assume the addition and multiplication are well-defined; whenever they are, all corresponding axioms hold. Due to the isometry, Thm. 78 implies gyroisometries over M n K . Theorem 96. On the radius model, the gyrodistance is identical to the geodesic distance, whereas the gyroinverse, gyration, and left gyrotranslation are gyroisome- tries. Radius GyroBN. Thm. 96 guarantees that GyroBN on the radius model normal- 94 Chapter 3. Riemannian Batch Normalization izes the sample mean and variance. Substituting the radius Riemannian operators from Tab. 2.14 and the gyro operators from Thms. 91 and 92 into Alg. 2 can directly yield the radius GyroBN. For K < 0 (hyperboloid), the Fréchet mean can be computed effi- ciently by Lou et al. [142, Alg. 3]; for K > 0 (sphere), we compute the Fréchet mean using the Karcher flow [116]. 3.3.4.6 Hyperbolic Beltrami–Klein The Beltrami–Klein model has recently emerged as a promising alternative to the Poincaré ball for representing hyperbolic geometry [147]. As reviewed in Sec. 2.9.5, the Poincaré ball admits the Möbius gyrovector space and the Beltrami–Klein model admits the Einstein gyrovector space. While prior studies mainly focused on the case K = −1 [147], we develop the Beltrami–Klein Riemannian structure under arbitrary negative curvature and relate it to the Einstein gyrospace. This allows us to estab- lish the equivalence between the gyro and Riemannian formulations and, ultimately, to instantiate GyroBN on this model. Beltrami–Klein Riemannian structure. The standard Einstein gyro operations on the Beltrami–Klein model are reviewed in Tab. 2.13. For K = −1, prior work related the Einstein operations to the Riemannian-form gyro operations in Eqs. (2.77) and (2.78) [147, Sec. 4.2]. We extend this equivalence to arbitrary K < 0 and derive the corresponding closed-form Beltrami–Klein Riemannian operators. To this end, we first establish the isometry between the Beltrami–Klein and Poincaré models. Proposition 97 (Beltrami–Klein isometries). [↓] The following maps are Rieman- nian isometries between the Beltrami–Klein and Poincaré ball models: π K n K →P n K : K n K ∋ x7−→ 1 1 + q 1 + K∥x∥ 2 x∈ P n K ,(3.65) π P n K →K n K : P n K ∋ x7−→ 2 1− K∥x∥ 2 x∈ K n K .(3.66) In particular, π P n K →K n K (0) = 0. Given x in the hyperbolic model H∈K n K , P n K and tangent vector v ∈ T x H, the differential maps of π K n K →P n K and π P n K →K n K are (π K n K →P n K ) ∗,x (v) = 1 1 + q 1 + K∥x∥ 2 v− K⟨x,v⟩ 1 + q 1 + K∥x∥ 2 2 q 1 + K∥x∥ 2 x, 95 3.3. Gyrogroup Batch Normalization (π P n K →K n K ) ∗,x (v) = 2 1− K∥x∥ 2 v + 4K⟨x,v⟩ 1− K∥x∥ 2 2 x. In particular, the differential maps at the zero vector are (π K n K →P n K ) ∗,0 (v) = 1 2 v,(3.67) (π P n K →K n K ) ∗,0 (v) = 2v.(3.68) Moreover, these isometries preserve gyroaddition and gyromultiplication: π P n K →K n K (x⊕ M y) = π P n K →K n K (x)⊕ E π P n K →K n K (y), ∀x,y ∈ P n K , π P n K →K n K (t⊙ M x) = t⊙ E π P n K →K n K (x), ∀t∈ R,∀x∈ P n K , (3.69) where ⊕ M and ⊙ M are Möbius operations, while ⊕ E and ⊙ E are the Einstein coun- terparts. Theorem 98 (Einstein by Beltrami–Klein). [↓] The Einstein gyro operations can be rewritten as Eqs. (2.77) and (2.78): x⊕ E y = Exp x (PT 0→x (Log 0 (y))), ∀x,y ∈ K n K ,(3.70) t⊙ E x = Exp 0 (t Log 0 (x)), ∀x∈ K n K ,∀t∈ R.(3.71) Thm. 98 demonstrates that the Einstein gyro-structure can be expressed by the Beltrami–Klein geometry. Conversely, the Beltrami–Klein geometry can also be formu- lated by the Einstein gyro-structure. Theorem 99 (Beltrami–Klein by Einstein). [↓] Given x,y ∈ K n K and v ∈ T x K n K , the distance, exponential, and logarithmic operators under the Beltrami–Klein ge- ometry are d(x,y) = 2 p |K| tanh −1 p |K| ∥−x⊕ E y∥ 1 + q 1 + K∥−x⊕ E y∥ 2 , Exp x (v) = x⊕ E Exp 0 1 q 1 + K∥x∥ 2 v− K⟨x,v⟩ 1 + q 1 + K∥x∥ 2 (1 + K∥x∥ 2 ) x , 96 Chapter 3. Riemannian Batch Normalization Log x (y) = 1 λ K ex (π P n K →K n K ) ∗,ex (Log 0 (−x⊕ E y)), where ex = π K n K →P n K (x). In particular, the exponential and logarithmic maps at the zero vector 0 are identical across the Beltrami–Klein and Poincaré ball models: Exp 0 (v) = tanh( p |K|∥v∥) v p |K|∥v∥ , ∀v ∈ T 0 H,(3.72) Log 0 (x) = tanh −1 ( p |K|∥x∥) x p |K|∥x∥ , ∀x∈H,(3.73) with H∈K n K , P n K . Remark 100. Since the Beltrami–Klein and Poincaré ball models share the same Exp 0 and Log 0 , it naturally follows that the Einstein and Möbius gyromultiplication coincide. As the Beltrami–Klein model is isometric to the Poincaré ball, Thm. 78 implies gyroisometries. Theorem 101. On the Beltrami–Klein model, the gyrodistance is identical to the geodesic distance, whereas the gyroinverse, gyration, and left gyrotranslation are gyroisometries. Beltrami–Klein GyroBN. Thm. 101 ensures that GyroBN on the Beltrami–Klein model normalizes the sample mean and variance. Owing to the isometry between the Beltrami–Klein and Poincaré models, the Fréchet mean can be computed via the Poincaré ball: map the data to the Poincaré model using π K n K →P n K , compute the Poincaré Fréchet mean [142, Alg. 1], and map the result back using π P n K →K n K . Together with the gyro and Riemannian operators in Sec. 3.3.4.6, we have all the ingredients to implement Alg. 2. 3.3.4.7 Correlation Manifolds As reviewed in Sec. 2.9.2, ECM, LECM, OLM, and LSM are correlation metrics in- duced from Euclidean, zero-curvature prototype spaces, and these four metrics have been instantiated for LieBN in Sec. 3.2.5.3. Here, we further handle the nonzero- curvature correlation metric, namely PHCM. Under PHCM, any correlation matrix can be identified with a product of hyperbolic spaces via its Cholesky decomposition. Given C ∈ Cor + (n), let L = Chol(C) be its Cholesky factor. The k-th row of L has 97 3.3. Gyrogroup Batch Normalization the form (L k1 ,...,L k,k−1 ,L k , 0,..., 0) with L k > 0, which belongs to the hyperbolic open hemisphere 9 HS k−1 = x∈ R k |∥x∥ = 1,x k > 0 .(3.74) As detailed in Sec. 5.4, HS n is isometric to the unit Poincaré ball P n =x∈ R n |∥x∥ < 1(3.75) by π HS n →P n " x x n+1 #! = x 1 + x n+1 .(3.76) Therefore, each correlation matrix can be identified with n− 1 Poincaré vectors: Cor + (n)∋ C7→ 10 ·0 L 21 L 22 ·0 . . . . . . . . . . . . L n1 L n2 · L n 7→ x 1 ∈ P 1 . . . x n−1 ∈ P n−1 .(3.77) Here, x i = π HS i →P i L (i+1,1) ,· ,L (i+1,i+1) ⊤ corresponds to the (i + 1)-th row of the Cholesky factor. Let P n−1 = Q n−1 i=1 P i denote the product of unit Poincaré balls. We denote the identification in Eq. (3.77) by Φ : Cor + (n) → P n−1 . GyroBN on the correlation manifold can then be realized via the Poincaré GyroBN applied row-wise: first map C to P n−1 via Φ, apply GyroBN i independently on each P i , and finally map back with Φ −1 . For a batch of activations C i N i=1 ⊂ Cor + (n), the process can be expressed as ∀i≤ N, C i Φ 7−→ x i 1 ∈ P 1 . . . x i n−1 ∈ P n−1 GyroBN 1 7−→ . . . GyroBN n−1 7−→ ex i 1 ∈ P 1 . . . ex i n−1 ∈ P n−1 Φ −1 7−→ e C i .(3.78) 3.3.4.8 Summary To conclude this section, Tab. 3.14 summarizes the key gyro operators needed to imple- ment GyroBN on representative manifolds. The correlation manifold is excluded, since its GyroBN is realized row-wise through the Poincaré ball. 9 Also known as the Jemisphere model, where the “J” is pronounced as in Spanish [33, Sec. 7]. 98 Chapter 3. Riemannian Batch Normalization U ⊕ Gr V⊖ Gr Ut⊙ Gr UFréchet mean exp(Ω U )Vexp (−Ω U )I p,n exp (tΩ U )I p,n Karcher flow (a) The ONB Grassmannian Operatorst n K M n K K n K x⊕ y (1− 2K⟨x,y⟩− K∥y∥ 2 )x + (1 + K∥x∥ 2 )y 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 Eq. (3.59)Tab. 2.13 ⊖x−x " x t −x s # −x t⊙ x tan K t tan −1 K p |K|∥x∥ p |K| x ∥x∥ Eq. (3.55)Tab. 2.13 Fréchet mean K < 0: [142, Alg. 1] K > 0: Karcher flow K < 0: [142, Alg. 3] K > 0: Karcher flow Via Poincaré ball (see Sec. 3.3.4.6) (b) Constant curvature spaces Table 3.14: Summary of operators for GyroBN across representative manifolds. 3.3.5 Experiments GyroBN layers are backbone-agnostic and can be integrated into different networks whenever the underlying manifold admits the required gyro operations. This subsection evaluates GyroBN on the Grassmannian, five CCSs, and the correlation manifold. The main findings are summarized as follows. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.2. • Numerical experiments (Sec. 3.3.5.1). The closed-form expression de- rived for radius gyroaddition in Eq. (3.59) significantly accelerates the compu- tation compared with its definition-based counterpart in Eq. (3.53), achieving 2×–3× speedups. Furthermore, visualizations demonstrate that GyroBN ef- fectively normalizes sample distributions across diverse geometries. • Performance (Secs. 3.3.5.2 to 3.3.5.4). On networks over the Grassman- nian, CCSs, and the correlation manifold, GyroBN generally improves back- bone networks, whereas existing RBN methods often degrade performance. Compared to these methods, GyroBN is generally faster or comparable in runtime, requires fewer or equal parameters, and improves robustness. 3.3.5.1 Numerical Experiments Efficiency of the closed-form radius gyroaddition. Recalling Sec. 3.3.4.5, we derive closed-form expressions for the radius gyrooperations to improve computational 99 3.3. Gyrogroup Batch Normalization GeometryHyperboloidSphere DimRiemannianClosed formRiemannianClosed form 16361.22121.68 (33.69%)323.58122.02 (37.71%) 32 363.41123.08 (33.87%)327.75123.71 (37.74%) 64387.94181.50 (46.78%)453.42181.05 (39.93%) 128574.37272.10 (47.37%)644.18270.21 (41.95%) 256 1149.58534.09 (46.46%)1137.62538.58 (47.34%) 10243364.231414.67 (42.05%)3551.691421.67 (40.03%) 20486479.952497.47 (38.54%)6930.692449.56 (35.34%) Table 3.15: Efficiency (in μs) of gyroaddition on the radius manifold: closed form versus Riemannian definition. Values in parentheses indicate the runtime of the closed-form implementation as a percentage of the corresponding Riemannian implementation. The best results are bold. efficiency. To assess this, we compare two variants of radius gyroaddition: (i) the definition-based operator Eq. (3.53), implemented via a composition of the Riemannian logarithm, parallel transport, and exponential map; and (i) the closed-form opera- tor Eq. (3.59). We report the mean wall-clock time (in μs), averaged over 100 runs with a batch size of 10,000, across varying dimensions. As shown in Tab. 3.15, the closed-form implementation consistently outperforms its definition-based counterpart, achieving speedups of roughly 2×–3× across the hyperboloid and sphere. Visualization of GyroBN on different geometries. To intuitively illustrate the effect of GyroBN, we visualize its behavior on the ONB Grassmannian Gr(1, 3), five CCSs with |K| = 1, and the correlation manifold Cor + (3). For each geometry, we randomly generate a batch of points, compute their batch mean, apply GyroBN, and then plot the normalized batch along with the resulting mean. The visualizations are constructed using the following embeddings: • Grassmannian. Since Gr(1, 3) is homeomorphic to the real projective space RP 2 , it is depicted as the unit hemisphere with antipodal points identified. • CCSs. The unit Poincaré ball P 3 and unit Beltrami–Klein ball K 3 are shown as the interior of the unit ball in R 3 . The unit hyperboloid H 2 is visualized as the upper sheet of a two-sheeted hyperboloid in R 3 . The unit sphere S 2 is embedded as a 2-sphere in R 3 , while the projected hypersphere D 3 coincides with R 3 itself. • Correlation manifold. Cor + (3) is embedded in R 3 as an open elliptope using its strictly lower triangular part. For better visualization, we fix the bias parameter to the gyro identity element and set the scaling parameter to 0.4 for the Grassmannian, correlation manifold, and sphere; 0.7 for the Poincaré ball, Beltrami–Klein ball, and hyperboloid; and 1 for the projected 100 Chapter 3. Riemannian Batch Normalization Figure 3.8: Visualization of GyroBN across different geometries. Blue and green points represent input and normalized data, respectively. Red and cyan points denote the input and output batch means. Black points mark the manifold boundary, and the gray surface depicts the manifold. hypersphere. As shown in Fig. 3.8, GyroBN consistently normalizes data distributions across these geometries. Notably, although the Poincaré and Beltrami–Klein inputs are identical, their GyroBN behavior and resulting sample distributions differ due to their distinct Riemannian metrics, underscoring that GyroBN faithfully respects the underlying geometry. 3.3.5.2 Experiments on Grassmannian Neural Networks Data sets and preprocessing. In line with previous work [108, 159], we evaluate Gy- roBN on three skeleton-based action recognition tasks, including HDM05 [153], NTU60 [178], and NTU120 [138] data sets, focusing on mutual actions for NTU60 and NTU120. 101 3.3. Gyrogroup Batch Normalization Each sequence is represented as a Grassmannian matrix of size 93× 10, 150× 10, and 150× 10 for HDM05, NTU60, and NTU120, respectively. Comparative methods. Since no Grassmannian-specific BN methods exist, we adapt two previous approaches to the Grassmannian: ManifoldNorm [34, Algs. 1–2] and the RBN method of Lou et al. [142, Alg. 2], which we denote LRBN for clarity. Although neither was originally designed for the Grassmannian, they can be adapted by employing Riemannian operators such as geodesics, exponential/logarithmic maps, and parallel transport. The key difference is that GyroBN can normalize data distributions across different geometries, whereas the other two methods cannot. Backbone networks. We adopt the recently proposed GyroGr architecture [159] as the backbone, which is briefly reviewed in Sec. A.2.2. GyroGr replaces the non-intrinsic FRMap + ReOrth block in GrNet [108] with Grassmannian left gyrotranslation, thereby improving numerical stability and performance. It consists of three basic components: left gyrotranslation, pooling [108], and the Projection Map (ProjMap) [108], where Pro- jMap maps Grassmannian points to symmetric matrices for classification. We consider both the 1-block and L-block variants. The 1-block version is structured as: gyrotrans- lation → pooling → ProjMap → classification, where the classification is implemented as an FC layer with softmax. The L-block version stacks L blocks of gyrotranslation and pooling, followed by a final ProjMap and classification layer. Since each pooling step approximately halves the dimensionality, we omit the pooling operation in the last block when L > 1. Following Huang et al. [108], the number of channels is fixed to 8. Implementation details. Following Nguyen and Yang [159], we use the Cayley map to approximate the matrix exponential of skew-symmetric matrices and apply the trivialization strategy reviewed in Sec. 2.7 to parameterize the Grassmannian variables in both the gyrotranslation and GyroBN layers. Specifically, each trainable Grassman- nian parameter is represented by Euclidean coordinates through the exponential map at the identity. This allows direct use of PyTorch optimizers [169] and avoids direct Riemannian updates of these parameters. In contrast, we find that the Grassman- nian LRBN benefits from Riemannian optimization. Thus, we employ Geoopt [125] to optimize its Grassmannian bias parameter. Similarly, we use Geoopt to update the orthogonal bias parameter in ManifoldNorm. For all models, the BN layer is inserted after the first pooling layer with a momentum of 0.1. Training uses SGD with a learn- ing rate of 5e −2 , batch size 30, and 400, 200, and 200 epochs for HDM05, NTU60, and NTU120, respectively. All models are optimized with a standard cross-entropy loss. Following previous normalization methods on matrix manifolds [123, 215] and the LieBN implementation in Sec. A.3.1, we adopt a single Fréchet mean iteration. 102 Chapter 3. Riemannian Batch Normalization Method HDM05 (47 × 10) NTU60 (75 × 10) NTU120 (75 × 10) AccFit Time #Params (M)AccFit Time #Params (M)AccFit Time #Params (M) GyroGr48.97±0.242.092.074470.13±0.1628.160.506253.76±0.1849.621.1812 GyroGr-ManifoldNorm49.67±0.7632.902.092168.56±0.43232.600.551251.41±0.38399.781.2262 GyroGr-LRBN48.64±0.7733.312.078167.77±0.52238.530.512250.56±0.22403.401.1872 GyroGr-GyroBN51.89±0.373.042.077372.60±0.0435.850.511455.47±0.1067.371.1864 Table 3.16: Comparison of GyroBN against other Grassmannian BN methods under the GyroGr backbone. Here, accuracy is reported as a percentage, fit time denotes the average training time per epoch (s/epoch), and #Params is reported in millions. Values in parentheses specify the dimension of the Grassmannian input to the BN layer. The largest number of parameters is marked in red. Main results. We compare GyroBN with ManifoldNorm and LRBN under the 1-block GyroGr backbone. The 5-fold results are presented in Tab. 3.16. We have the following four findings, which highlight the effectiveness of GyroBN in facilitating network training. • Improved accuracy. Across all three data sets, GyroBN consistently improves performance, enhancing the accuracy of the vanilla GyroGr by 2.92, 2.47, and 1.71 percentage points on HDM05, NTU60, and NTU120, respectively. In con- trast, both ManifoldNorm and LRBN often degrade performance, particularly on NTU60 and NTU120. This advantage comes from the theoretical guarantee of GyroBN for normalizing sample statistics, which is absent in the other two methods (see Tab. 3.13). • Enhanced efficiency. GyroBN is substantially more efficient than Manifold- Norm and LRBN. The efficiency gain is mainly attributed to: (i) replacing com- putationally expensive Riemannian operators (e.g., parallel transport, exponen- tial/logarithmic maps) with simpler gyro operations, (i) reducing matrix multi- plications from n× p to (n− p)× p or p× p (see Thm. 86), and (i) applying trivialization to avoid costly Riemannian optimization. • Improved parameter economy. GyroBN also requires fewer parameters than LRBN and ManifoldNorm. The key difference lies in the bias parameter. Due to the trivialization, GyroBN only needs an (n− p)× p Euclidean matrix, whereas LRBN requires an n× p Grassmannian matrix and ManifoldNorm requires an n× n orthogonal matrix. • Stronger generalization. As shown in Fig. 3.9, we observe that GyroBN can narrow the gap between training and testing accuracy, indicating a stronger gen- eralization ability. 103 3.3. Gyrogroup Batch Normalization Figure 3.9: Training and testing curves of 1-block GyroGr on two NTU data sets. 3.3.5.3 Experiments on Networks over CCSs Data sets. Following Lou et al. [142], we focus on the link prediction task on four graph data sets: Cora [177], Disease [4], Airport [229], and PubMed [155]. Comparative methods. We compare GyroBN with LRBN [142, Alg. 2]. As the original LRBN is only implemented in the Poincaré ball, we extend it to the other four CCSs. In particular, the core difference between LRBN and GyroBN lies in normaliza- tion: GyroBN can normalize sample statistics across different geometries, while LRBN lacks this guarantee. Backbone networks. We use a Hyperbolic Neural Network (HNN) [76] for the Poincaré ball and Klein HNN (KNN) [147] for the Beltrami–Klein. For the sphere and projected hypersphere, we mimic the transformation and activation in HNN [76, Sec. 3.2] to build the corresponding layers. The layers on the four spaces above can be expressed as Transformation: x k = Exp e M k Log e (x k−1 ) ⊕ b k , with b k ∈N, and M k ∈ R m×n , Activation: x k = Exp e φ Log e (x k−1 ) , with φ as an activation, where e is the origin, ⊕ is the gyroaddition, and N ∈ P n K , K n K , S n K , D n K . For the 104 Chapter 3. Riemannian Batch Normalization hyperboloid, we use the Lorentz fully-connected layer [45, Eq. 3] and Lorentz activation layer [15, Eq. 13], which are briefly reviewed in Sec. A.2.4. The above backbone network is referred to as HNN, KNN, SNN, PHNN, and LNN, respectively. We collectively call them Constant Curvature Neural Networks (CCNNs). Implementation details on CCNNs. We follow the official implementations of HGCN 10 [38], LRBN 11 [142], and HCNN 12 [15] to conduct experiments, where we adopt the same training settings as Lou et al. [142, Sec. H.1]. Specifically, the baseline encoder is a CCNN with two transformation layers: the first maps the input feature dimension to 128, and the second maps 128 to 128. After each transformation layer, we use a ReLU activation, except for the Cora data set where activation is omitted. A BN layer, GyroBN or LRBN, is inserted after each transformation layer. The curvature is set as |K| = 1. For the manifold-valued bias parameter in GyroBN and LRBN, we apply the exponential map Exp e (v) to trivialize it via a Euclidean parameter v. The optimization is performed with Adam [121], using a learning rate of 1e −2 and a weight decay of 1e −3 , except for the Cora data set, where weight decay is set to 0. The Fréchet mean iterations are performed until convergence. Main results. We compare GyroBN with LRBN across five CCSs under the CCNN backbone. Tab. 3.17 reports the 5-fold average testing AUC on four data sets. We highlight the following findings. • Improved performance. GyroBN consistently improves performance over the vanilla CCNNs across all data sets and geometries, whereas LRBN degrades per- formance in several cases (highlighted in red). The gains are especially pronounced on the sphere (SNN), where GyroBN achieves improvements of 17.65 (Disease), 7.49 (Airport), and 13.37 (PubMed) percentage points. This contrast underscores the advantage of GyroBN’s theoretical guarantee of normalizing sample statistics. • Efficiency. As shown in the “Fit Time” rows of Tab. 3.17, GyroBN is more efficient than LRBN on the Poincaré ball, Beltrami–Klein and projected hyper- sphere, due to the simplicity of gyro operations. On the hyperboloid and sphere, GyroBN and LRBN exhibit comparable efficiency. • Parameter equivalence. GyroBN and LRBN require the same number of pa- rameters, only marginally more than those of the vanilla backbone. Thus, the 10 https://github.com/HazyResearch/hgcn 11 https://github.com/CUAI/Differentiable-Frechet-Mean 12 https://github.com/kschwethelm/HyperbolicCV 105 3.3. Gyrogroup Batch Normalization SpacePoincaré Ball P n K Hyperboloid H n K Beltrami–Klein K n K MethodHNNHNN-LRBNHNN-GyroBNLNNLNN-LRBNLNN-GyroBNKNNKNN-LRBNKNN-GyroBN Disease ROC79.21 ± 2.1476.58 ± 2.1581.18 ± 0.9387.71 ± 1.4285.31 ± 0.9588.87 ± 0.3381.31 ± 1.3778.98 ± 1.5181.56 ± 0.70 Fit Time0.0270.0880.0840.0120.0570.0570.0200.0890.077 #Params (M)0.01800.01830.01830.01840.01870.01870.01800.01830.0183 Airport ROC94.63 ± 0.1994.17 ± 0.4095.40 ± 0.1793.86 ± 0.2193.05 ± 1.0095.06 ± 0.1295.00 ± 0.0594.47 ± 0.3396.14 ± 0.05 Fit Time0.0540.1220.1190.0470.0890.0920.0560.1360.126 #Params (M)0.01820.01840.01840.01860.01880.01880.01820.01840.0184 PubMed ROC95.02± 0.4293.40 ± 0.2095.83 ± 0.1195.36 ± 0.1095.88 ± 0.0995.89 ± 0.1195.87 ± 0.1189.83 ± 0.2796.23 ± 0.12 Fit Time0.1250.3420.3350.1110.2420.2490.1250.3530.337 #Params (M)0.08060.08090.08090.08150.08180.08180.08060.08090.0809 Cora ROC89.96 ± 0.4493.47 ± 0.4994.32 ± 0.2291.84 ± 1.0192.62 ± 0.1793.66 ± 0.3090.03 ± 0.3293.38 ± 0.1293.48 ± 0.25 Fit Time0.0320.0910.0760.1140.2440.2420.0380.0910.076 #Params (M)0.20010.20030.20030.20190.20210.20210.20010.20030.2003 (a) Results on three hyperbolic spaces. SpaceProjected Hypersphere D n K Sphere S n K MethodPHNNPHNN-LRBNPHNN-GyroBNSNNSNN-LRBNSNN-GyroBN Disease ROC69.70 ± 2.0160.25 ± 1.2572.26 ± 0.6154.19 ± 2.2153.38 ± 4.0771.84 ± 0.89 Fit Time0.0220.0660.0630.0290.0420.043 #Params (M)0.01800.01830.01830.01810.01830.0183 Airport ROC89.60 ± 0.9987.06 ± 0.4690.44 ± 0.9383.63 ± 0.7786.14 ± 0.7991.12 ± 1.57 Fit Time0.0560.1020.0950.0530.0700.071 #Params (M)0.01820.01840.01840.01820.01840.0184 PubMed ROC89.86 ± 0.3990.06 ± 0.2392.06 ± 0.6479.94 ± 1.7690.10 ± 0.2193.31 ± 0.08 Fit Time0.1210.1750.1740.1200.1400.156 #Params (M)0.08060.08090.08090.08060.08090.0809 Cora ROC92.88 ± 0.2692.03 ± 0.4293.26 ± 0.4292.10 ± 0.4082.01 ± 0.7193.16 ± 0.32 Fit Time0.0260.0670.0630.0250.0440.045 #Params (M)0.20010.20030.20030.20010.20030.2003 (b) Results on two spherical spaces. Table 3.17: Comparison of GyroBN against LRBN across five CCSs. ROC is the testing AUC reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. When LRBN degenerates the backbone network, the results are highlighted with red. performance gains of GyroBN cannot be attributed to parameter size, but rather to its principled normalization mechanism. 3.3.5.4 Experiments on Correlation Neural Networks Data sets and preprocessing. Following the CorNet protocol in Sec. 5.4, we use the Radar, HDM05, and FPHA data sets and model each input sequence as multichannel correlation matrices. Comparative methods. Similar to Sec. 3.3.5.3, we compare GyroBN against LRBN. Backbone networks. We adopt the PHCM-based CorNet detailed in Sec. 5.4 as the backbone network. CorNet identifies each correlation matrix C ∈ Cor + (n) with a collection of Poincaré vectors and applies Poincaré layers. Specifically, C is mapped to the poly-Poincaré space P n−1 via Eq. (3.77). The resulting multi-channel Poincaré vectors are then merged into a single Poincaré representation through β-concatenation [181, Sec. 3.3]. Then, a Poincaré FC layer followed by a Poincaré MLR layer constructs the network. Implementation details. Since CorNet operates in the Poincaré geometry, both 106 Chapter 3. Riemannian Batch Normalization Methods RadarHDM05FPHA AccFit Time #Params (M)AccFit Time #Params (M)AccFit Time #Params (M) CorNet96.56 ± 0.862.120.043882.26 ± 0.920.740.409190.03 ± 0.630.701.1210 CorNet-LRBN92.85 ± 2.462.330.0439N/A1.640.409481.53 ± 0.721.121.1213 CorNet-GyroBN97.67 ± 0.362.190.043980.82 ± 0.861.330.409492.88 ± 0.201.071.1213 Table 3.18: Comparison of CorNet with or without RBN layers. Accuracy is reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. GyroBN and LRBN are instantiated in the Poincaré model and applied after the Poincaré FC layer. To stabilize training, we scale the learning rate of the scaling parameter s by factors of 0.1 and 0.5 for Radar and HDM05, respectively. On FPHA, we further regularize the scaling by clamping: min s √ v 2 +ε , 4 . For better efficiency, the number of Fréchet mean iterations is set to 2. Main results. We summarize the comparison of CorNet with or without normal- ization layers in Tab. 3.18. Overall, GyroBN demonstrates clear benefits on Radar and FPHA with negligible parameter cost and modest efficiency trade-offs. On HDM05, however, neither GyroBN nor LRBN improves the baseline, with LRBN even diverging and failing to converge. 3.4 Conclusion This chapter developed a unified approach to RBN across broad families of manifolds. We first presented LieBN, which operates under left-, right-, and bi-invariant metrics. LieBN uses the group translations for centering and biasing and performs scaling in the tangent space at the identity element. These operations provide theoretical con- trol over Riemannian sample and population statistics. Their instantiations on SPD, rotation, and full-rank correlation manifolds further demonstrate how a choice of Lie group structure and invariant metric yields a concrete normalization layer, including the power-deformed SPD structures and CRIM developed in this chapter. The chapter then introduced pseudo-reductive gyrogroups as a relaxation of classi- cal gyrogroups, providing a broader algebraic foundation for Riemannian normalization. Building on this structure, we developed GyroBN and showed that gyroisometric gy- rations enable theoretical control over sample statistics. Since every Lie group is a gyrogroup with identity gyrations, GyroBN recovers LieBN as a special case. GyroBN replaces group subtraction, addition, and tangent-space scaling with gyrosubtraction, gyroaddition, and scalar gyromultiplication, thereby extending the normalization prin- ciple to broader geometries. Finally, we instantiated the framework on the Grassman- nian, five CCS models, and the full-rank correlation manifold. Experiments across all 107 3.4. Conclusion considered geometries support the effectiveness of this extension. Together, LieBN and GyroBN extend normalization with theoretical control of in- trinsic statistics from Lie groups to pseudo-reductive gyrogroups. Gyrogroups substan- tially broaden the scope beyond Lie groups but do not encompass all manifolds. A natural direction for future work is therefore to extend this normalization principle to manifolds that do not admit suitable gyro-structures. The next chapter considers intrinsic classification. 108 Chapter 4 Riemannian Multinomial Logistic Regression 4.1 Introduction Although Riemannian networks demonstrated success in many applications, most ap- proaches still rely on Euclidean spaces for classification, such as tangent spaces [106, 107, 31, 156, 207, 209, 157, 158, 123, 208, 47], ambient Euclidean spaces [205, 183, 184], or coordinate systems [36]. 1 However, these strategies distort the intrinsic geometry of the manifold, undermining the effectiveness of Riemannian networks. Researchers have recently started directly developing Riemannian Multinomial Logistic Regression (RMLR) on manifolds. Inspired by the idea of hyperplane margin [130], Ganea et al. [76] developed a hyperbolic MLR in the Poincaré ball for HNNs. Motivated by HNNs, Nguyen and Yang [159] developed three kinds of gyro SPD MLRs based on three distinct gyro-structures of the SPD manifold. Nguyen et al. [161] proposed gyro MLRs for the Symmetric Positive Semidefinite (SPSD) manifold based on the product of gyro spaces. However, these classifiers often rely on manifold-specific geometric structures, limiting their generalizability to other geometries. For instance, the hyperbolic MLR [76] relies on the generalized law of sines, while the gyro MLRs [159, 161] rely on gyro-structures. This chapter proceeds in two stages. We first study SPD manifolds endowed with pullback Euclidean metrics, whose flat geometry reduces the point-to-hyperplane infi- mum to a Euclidean problem and yields closed-form intrinsic MLRs. We then extend the classification principle to general Riemannian manifolds. Rather than evaluating 1 Notice that there are also works designing classifiers for grid-based manifold-valued data [37], but we focus on non-gridded data in line with many previous SPD networks. 109 4.2. Multinomial Logistic Regression on SPD Manifolds the potentially intractable point-to-hyperplane infimum on each manifold, the general RMLR adopts a Riemannian-trigonometric formulation that requires only a well-defined Riemannian logarithm. Since the SPD classifiers in Sec. 4.2 are special cases of the gen- eral RMLR framework in Sec. 4.3, and the two parts use the same experimental settings, all experiments 2 are presented together in Sec. 4.4. Throughout this chapter and its accompanying appendix material, all parameters are assumed to satisfy the standing admissibility conditions. 3 4.2 Multinomial Logistic Regression on SPD Mani- folds 4.2.1 Introduction SPD matrices are commonly encountered in a diverse range of scientific fields, such as medical imaging [36, 37], signal processing [8, 105, 32, 31], elasticity [152, 89], question answering [141, 157], graph classification [42], and computer vision [106, 94, 232, 34, 41, 230, 37, 48, 183, 156, 158, 185]. Despite their ubiquitous presence, traditional learning algorithms are ineffective in handling the non-Euclidean geometry of SPD matrices. To address this limitation, several Riemannian metrics [171, 9, 137] have been proposed. With these Riemannian metrics, various machine learning techniques can be generalized to SPD manifolds. Inspired by the great success of deep learning [102, 127, 97], several deep networks have been developed on SPD manifolds. Despite their promising performance, many approaches still rely on Euclidean spaces for classification, such as tangent spaces [106, 31, 156, 207, 157, 158, 123, 208, 47], ambient Euclidean spaces [205, 183, 184], and coordinate systems [36]. However, these strategies distort the intrinsic geometry of the SPD manifold, undermining the effectiveness of SPD neural networks. Notably, there are also some similarity-based classifiers originally designed for shallow learning methods [77, 94, 48]. Although these classifiers can be extended to deep SPD neural networks [209, 211], the calculation of pairwise distances might undermine training efficiency. Recently, motivated by HNNs [76], three kinds of SPD Multinomial Logistic Regression (MLR) classifiers based on the gyro-structures induced by LEM, LCM, and 2 The code is available at https://github.com/GitZH-Chen/RMLR. 3 All normal vectors are nonzero. Whenever a displayed formula contains 1/θ or 1/θ 2 , we assume θ ̸= 0, except when an explicit limit as θ → 0 is considered. The inner-product parameters satisfy (α,β)∈ST, as defined in Eq. (2.93). These conditions are not repeated below. 110 Chapter 4. Riemannian Multinomial Logistic Regression AIM were developed by Nguyen and Yang [159]. However, the proposed SPD MLRs rely on the gyro-structures, limiting their generality. Besides, Chakraborty et al. [37] also introduced an invariant layer for manifold-valued data mimicking the invariant FC layer in Convolutional Neural Networks (CNNs). However, it is designed for gridded manifold-valued data, which is not the primary data type encountered in many other SPD networks. Following the convention of most SPD networks, we only focus on non-gridded cases. In fact, SPD MLR can be directly derived under LEM and LCM without the as- sistance of gyro-structures. More generally, LEM and LCM are pullback Euclidean metrics, which are metrics pulled back from the Euclidean space. This section devel- ops a unified construction of SPD MLR across pullback Euclidean metrics. On the empirical side, we focus on the parameterized LEM and LCM defined in Sec. 3.2.5.1, which generalize the standard LEM and LCM by the pullback of matrix power. Their deformation behavior is established in Thm. 67. We showcase our SPD MLRs under these parameterized metrics. Besides, our framework encompasses the gyro SPD MLRs induced by the standard LEM and LCM in Nguyen and Yang [159]. More importantly, our framework also provides an intrinsic explanation for the commonly used LogEig classifier on SPD manifolds, which consists of successive matrix logarithm, FC, and softmax layers. The main contributions are summarized as follows: (1) We develop a unified construction of SPD MLR under pullback Euclidean metrics and instantiate it under two parameterized metric families. (2) Our framework offers an intrinsic explanation of the most popular LogEig classifier, which stacks matrix logarithm, the FC layer, and softmax. Outline. Sec. 4.2.2 develops SPD MLRs under pullback Euclidean metrics. The deformed-metric instantiation follows in Sec. 4.2.3. Sec. 4.2.4 reinterprets the existing LogEig classifier intrinsically. The corresponding experiments are presented together with the general RMLR experiments in Sec. 4.4. Proofs are deferred to Sec. B.4. 4.2.2 SPD Multinomial Logistic Regression This section first reformulates the Euclidean MLR. Then, we deal with SPD MLR under an arbitrary pullback Euclidean metric on SPD manifolds. We use the pullback metric in Thm. 33 with a Euclidean codomain. The induced abelian Lie-group structure, bi- invariant metric, distance, and Riemannian operators required below are established later in Thm. 149. 111 4.2. Multinomial Logistic Regression on SPD Manifolds 4.2.2.1 Reformulation of Euclidean MLR The Euclidean MLR was first reformulated by Lebanon and Lafferty [130] from the perspective of distances to margin hyperplanes. Hyperbolic MLR was designed based on this reformulation [76]. Nguyen and Yang [159] further proposed three gyro SPD MLRs based on the gyro-structures induced by AIM, LEM, and LCM. We now briefly review the reformulation of Euclidean MLR. Given C classes, MLR in R n computes the following softmax probabilities: ∀k ∈1,...,C, p(y = k | x)∝ exp (⟨a k ,x⟩− b k ),(4.1) where b k ∈ R and x,a k ∈ R n . As shown in Lebanon and Lafferty [130, Sec. 5] and Ganea et al. [76, Sec. 3.1], Eq. (4.1) can be reformulated as p(y = k | x)∝ exp (sign (⟨a k ,x− p k ⟩)∥a k ∥d (x,H a k ,p k )), (4.2) where ⟨a k ,p k ⟩ = b k , and H a k ,p k is referred to as a hyperplane, defined as H a k ,p k =x∈ R n |⟨a k ,x− p k ⟩ = 0.(4.3) As reviewed in Tab. 2.3, Log p x is the natural generalization of the directional vector −→ px = x−p starting at p and ending at x, while the Riemannian metric at p corresponds to the inner product. Therefore, the MLR in Eq. (4.2) and hyperplane in Eq. (4.3) can be readily generalized to the SPD manifold (S n ++ ,g). Definition 102 (SPD hyperplanes). Given P ∈ S n ++ and A ∈ T P S n ++ \0, we define the SPD hyperplane as ̃ H A,P = S ∈S n ++ | g P (Log P S,A) =⟨Log P S,A⟩ P = 0 ,(4.4) where P and A are referred to as shift and normal matrices, respectively. Definition 103 (SPD MLR). The SPD MLR is defined as p(y = k | S)∝ exp sign A k , Log P k (S) P k ∥A k ∥ P k d S, ̃ H A k ,P k ,(4.5) where P k ∈ S n ++ , A k ∈ T P k S n ++ \0, ⟨·,·⟩ P k = g P k , ∥·∥ P k is the norm on T P k S n ++ induced by g at P k , and ̃ H A k ,P k is a margin hyperplane inS n ++ as defined in Eq. (4.4). 112 Chapter 4. Riemannian Multinomial Logistic Regression d S, ̃ H A k ,P k denotes the margin distance between S and the SPD hyperplane ̃ H A k ,P k , which is formulated as d S, ̃ H A k ,P k =inf Q∈ ̃ H A k ,P k d(S,Q),(4.6) where d(S,Q) is the geodesic distance induced by g. In geometry, the hyperplane in Eq. (4.3) is actually a regular submanifold of the trivial manifold R n . As for our definition of SPD hyperplanes, we have a similar result. Proposition 104 (Submanifolds). [↓] The SPD hyperplane, as defined in Eq. (4.4), is a regular submanifold of the SPD manifold if Log P : S n ++ → T P S n ++ is a global diffeomorphism. Thm. 104 rationalizes Thm. 102, as both the SPD hyperplane and Euclidean hyper- plane are submanifolds. Nevertheless, we still follow the nomenclature of Ganea et al. [76], Lebanon and Lafferty [130] and call ̃ H A,P an SPD hyperplane. 4.2.2.2 SPD MLRs under Pullback Euclidean Metrics For our SPD MLR in Thm. 103, under most Riemannian metrics on SPD manifolds, all the operators involved in Eq. (4.5) have closed-form expressions except the margin distance in Eq. (4.6). Therefore, the only difficulty lies in the calculation of the margin distance. This subsection proposes a general expression for SPD MLRs under pullback Euclidean metrics, which are defined by a pullback from Euclidean metric, as detailed in Thm. 149. We choose the identity matrix I as a fixed anchor. We choose pullback Euclidean metrics as our starting metrics mainly because of their extensive inclusion and easy computation. Several Riemannian metrics, including LEM, LCM, and their variants [196, 194], are pullback Euclidean metrics. Besides, due to their fast and simple calculation, the margin distance under a pullback Euclidean metric has a closed-form expression, while obtaining distances to hyperplanes under other metrics, such as AIM, would be complicated. We start by calculating the margin distance in Eq. (4.6) under a given pullback Euclidean metric. Lemma 105. [↓] Given a pullback Euclidean metric g, the margin distance defined 113 4.2. Multinomial Logistic Regression on SPD Manifolds in Eq. (4.6) has a closed-form solution: d S, ̃ H A k ,P k = d φ(S),H φ ∗,P k (A k ),φ(P k ) (4.7) = |⟨φ(S)− φ(P k ),φ ∗,P k (A k )⟩| ∥A k ∥ P k ,(4.8) where |·| is the absolute value. Putting Eq. (4.8) into Eq. (4.5), we obtain our SPD MLR under a given pullback Euclidean metric: p(y = k | S)∝ exp A k , Log P k (S) P k (4.9) = exp (⟨φ(S)− φ(P k ),φ ∗,P k (A k )⟩),(4.10) where S,P k ∈ S n ++ and A k ∈ T P k S n ++ \0. When P k is fixed, A k ∈ T P k S n ++ indeed lies in a Euclidean space. However, P k would vary during training, making A k non- Euclidean. To remedy this issue, A k can be generated from a Euclidean parameter in a fixed tangent space through one of three mechanisms: Riemannian parallel transport [69], vector transport 4 [1], or the differential of a group translation, including Lie group translation [197, Sec. 20] and gyrogroup translation [199]. We focus on Riemannian parallel transport and the differential of a Lie group translation, for which we establish an equivalence under pullback Euclidean metrics. Under parallel transport, we write A k = PT Q→P k ( ̃ A k ) with ̃ A k ∈ T Q S n ++ as a Euclidean parameter. This is the solution also adopted by HNNs [76], where the tangent point is the zero vector. Since the Lie groups associated with pullback Euclidean metrics are abelian, we only consider the left translation. We have the following two lemmas to show the relation between parallel transport and the differential of left translation. Lemma 106. [↓] Given a pullback Euclidean metric, any parallel transport is equiv- alent to the differential map of a left translation and vice versa. Lemma 107. [↓] Given two fixed SPD matrices Q 1 ,Q 2 ∈S n ++ , we have the follow- 4 Vector transport is often used in optimization as a substitute for parallel transport because its expressions are typically simpler and cheaper to compute [27, Sec. 10.5]. 114 Chapter 4. Riemannian Multinomial Logistic Regression ing equivalence for parallel transports under a pullback Euclidean metric: ∀ ̃ A 1,k ∈ T Q 1 S n ++ , ∃! ̃ A 2,k ∈ T Q 2 S n ++ , s.t. PT Q 1 →P k ̃ A 1,k = PT Q 2 →P k ̃ A 2,k . (4.11) Thm. 106 indicates that under pullback Euclidean metrics, the above two solutions are equivalent, while Thm. 107 implies that anchor points can be arbitrarily chosen. Therefore, without loss of generality, we generate A k from the tangent space at the identity matrix I by parallel transport, i.e., A k = PT I→P k ( ̃ A k ) with ̃ A k ∈ T I S n ++ ∼ = S n . Together with Eq. (6.9), Eq. (4.10) can be further simplified. Theorem 108 (SPD MLR under a pullback Euclidean metric). [↓] Under any pullback Euclidean metric, SPD MLR and SPD hyperplane are p(y = k | S)∝ exp D φ(S)− φ(P k ),φ ∗,I ( ̃ A k ) E ,(4.12) ̃ H ̃ A k ,P k = n S ∈S n ++ | D φ(S)− φ(P k ),φ ∗,I ( ̃ A k ) E = 0 o ,(4.13) where ̃ A k ∈ T I S n ++ \0 ∼ = S n \0 is a symmetric matrix, and P k ∈S n ++ is an SPD matrix. 4.2.3 SPD MLRs under Deformed LEM and LCM In this section, we first review the deformed LEM and LCM, and then showcase our SPD MLR in Thm. 108 under these deformed metrics. As reviewed in Sec. 3.2.5.1, inspired by the deforming utility of the matrix power function [192, 194], we define (θ,α,β)-LEM and θ-LCM as the pullback metrics of (α,β)-LEM and LCM by the matrix power function (·) θ and scaled by 1 θ 2 for θ ̸= 0. As shown in Thm. 67, (θ,α,β)-LEM is equal to (α,β)-LEM and θ-LCM interpolates between the standard LCM for θ = 1 and an LEM-like metric as θ → 0. Besides, both (α,β)-LEM and θ-LCM are pullback Euclidean metrics, as summa- rized in Tab. 3.6. Therefore, the SPD MLRs under these two families of metrics can be directly obtained by Thm. 108. Corollary 109 (SPD MLRs under the deformed LEM and LCM). [↓] The SPD 115 4.2. Multinomial Logistic Regression on SPD Manifolds MLR under (α,β)-LEM is p(y = k | S)∝ exp D log(S)− log(P k ), ̃ A k E (α,β) ,(4.14) where ̃ A k ∈ T I S n ++ ∼ = S n and P k ∈S n ++ . The SPD MLR under θ-LCM is p(y = k | S)∝ exp 1 θ ⟨X,Y⟩ ,(4.15) with X and Y defined as X =⌊ ̃ K⌋−⌊ ̃ L k ⌋ + h Dlog D( ̃ K) − Dlog D( ̃ L k ) i ,(4.16) Y =⌊ ̃ A k ⌋ + 1 2 D( ̃ A k ),(4.17) where ̃ K = Chol(S θ ), ̃ L k = Chol(P θ k ), and D( ̃ A k ) denotes a diagonal matrix with the diagonal elements of ̃ A k . S 2 ++ can be visualized as an open cone in R 3 by the condition that P = x y y z ! ∈ S 2 is positive definite if and only if x,z > 0 ∧ xz > y 2 . Fig. 4.1 illustrates SPD hyperplanes induced by (α,β)-LEM and θ-LCM. Remark 110. This construction incorporates the results with respect to LEM and LCM presented by Nguyen and Yang [159]. For (α,β)-LEM, when (α,β) = (1, 0), (α,β)-LEM becomes the standard LEM. Our margin distance to the hyperplane in Thm. 105 becomes the pseudo-gyrodistance under LEM [159, Thm. 2.23]. For θ-LCM, when θ = 1, θ-LCM becomes the standard LCM. Thm. 105 then becomes the pseudo-gyrodistance induced by LCM [159, Thm. 2.24]. However, our frame- work does not require gyro-structures and directly obtains the margin distance and SPD MLR based on the Riemannian metric. 4.2.4 Rethinking the Existing LogEig Classifier Many existing SPD neural networks [106, 31, 160, 207, 156, 208, 47] rely on a Euclidean MLR in the codomain of matrix logarithm, i.e., a matrix logarithm followed by an FC layer and a softmax layer. For simplicity, we call this classifier LogEig MLR. The existing explanation of LogEig MLR is that it approximates manifolds by a tangent 116 Chapter 4. Riemannian Multinomial Logistic Regression Figure 4.1: Conceptual illustration of SPD hyperplanes induced by (α,β)-LEM and θ-LCM. In each subfigure, the black dots are SPSD matrices, denoting the boundary of S 2 ++ , while the blue, red, and yellow dots denote three SPD hyperplanes. space. However, our framework can offer a novel intrinsic explanation for this widely used MLR. When (α,β) = (1, 0) for (α,β)-LEM, the SPD MLR in Eq. (4.14) is very similar to the LogEig MLR. However, due to the nonlinearity of log(·) and the non-Euclideanness of the SPD parameter P k , SPD MLR cannot be hastily viewed as equivalent to LogEig MLR. Nevertheless, under special circumstances, Eq. (4.14) is indeed equivalent to a LogEig MLR. Proposition 111. [↓] The LEM-based SPD MLR is equivalent to a LogEig MLR with parameters in the FC layer optimized by Euclidean SGD when SPD mani- folds are endowed with the standard LEM, the SPD parameter P k in Eq. (4.14) is optimized by LEM-based RSGD, and the Euclidean parameter ̃ A k is optimized by Euclidean SGD. Thm. 111 implies that, when optimized by LEM-based RSGD, the LEM-based SPD MLR is equivalent to the Euclidean MLR in the codomain of matrix logarithm. Never- theless, a substantial body of prior work underscores the theoretical and empirical supe- riority of AIM-based optimization over its LEM-based counterpart [186, 92]. Therefore, we adopt the AIM-based optimizer to update the involved SPD parameters. 4.3 Extension to General Riemannian Manifolds 4.3.1 Introduction The SPD MLR framework developed in Sec. 4.2 constructs an intrinsic classifier from the geodesic distance between an SPD matrix and a margin hyperplane. Its central 117 4.3. Extension to General Riemannian Manifolds quantity is the point-to-hyperplane distance in Eq. (4.6), defined as the infimum of the geodesic distance over all points on the hyperplane. For the pullback Euclidean metrics considered in Sec. 4.2.2.2, this optimization admits a closed-form solution. On a general Riemannian manifold, however, directly evaluating this infimum can require solving a difficult, potentially non-convex optimization problem, which prevents an extension of the preceding framework. We circumvent this obstacle by reinterpreting the Euclidean margin distance from a trigonometric perspective rather than directly solving the infimum. This alternative characterization expresses the margin distance through the angle between geodesics and the geodesic distance from the input to the hyperplane anchor. Lifting this characteri- zation to Riemannian manifolds yields a closed-form RMLR. Accordingly, the resulting framework only requires an explicit expression of the Riemannian logarithm, which is the minimal geometric requirement for extending Euclidean MLR to manifolds. Since this requirement is satisfied by many manifolds commonly encountered in machine learning, the framework applies broadly across different geometries. This reliance on an operator shared by many manifolds, rather than on additional manifold-specific structure, makes RMLR a unified classification module across geometries. We instantiate the framework on SPD manifolds and rotation matrices. On the SPD manifold, we systematically develop SPD MLRs under five families of power-deformed metrics and provide a complete theoretical discussion of their geometric properties. On the Lie group SO(n), we construct a Lie MLR under the widely used bi-invariant metric, providing the first extension of Euclidean MLR to Lie groups. Moreover, the framework incorporates several existing Riemannian MLRs, including gyro SPD MLRs [159], the SPD MLRs developed in Sec. 4.2, and gyro SPSD MLRs [161]. Our SPD MLRs are validated on four SPD backbone networks, including SPDNet [106] on the radar and human action recognition tasks and TSMNet [123] on the EEG classification tasks for the Riemannian feedforward network, RResNet [117] on the human action recognition task for the Riemannian residual network, and SPDGCN [231] on the node-classification task for the Riemannian graph neural network. Our Lie MLR is validated on the classic LieNet [107] backbone for the human action recognition task. Compared with previous non-intrinsic classifiers, our MLRs achieve consistent performance gains. In particular, our SPD MLRs outperform the previous classifiers by 14.23 percentage points on SPDNet and 13.72 percentage points on RResNet for human action recognition, and 4.46 percentage points on TSMNet for EEG inter-subject classification. Furthermore, our Lie MLR can improve both the training stability and performance. In summary, our main theoretical contributions are the 118 Chapter 4. Riemannian Multinomial Logistic Regression following: (1) We develop a unified RMLR framework for general Riemannian manifolds that requires only the Riemannian logarithm and incorporates several existing manifold- specific RMLRs as special cases. (2) We systematically propose five families of SPD MLRs based on different geometries of the SPD manifold. (3) We propose a novel Lie MLR for deep neural networks on SO(n). Outline. Sec. 4.3.2 revisits the existing Riemannian MLRs and proposes the general RMLR framework. Sec. 4.3.3 instantiates the framework on SPD manifolds under five families of power-deformed metrics. Sec. 4.3.4 presents the Lie MLR on SO(n). Sec. 4.4 reports experiments on Riemannian feedforward, residual, and graph networks, direct SPD classification, and LieNet. Proofs are deferred to Sec. B.5. 4.3.2 Riemannian Multinomial Logistic Regression Inspired by Lebanon and Lafferty [130], prior work extended the Euclidean MLR to hy- perbolic, SPD, and SPSD manifolds [76, 159, 161]. Sec. 4.2 further develops SPD MLRs under pullback Euclidean metrics. However, these classifiers rely on specific Rieman- nian properties, such as the generalized law of sines, gyro-structures, and flat metrics, which limit their generality. In this section, we first revisit several existing MLRs and then propose our Riemannian classifiers with minimal geometric requirements. 4.3.2.1 Revisiting Existing Multinomial Logistic Regressions The Euclidean MLR and its reformulation through margin distances to hyperplanes have already been reviewed in Sec. 4.2.2.1. By abstracting the expressions for the SPD MLR and SPD hyperplanes in Eqs. (4.4) and (4.5), we can readily extend this construction from the SPD manifold to a general Riemannian manifold M: p(y = k | S)∝ exp sign(⟨ ̃ A k , Log P k (S)⟩ P k )∥ ̃ A k ∥ P k ̃ d(S, ̃ H ̃ A k ,P k ) ,(4.18) ̃ H ̃ A k ,P k = n S ∈M| g P k (Log P k S, ̃ A k ) = 0 o ,(4.19) 119 4.3. Extension to General Riemannian Manifolds where P k ∈M, ̃ A k ∈ T P k M\0, g P k is the Riemannian metric at P k , and Log P k is the Riemannian logarithm at P k . The margin distance is defined as an infimum: ̃ d(S, ̃ H ̃ A k ,P k ) =inf Q∈ ̃ H ̃ A k ,P k d(S,Q).(4.20) The MLRs in Lebanon and Lafferty [130], Ganea et al. [76], Nguyen and Yang [159] and Sec. 4.2 can be viewed as different implementations of Eqs. (4.18) to (4.20). To calculate the MLR in Eq. (4.18), one has to compute the associated Riemannian metrics, logarithmic maps, and margin distance. The associated Riemannian metrics and logarithmic maps often have closed-form expressions on the frequently encountered manifolds in machine learning. However, the computation of the margin distance can be challenging. On the Poincaré ball of hyperbolic manifolds, the generalized law of sines simplifies the calculation of Eq. (4.20) [76]. However, the generalized law of sines is not universally guaranteed on other manifolds. For pullback Euclidean metrics on SPD manifolds, Thm. 105 provides a closed-form solution for the margin distance. For curved manifolds, solving Eq. (4.20) would become a non-convex optimization prob- lem. To address this challenge, Nguyen and Yang [159] defined gyro-structures on the SPD manifold and proposed a pseudo-gyrodistance to calculate the margin distance. Similarly, Nguyen et al. [161] proposed a pseudo-gyrodistance on the SPSD manifold based on the gyro product space. However, gyro-structures do not necessarily exist in general geometries. In summary, the aforementioned methods often rely on specific properties of their associated Riemannian metrics, which usually do not generalize to general geometries. 4.3.2.2 Riemannian Multinomial Logistic Regression Recalling Eqs. (4.18) and (4.19), the minimum requirement for extending Euclidean MLR to manifolds is the well-definedness of Log P k (S) for each k. In this subsection, we will develop Riemannian MLR, which depends solely on the Riemannian logarithm, without additional requirements, such as gyro-structures and the generalized law of sines. In the following, we always assume the well-definedness of the Riemannian log- arithm. We start by reformulating the Euclidean margin distance to the hyperplane from a trigonometric perspective and then present our Riemannian MLR. As we discussed before, obtaining the margin distance of Eq. (4.20) could be chal- lenging. Inspired by Nguyen and Yang [159], we resort to the perspective of trigonom- etry to reinterpret Euclidean margin distance. In Euclidean space, the margin distance 120 Chapter 4. Riemannian Multinomial Logistic Regression is equivalent to d (x,H a,p ) = sin (∠xpy ∗ )d(x,p), with y ∗ = argmax y∈H a,p \p cos∠xpy.(4.21) We extend Eq. (4.21) to manifolds by the Riemannian trigonometry and geodesic dis- tance, the counterparts of Euclidean trigonometry and distance. Definition 112 (Riemannian margin distance). Let ̃ H ̃ A,P be a Riemannian hyper- plane defined in Eq. (4.19), and S ∈M. The Riemannian margin distance from S to ̃ H ̃ A,P is defined as d S, ̃ H ̃ A,P = sin (∠SPY ∗ )d(S,P ),(4.22) where d(S,P ) is the geodesic distance, and Y ∗ = argmax Y∈ ̃ H ̃ A,P \P cos∠SPY.(4.23) The initial velocities of geodesics define cos∠SPY : cos∠SPY = ⟨Log P Y, Log P S⟩ P ∥ Log P Y∥ P ∥ Log P S∥ P ,(4.24) where ⟨·,·⟩ P is the Riemannian metric at P, and ∥·∥ P is the associated norm. The Riemannian margin distance in Thm. 112 has a closed-form expression. Theorem 113. [↓] The Riemannian margin distance defined in Thm. 112 is given by d(S, ̃ H ̃ A,P ) = |⟨Log P S, ̃ A⟩ P | ∥ ̃ A∥ P .(4.25) Putting Eq. (4.25) into Eq. (4.18), we can obtain a closed-form expression for Rie- mannian MLR. Theorem 114 (RMLR). [↓] Given a Riemannian manifold (M,g), the Rieman- nian MLR induced by g is p(y = k | S ∈M)∝ exp ⟨Log P k S, ̃ A k ⟩ P k ,(4.26) 121 4.3. Extension to General Riemannian Manifolds where P k ∈M, ̃ A k ∈ T P k M\0, and Log is the Riemannian logarithm. As discussed for SPD MLRs in Sec. 4.2.2.2, a tangent vector attached to a train- able base point can be generated from a Euclidean parameter in a fixed tangent space through Riemannian parallel transport, vector transport, or the differential of a group translation. Following the hyperbolic and gyro MLRs [76, 159], we focus on parallel transport and Lie group translation: ̃ A k = Γ Q→P k A k ,(4.27) ̃ A k = L P k ⊙Q −1 ⊙ ∗,Q A k ,(4.28) where Q ∈ M is a fixed point, A k ∈ T Q M\0, Γ is the parallel transport along the geodesic connecting Q and P k , and L P k ⊙Q −1 ⊙ ∗,Q denotes the differential map at Q of left translation L P k ⊙Q −1 ⊙ with P k ⊙ Q −1 ⊙ denoting the Lie group product and inverse. In this way, A k lies in a fixed tangent space and, therefore, can be optimized by a Euclidean optimizer. Remark 115. We make the following remarks regarding our Riemannian MLR. (a) The reformulations of Eq. (4.21) in gyro MLR [159, 161] and in our work are different. Gyro MLR adopts gyro trigonometry and gyro distance to reformulate Eq. (4.21), while our method directly uses Riemannian trigonometry and geodesic distance. (b) Compared with the hyperbolic, gyro SPD, and gyro SPSD MLRs [76, 159, 161] and the SPD MLRs in Sec. 4.2, our framework enjoys broader applicability, as our framework only requires the Riemannian logarithm. This property is commonly satisfied by most manifolds encountered in machine learning, such as the five met- rics on SPD manifolds mentioned in Sec. 2.9.1, the invariant metric on SO(n) [28], and the hyperbolic and spherical models in Sec. 2.9.5. Besides, several existing MLRs on different geometries are special cases of our Riemannian MLR, which are detailed in Tab. 4.1. (c) The well-definedness of the Riemannian logarithm is a much weaker require- ment compared to the existence of the gyro-structure. The gyro-structure not only requires the Riemannian logarithm but also implicitly requires geodesic complete- ness [159, Eqs. (1)–(2)]. For instance, on SPD manifolds, EM and BWM [196] are incomplete, undermining the well-definedness of gyro operations. 122 Chapter 4. Riemannian Multinomial Logistic Regression MLRGeometriesRequirements Incorporated by Our MLR Euclidean MLR (Eq. (4.1))Euclidean geometryN/A✓(Sec. A.4.1.1) Gyro SPD MLRs [159]AIM, LEM & LCM on S n ++ Gyro-structures✓(Thm. 118) Gyro SPSD MLRs [161]SPSD product gyro spaces Gyro-structures ✓(Sec. A.4.1.2) Flat SPD MLRs (Sec. 4.2)(α,β)-LEM & (θ)-LCM on S n ++ Pullback metrics from the Euclidean space ✓(Thm. 118) OursGeneral geometriesRiemannian logarithmN/A Table 4.1: Several MLRs on different geometries are special cases of our MLR. (",$,%)-AIM (",$,%)-EM (",$,%)-LEM Riemannian Metrics on SPD Manifolds "-LCM O'-InvariantMetrics ("-BWM Figure 4.2: Illustration of the deformation (left) and Venn diagram (right) of metrics on SPD manifolds, where IEM, SREM, and 1 4 PAM denote Inverse Euclidean Metric, Square Root Euclidean Metric, and Polar Affine Metric scaled by 1 /4, respectively. 4.3.3 SPD Multinomial Logistic Regressions This section showcases our RMLR framework on the SPD manifold. We first systemati- cally discuss the power-deformed geometries of SPD manifolds. Based on these metrics, we will develop five families of deformed SPD MLRs. 4.3.3.1 Deformed Geometries of SPD Manifolds As discussed in Sec. 2.9.1, there are five popular Riemannian metrics on SPD manifolds. These metrics can all be extended to power-deformed metrics. For a metric g on S n ++ , the power-deformed metric is defined as ̃g P (V,W ) = 1 θ 2 g P θ ((φ θ ) ∗,P (V ), (φ θ ) ∗,P (W )),∀P ∈S n ++ ,V,W ∈ T P S n ++ ,(4.29) where φ θ (P ) = P θ is the matrix power, and (φ θ ) ∗,P is the differential map. The de- formed metric ̃g can interpolate between a LEM-like metric (θ → 0) and g (θ = 1) [194]. Previous work power-deformed (α,β)-AIM and BWM into (θ,α,β)-AIM [192] and 2θ-BWM [194], respectively. Earlier, we introduced (θ,α,β)-LEM and θ-LCM in Sec. 3.2.5.1. They are the power deformations of (α,β)-LEM and LCM, respectively. By Thm. 67, (θ,α,β)-LEM is equal to (α,β)-LEM. We therefore use (α,β)-LEM for this family. Here, we further define the power deformation of (α,β)-EM through Eq. (4.29), 123 4.3. Extension to General Riemannian Manifolds NameProperties (θ,α,β)-LEMBi-invariance, O(n)-invariance, Geodesic Completeness (θ,α,β)-AIMLie Group Left-Invariance, O(n)-invariance, Geodesic Completeness (θ,α,β)-EMO(n)-Invariance θ-LCMLie Group Bi-Invariance, Geodesic Completeness 2θ-BWMO(n)-Invariance Table 4.2: Properties of deformed metrics on SPD manifolds (θ ̸= 0 and min(α,α + nβ) > 0). Figure 4.3: Conceptual illustration of SPD hyperplanes induced by five families of Riemannian metrics. The black dots denote the boundary of S 2 ++ . denoted by (θ,α,β)-EM. We have the following result for (θ,α,β)-EM. Proposition 116. [↓] (θ,α,β)-EM interpolates between (α,β)-LEM (θ → 0) and (α,β)-EM (θ = 1). So far, all five popular Riemannian metrics on SPD manifolds have been generalized to power-deformed families of metrics. We summarize their associated properties in Tab. 4.2 and present their theoretical relation in Fig. 4.2. We leave technical details in Sec. A.4.1.3. 4.3.3.2 Five Families of SPD Multinomial Logistic Regressions This subsection presents five families of specific SPD MLRs using our general frame- work in Thm. 114 and the metrics discussed in Sec. 4.3.3.1. We focus on generating ̃ A k by parallel transport from the identity matrix, except for 2θ-BWM. Since the par- allel transport under 2θ-BWM would undermine numerical stability (please refer to Sec. A.4.1.4 for more details), we resort to a newly developed Lie group operation [195]: S 1 ⊙ S 2 = L 1 S 2 L ⊤ 1 , ∀S 1 ,S 2 ∈S n ++ ,(4.30) where L 1 = Chol(S 1 ) is the Cholesky factor. 124 Chapter 4. Riemannian Multinomial Logistic Regression Theorem 117 (SPD MLRs). [↓] By abuse of notation, we omit the subscripts k of A k and P k . Given an SPD feature S, the SPD MLRs, p(y = k | S ∈ S n ++ ), are proportional to (α,β)-LEM : exp ⟨log(S)− log(P ),A⟩ (α,β) ,(4.31) (θ,α,β)-AIM : exp 1 θ D log P − θ 2 S θ P − θ 2 ,A E (α,β) ,(4.32) (θ,α,β)-EM : exp 1 θ ⟨S θ − P θ ,A⟩ (α,β) ,(4.33) θ-LCM : exp 1 θ * ⌊ ̃ K⌋−⌊ ̃ L⌋ + h Dlog(D( ̃ K))− Dlog(D( ̃ L)) i , ⌊A⌋ + 1 2 D(A) + , (4.34) 2θ-BWM : exp " 1 4θ * (P 2θ S 2θ ) 1 2 + (S 2θ P 2θ ) 1 2 − 2P 2θ , L P 2θ ̄ LA ̄ L ⊤ +# ,(4.35) where A∈ T I S n ++ \0 is a symmetric matrix, log(·) is the matrix logarithm, L P [V ] is the solution to the matrix linear system L P [V ]P + PL P [V ] = V , known as the Lyapunov operator, Dlog(·) is the diagonal element-wise logarithm,⌊·⌋ is the strictly lower part of a square matrix, and D(·) is a diagonal matrix with diagonal elements of a square matrix. Besides, log ∗,P is the differential map at P, ̃ K = Chol(S θ ), ̃ L = Chol(P θ ), and ̄ L = Chol(P 2θ ). The Lyapunov operator in Eq. (4.35) requires the eigendecomposition. However, the backpropagation of eigendecomposition involves 1 /(σ i −σ j ) [111], undermining numerical stability. Therefore, we propose a numerically stable backpropagation for the Lyapunov operator, detailed in Sec. A.4.1.4. As 2× 2 SPD matrices can be embedded into R 3 as an open cone [220], we illustrate SPD hyperplanes induced by five families of metrics in Fig. 4.3. Remark 118. Our SPD MLRs extend the gyro SPD MLRs of Nguyen and Yang [159] and the flat SPD MLRs in Sec. 4.2. The pseudo-gyrodistance to an SPD hyperplane in Nguyen and Yang [159, Thms. 2.23–2.25] is incorporated by our Thm. 113, while the flat SPD MLRs under (α,β)-LEM and θ-LCM in Thm. 109 are special cases of our Thm. 117. Furthermore, our approach extends the scope of prior work because neither the framework in Sec. 4.2 nor the gyro SPD MLRs of Nguyen and Yang [159] cover SPD MLRs based on (θ,α,β)-EM and 2θ-BWM. The gyro operations in 125 4.3. Extension to General Riemannian Manifolds Nguyen and Yang [159, Eq. (1)] implicitly require geodesic completeness, whereas (θ,α,β)-EM and 2θ-BWM are incomplete. As neither (θ,α,β)-EM nor 2θ-BWM belong to pullback Euclidean metrics, the framework in Thm. 108 cannot be applied to these metrics. To the best of our knowledge, our work is the first to apply PEM and BWM to establish Riemannian neural networks, opening up new possibilities for utilizing these metrics in machine learning applications. Besides, neither the gyro SPD MLRs of Nguyen and Yang [159] nor the flat SPD MLRs in Sec. 4.2 cover the deformed metrics for building SPD MLRs. 4.3.4 Lie Multinomial Logistic Regression This section introduces our Lie MLR on SO(n) based on the general RMLR framework in Thm. 114. The Riemannian metric on SO(n) is assumed to be the invariant metric in Tab. 2.12. We employ the vector transport on SO(n) given by Boumal and Absil [28, Tab. 1], which coincides with the differential of left translation in Eq. (4.28). Lemma 119. [↓] T Q→P (H) = (L PQ −1 ) ∗,Q (H) = PQ ⊤ H, ∀P,Q∈ SO(n), H ∈ T Q SO(n). (4.36) Similar to SPD MLRs, we set Q = I. The Lie MLR on SO(n) is presented in the following. Theorem 120. [↓] The Lie MLR on SO(n) is given by p(y = k | R∈ SO(n))∝ exp log P ⊤ k R ,A k ,(4.37) where P k ∈ SO(n) and A k ∈ so(n). We refer to the Riemannian hyperplanes (Eq. (4.19)) on SO(n) as Lie hyperplanes. As SO(3) is homeomorphic to 3-dimensional real projective space RP 3 [95], Fig. 4.4 illustrates Lie hyperplanes in the closed ball in R 3 of radius π. 126 Chapter 4. Riemannian Multinomial Logistic Regression Figure 4.4: Conceptual illustration of a Lie hyperplane. Each pair of antipodal black dots corresponds to a rotation matrix with an Euler angle of π, while the green dots denote a Lie hyperplane. 4.4 Experiments ArchitecturesLogEig MLR (θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWMθ-LCM (1,1,0)(1,1,0)(1,1, 1 /8)(1,1,0)(1,1,1)(0.5)(0.25)(1)(0.5) 2-Block92.88±1.0594.53±0.9594.24±0.5594.93±0.6093.55±1.2195.64±0.8392.22±0.8394.99±0.4793.49±1.2594.59±0.82 5-Block93.47±0.4594.32±0.9495.11±0.8295.01±0.8494.60±0.7095.87±0.5893.69±0.6694.84±0.6893.93±0.9895.16±0.67 Table 4.3: Comparison of SPDNet with LogEig against SPD MLRs on the Radar data set. The best results are bold. ArchitecturesLogEig MLR (θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWMθ-LCM (1,1,0)(1,1,0)(0.5,1.0, 1 /30)(1,1,0)(0.5)(1)(0.5) 1-Block57.42±1.3158.07±0.6466.32±0.6371.65±0.8856.97±0.6170.24±0.9263.84±1.3165.66±0.73 2-Block60.69±0.6660.72±0.6266.40±0.8770.56±0.3960.69±1.0270.46±0.7162.61±1.4665.79±0.63 3-Block60.76±0.8061.14±0.9466.70±1.2670.22±0.8160.28±0.9170.20±0.9162.33±2.1565.71±0.75 Table 4.4: Comparison of SPDNet with LogEig against SPD MLRs on the HDM05 data set. ClassifiersLogEig MLR (θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWMθ-LCM (1,1,0)(0.5,1,0.05)(1,1,0)(1,1,0)(0.5)(1)(1.5) Balanced Acc.53.83±9.7753.36±9.9255.27±8.6854.48±9.2153.51±10.0255.54±7.4555.71±8.5756.43±8.79 Table 4.5: Inter-session experiments of TSMNet with different MLRs on the Hinss2021 data set. 127 4.4. Experiments ClassifiersLogEig MLR (θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWMθ-LCM (1,1,0)(1.5,1,0)(1,1,0)(1.5,1, 1 /20)(1,1,0)(0.5)(0.75)(1)(0.5) Balanced Acc.49.68±7.8850.65±8.1351.15±7.8350.02±5.8151.38±5.7751.41±7.9850.26±7.2351.67±8.7352.93±7.7654.14±8.36 Table 4.6: Inter-subject experiments of TSMNet with different MLRs on the Hinss2021 data set. We first instantiate our SPD MLRs in four SPD neural networks: SPDNet [106] and TSMNet [123] for Riemannian feedforward networks, RResNet [117] for Riemannian residual networks, and SPDGCN [231] for Riemannian graph neural networks. Then, we proceed with experiments of our Lie MLR under the classic LieNet architecture [107]. The classifier in all the above networks is the LogEig MLR (matrix logarithm + FC + softmax), a Euclidean MLR on the tangent space at the identity matrix. We substitute the original non-intrinsic LogEig MLR in each baseline model with our RMLRs. Notably, the gyro SPD MLRs [159] are special cases of our SPD MLRs under the standard AIM, LEM, and LCM ((θ,α,β) = (1, 1, 0)), while the flat SPD MLRs in Sec. 4.2 are incorporated by our SPD MLRs under (α,β)-LEM and θ-LCM. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.3. 4.4.1 Experiments on the Proposed SPD MLRs In the following, we abbreviate SPD MLR-metric as metric. For instance, (θ,α,β)-AIM denotes the baseline endowed with the SPD MLR induced by (θ,α,β)-AIM, with (1, 1, 0) as the value of (θ,α,β). Experiments on the Riemannian feedforward network. We evaluate our SPD MLRs for Riemannian feedforward networks under the SPDNet and TSMNet backbones. Following Huang and Van Gool [106], Brooks et al. [31], on SPDNet, we use the Radar data set [31] for radar recognition and the HDM05 data set [153] for human action recognition. TSMNet [123] is one of the state-of-the-art methods for the EEG classification task. Following Kobler et al. [123], we use the Hinss2021 [101] data set. For each family of SPD MLRs, we report the SPD MLR induced by the standard metric (θ = 1,α = 1,β = 0) and the one induced by the deformed metric with the best (θ,α,β). Besides, if the standard SPD MLR is already saturated, we only report the results of the standard one. Under each metric, we highlight the results of our SPD MLR under the best hyperparameters in bold. (1) Radar. In line with Brooks et al. [31], we evaluate our classifiers under two network architectures: 2-Block and 5-Block configurations. The 10-fold results (mean±std) are presented in Tab. 4.3. Note that the SPD MLR induced by standard AIM is sat- 128 Chapter 4. Riemannian Multinomial Logistic Regression Data Sets LogEig MLR(θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWMθ-LCM HDM0558.17 ± 2.0760.23 ± 1.2671.89 ± 0.60 (↑ 13.72)59.44 ± 0.8769.85 ± 0.2365.76 ± 0.96 NTU6045.22 ± 1.2348.94 ± 0.6852.24 ± 1.2546.99 ± 0.4150.56 ± 0.5953.63 ± 0.95 (↑ 8.41) Table 4.7: Comparison of LogEig against SPD MLRs under the RResNet architecture. urated. Generally speaking, our SPD MLRs achieve superior performance against the vanilla LogEig MLR. Moreover, for most families of metrics, the associated SPD MLRs with proper (θ,α,β) outperform the standard SPD MLR, demonstrat- ing the effectiveness of our parameterization. Besides, among all SPD MLRs, the ones induced by (α,β)-LEM achieve the best performance. (2) HDM05. Following Huang and Van Gool [106], three architectures are adopted: 1-Block, 2-Block and 3-Block configurations. The 10-fold results (mean±std) are presented in Tab. 4.4. Note that the standard SPD MLRs under AIM, LEM, and BWM are already saturated on this data set. As on the Radar data set, similar observations can be made on this data set. Our SPD MLRs can bring consistent performance gains for SPDNet, and properly selected hyperparameters can bring further improvement. Particularly, among all the SPD MLRs, the ones based on the 2θ-BWM and (θ,α,β)-EM achieve the best performance. Compared to the vanilla LogEig MLR, the highest performance improvement is 14.23 percentage points, highlighting our approach’s effectiveness. Notably, since 2θ-BWM and (θ,α,β)-EM are geodesically incomplete and not pulled back from a Euclidean space, the SPD MLR under these two metrics cannot be derived by the framework of gyro or flat MLR. This contrast confirms the applicability of our theoretical framework to a broader range of geometries. (3) Hinss2021. The results (mean±std) of leave-5%-out cross-validation are reported in Tabs. 4.5 and 4.6. Once again, our intrinsic classifiers demonstrate improved performance compared to the LogEig MLR in both inter-session and inter-subject scenarios. Besides, the SPD MLRs based on θ-LCM achieve the best performance, outperforming the vanilla classifier by 2.60 percentage points for inter- session and by 4.46 percentage points for inter-subject. This finding high- lights the versatility of our framework. Experiments on the Riemannian residual network. Following Katsman et al. [117], we use the HDM05 and NTU60 [178] data sets on the RResNet backbone. For the hyperparameter (θ,α,β) in our SPD MLRs, we borrow the best ones from Tab. 4.4. Tab. 4.7 reports the 10-fold and 5-fold results on the HDM05 and NTU60 data sets, 129 4.4. Experiments Classifiers DiseaseCoraPubmed Mean±STDMaxMean±STD Max Mean±STD Max LogEig MLR90.55 ± 4.8396.8578.04 ± 1.2779.670.99 ± 5.1277.6 (θ,α,β)-AIM94.84 ± 2.2798.4379.79 ± 1.4481.677.83 ± 1.0880 (θ,α,β)-EM90.87 ± 5.1498.0379.05 ± 1.238178.16 ± 2.4179.5 (α,β)-LEM96.33 ± 2.1998.8279.89 ± 0.9981.878.16 ± 2.4179.5 2θ-BWM91.93 ± 3.6496.8573.46 ± 2.1877.773.22 ± 4.0678.1 θ-LCM93.01 ± 2.1498.4377.59 ± 1.2080.174.46 ± 5.8178.9 Table 4.8: Comparison of LogEig against SPD MLRs under the SPDGCN architecture. ClassifiersRadarHDM05 Hinss2021 Inter-sessionInter-subject LogEig MLR91.93 ± 1.3048.43 ± 1.2539.76 ± 7.6044.66 ± 7.17 (θ,α,β)-AIM95.21 ± 0.8149.17 ± 1.0841.14 ± 7.2645.89 ± 6.52 (θ,α,β)-EM92.25 ± 1.2061.60 ± 0.6945.78 ± 8.51 (↑ 6.02)45.84 ± 4.75 (α,β)-LEM95.09 ± 0.5749.05 ± 0.9140.88 ± 7.4646.02 ± 5.96 (↑ 1.36) 2θ-BWM94.89 ± 0.4166.77 ± 1.34 (↑ 18.34)44.84 ± 8.0045.21 ± 7.44 θ-LCM95.67 ± 0.61 (↑ 3.74)58.66 ± 0.5143.17 ± 6.2145.10 ± 6.20 Table 4.9: Comparison of LogEig against SPD MLRs for direct classification. respectively. The SPD MLRs still consistently outperform the vanilla LogEig MLR. Besides, similar to the SPD MLRs under the SPDNet backbone for action recognition (Tab. 4.4), the SPD MLR based on θ-LCM, 2θ-BWM, or (θ,α,β)-EM outperforms the vanilla LogEig MLR by a large margin. In particular, the highest performance improvements are 13.72 and 8.41 percentage points on these two data sets. Experiments on the Riemannian graph network. We use SPDGCN [231] as the backbone network for the Riemannian graph network. Following Zhao et al. [231], we use the Disease [4], Cora [177], and Pubmed [155] data sets for node classification. The 10-fold average and maximum results of the vanilla LogEig MLR against our SPD MLR with the best (θ,α,β) are reported in Tab. 4.8. Similar to the previous results, our SPD MLRs generally outperform the LogEig MLR. Besides, the SPD MLR based on (α,β)-LEM generally achieves the best performance for SPDGCN. Ablations of SPD MLRs on direct classification. For a more straightforward comparison, we compare LogEig against our SPD MLRs for direct classification. We adopt the Radar, HDM05, and Hinss2021 data sets. We follow the preprocessing of SPDNet and TSMNet to model features into the SPD manifold and directly use LogEig or our SPD MLRs for classification. The average results are presented in Tab. 4.9. The hyperparameters (θ,α,β) are borrowed from Tabs. 4.3 to 4.6. Our SPD MLRs consistently outperform the vanilla LogEig MLR. In particular, on the HDM05 data set, the highest performance improvement by our SPD MLRs is 18.34 percentage points, surpassing the non-intrinsic LogEig MLR by a large margin. 130 Chapter 4. Riemannian Multinomial Logistic Regression Classifiers G3DHDM05 Mean±STD Max Mean±STD Max LogEig MLR87.91±0.9089.7376.92±1.2779.11 Lie MLR89.13±1.792.1278.24±1.0380.25 Table 4.10: Results of LogEig MLR against Lie MLR under the LieNet architecture. 4.4.2 Experiments on the Proposed Lie MLR We apply our Lie MLR to the classic SO(n) network, i.e., LieNet [107], where fea- tures are on the Lie group of SO(3)×·× SO(3). More precisely, this feature space is a product manifold of SO(3) factors, and Thm. 120 extends naturally under the product metric, with each class logit obtained by summing the factorwise inner prod- ucts. Following LieNet [107], we use G3D [23] and HDM05 [153] data sets. We also extend the Riemannian optimization package Geoopt [125] to SO(3), allowing for the direct Riemannian optimization reviewed in Sec. 2.7. We find that RSGD performs best for LieNet. Tab. 4.10 presents the 10-fold average results of LieNet with or with- out Lie MLR. Note that on the HDM05 data set, LieNet might fail to converge, with the validation accuracy fluctuating between 70% and 75%. Therefore, we select the 10 best-performing folds out of 20 experimental folds. It can be observed that our Lie MLR can improve the performance of LieNet. Besides, our Lie MLR can also improve the training stability. On the HDM05 data set, LieNet fails to converge in 8 out of 20 folds. However, when endowed with our Lie MLR, LieNet+LieMLR only encounters convergence failures in 2 folds. 131 4.5. Conclusion 4.5 Conclusion This chapter developed a unified approach to intrinsic classification in two stages, progressing from a structured family of flat SPD geometries to general Riemannian manifolds. The first part considered SPD manifolds endowed with pullback Euclidean metrics. Their flat geometry reduces the infimum defining the geodesic distance from an SPD point to a margin hyperplane to a Euclidean point-to-hyperplane problem. This yields a closed-form margin distance and, consequently, a unified construction of SPD MLR. We instantiated this construction under deformed LEM and LCM and showed that, under the corresponding optimization scheme, its LEM instance recovers the widely used LogEig classifier, thereby providing an intrinsic interpretation of the existing pipeline. The second part addressed the central obstacle to extending this construction beyond flat geometries. On a general Riemannian manifold, evaluating the point-to-hyperplane distance through its infimum can require solving a difficult, potentially non-convex optimization problem and may not admit a closed-form solution. Instead of solving this minimization problem, we replaced the infimum-based margin formulation with a Riemannian-trigonometric one that combines the geodesic distance from an input to the hyperplane anchor with the angle between the corresponding geodesics. This reformulation yields a closed-form RMLR that requires only a well-defined Riemannian logarithm, extending the classification principle from flat SPD geometries to a broad range of Riemannian manifolds. We instantiated the general framework as five families of SPD MLRs under power- deformed metrics and as a Lie MLR on SO(n). Experiments across Riemannian feed- forward, residual, graph, and Lie-group networks demonstrated the broad applicability of the framework. 132 Chapter 5 Riemannian Neural Networks 5.1 Introduction The preceding chapters developed two fundamental network modules through unified geometric formulations that can be instantiated across different manifolds. Such uni- fied constructions make essential modules reusable across manifold families, but not every neural component admits a sufficiently tractable or effective formulation based only on broadly shared Riemannian properties. The general RMLR in Sec. 4.3.2.2, for example, achieves broad applicability by replacing the potentially intractable point-to- hyperplane infimum with a Riemannian-trigonometric formulation that requires only a well-defined Riemannian logarithm. In contrast, hyperbolic and flat correlation geome- tries permit exact evaluation of the corresponding point-to-hyperplane infima, while Busemann functions and horospheres provide an alternative hyperbolic decision prin- ciple. These examples illustrate how additional geometric or algebraic structure of a particular manifold can support more direct and better-tailored modules and architec- tures. This chapter therefore studies manifold-specific Riemannian network design through three complementary routes. In Sec. 5.2, we introduce the unconstrained Proper Veloc- ity (PV) model, establish its Riemannian toolkit, and construct MLR, fully connected, convolutional, activation, and normalization layers. In Sec. 5.3, we exploit Busemann functions and horospheres to derive intrinsic and batch-efficient BMLR and BFC layers for the Poincaré and Lorentz models. Finally, Sec. 5.4 exploits the specific geometries of full-rank correlation manifolds to construct MLR, fully connected, and convolutional layers together with Riemannian backpropagation. 133 5.2. Proper Velocity Neural Networks 5.2 Proper Velocity Neural Networks 5.2.1 Introduction Hyperbolic representations have recently delivered strong performance across different applications because the exponential volume growth of negatively curved manifolds enables low-distortion embeddings of tree-like and hierarchical structure [163]. These advantages have been validated in computer vision [78, 120, 72, 201, 79, 15, 99, 16, 189, 140, 217], graph learning [38, 13, 75, 188], multimodal learning [67, 166], rec- ommendation systems [223], astronomy [44], genome sequence learning [119], natural language processing [163, 76, 164, 90, 98, 222], and brain signal decoding [135]. Re- cently, the focus has shifted from hyperbolic embeddings to building HNNs that operate entirely within hyperbolic space. As reviewed in Sec. 2.9.5, hyperbolic geometry admits multiple models, so the choice of representation is central to the design of hyperbolic networks. Most recent works rely on the Poincaré ball and hyperboloid models, which provide convenient Riemannian or gyrovector structures, thereby facilitating neural network construction. However, both models are constrained spaces, which can lead to numerical instabilities. In particular, as embeddings in the Poincaré ball approach the boundary, numerical computations become unstable and might cause gradients to vanish [91]. On the other hand, the Proper Velocity (PV) model originates from Einstein’s special relativity, where proper velocity provides a natural parameterization for relativistic velocity addition [200, Ch. 10]. Algebraically, PV admits a gyrovector space [200, Ch. 6], analogous to the Möbius gyrovector space of the Poincaré ball. Unlike the constrained Poincaré ball and hyperboloid models, PV offers an unconstrained representation that alleviates numerical instabilities. These properties have made the PV model successful in relativistic physics and motivate its exploration as a stable alternative geometry for HNNs. However, its Riemannian operators, including exponential and logarithmic maps and parallel transport, remain largely unexplored, despite being fundamental for constructing neural networks. Inspired by the above discussions, we propose Proper Velocity Neural Networks (PVNNs). To this end, we first establish the complete Riemannian geometry of PV by deriving closed-form expressions for the exponential map, logarithmic map, geodesic distance, and parallel transport. Building on this foundation, we extend several fun- damental neural layers into PV space, including MLR classification, FC, convolutional, activation, and BN layers. Together, these layers form a complete PVNN framework 134 Chapter 5. Riemannian Neural Networks from which different network architectures can be constructed. We validate the frame- work through four sets of experiments, including numerical stability, image classifica- tion, graph learning, and genomic sequence learning, demonstrating both the stability of PV embeddings and effectiveness of PVNNs. To our knowledge, the PV model has remained largely unexplored in machine learning, and our work provides the first sys- tematic study of its use for representation learning. In summary, our contributions are threefold: (1) We establish the complete Riemannian geometric toolkit of the PV manifold, de- riving closed-form operators that enable its use as a new alternative to classical hyperbolic models. (2) We develop fundamental building blocks in PV space, including MLR, FC, convo- lutional, activation, and BN layers. (3) We validate the stability and effectiveness of PVNNs through experiments on four tasks: numerical stability, image classification, graph node classification, and ge- nomic sequence learning. 1 Outline. In Sec. 5.2.2, we introduce the PV model and its gyrovector operations. In Sec. 5.2.3, we develop the Riemannian geometry and closed-form operators of PV space. In Sec. 5.2.4, we construct the core layers of PVNNs. In Sec. 5.2.5, we connect PV constructions to hyperboloid neural layers, and in Sec. 5.2.6, we evaluate their numerical stability and effectiveness. Proofs are deferred to Sec. B.6. 5.2.2 Preliminaries PV Space [200]. As shown in Sec. 2.9.5, hyperbolic space is a space with constant negative curvature K < 0 and admits several models one can work with. The popular models include the Poincaré ball and the hyperboloid (also known as the Lorentz model). The PV model PV n K = R n is an alternative representation of hyperbolic geometry, which was initially named the Ungar gyrovector space and is used to describe algebraic structures of relativistic proper velocities [200]. Unlike the bounded Poincaré ball or the constrained hyperboloid, the PV model is an unconstrained space, offering better numerical stability. Its Riemannian metric is given by Sec. B.6.1: g x (u,v) =⟨u,v⟩ + Kβ 2 x ⟨x,u⟩⟨x,v⟩, ∀x∈ PV n K ,∀u,v ∈ T x PV n K .(5.1) 1 The code is available at https://github.com/NickyoyoSu/PVNN. 135 5.2. Proper Velocity Neural Networks Here, β x = 1 √ 1−K∥x∥ 2 is the relativistic beta factor. In Ungar’s notation, the curvature is parametrized by a positive constant s with s 2 =−1/K, where s plays the role of the vacuum speed of light in special relativity [200, Sec. 3.8]. PV Gyrovector [200]. From an algebraic point of view, the PV space forms a gyrovector space [200, Def. 6.2], which extends the Euclidean vector space to manifolds. Given x,y,z ∈ PV n K and t ∈ R, PV gyroaddition ⊕ U and scalar gyromultiplication ⊗ U [200, Chs. 3.11 and 6.20] are defined as 2 x⊕ U y = x + y + 1− β y β y − K β x 1 + β x ⟨x,y⟩ x,(5.2) t⊗ U y = sinh t sinh −1 √ −K∥y∥ y √ −K∥y∥ ,(t⊗ U 0 = 0).(5.3) In particular, the PV inverse is ⊖ U x = −x, and the PV identity is the zero vector: 0⊕ U x = x⊕ U 0 = x. PV Gyration. As shown by Ungar [200, Eqs. 3.220 and 3.221], the PV gyration for any x,y,z ∈ PV n K is given by gyr[x,y]z = z + Ax + By D ,(5.4) where the coefficients are A = (1− β 2 y )K⟨x,z⟩− (1 + β x )(1 + β y )β x β y K⟨y,z⟩ + 2β 2 x β 2 y K 2 ⟨x,y⟩⟨y,z⟩, (5.5) B = (1− β 2 x )β 2 y K⟨y,z⟩ + (1 + β x )(1 + β y )β x β y K⟨x,z⟩,(5.6) D = (1 + β x )(1 + β y ) (1− β x β y K⟨x,y⟩ + β x β y ).(5.7) Here, β x = 1 √ 1−K∥x∥ 2 is the relativistic beta factor. 5.2.3 Proper Velocity Geometry 5.2.3.1 From Gyro Isomorphism to Riemannian Isometry The Poincaré ball also admits a gyrovector space, named the Möbius gyrovector space, as reviewed in Sec. 2.9.5. Algebraically, the PV and Möbius gyrovector spaces are isomorphic. We further show that PV and the Poincaré ball are geometrically isometric. 2 The subscript U refers to the initial letter of Ungar. 136 Chapter 5. Riemannian Neural Networks The following bijections define the gyrovector space isomorphism [200, Tab. 6.1]: π PV n K →P n K : PV n K ∋ x7→ β x 1 + β x x∈ P n K , π P n K →PV n K : P n K ∋ y 7→ 2γ 2 y y ∈ PV n K , (5.8) where γ y = 1 √ 1+K∥y∥ 2 is the gamma factor. The isomorphism preserves the gyro operations: π PV n K →P n K (x⊕ U y) = π PV n K →P n K (x)⊕ M π PV n K →P n K (y), ∀x,y ∈ PV n K ,(5.9) π PV n K →P n K (r⊗ U x) = r⊙ M π PV n K →P n K (x), ∀x∈ PV n K ,∀r ∈ R,(5.10) where ⊙ M and ⊕ M are the Möbius gyro operations reviewed in Sec. 2.9.5. Lemma 121 (Differentials). [↓] The differentials of π PV n K →P n K and π P n K →PV n K are d x π PV n K →P n K (v) = K β 3 x (1 + β x ) 2 ⟨x,v⟩x + β x 1 + β x v, ∀x∈ PV n K ,∀v ∈ T x PV n K , d y π P n K →PV n K (w) =−4Kγ 4 y ⟨y,w⟩y + 2γ 2 y w, ∀y ∈ P n K ,∀w ∈ T y P n K . Let id be the identity map. The differentials at the origin 0 are d 0 π PV n K →P n K = 1 2 id, d 0 π P n K →PV n K = 2 id.(5.11) Based on Thm. 121, we can prove that the above isomorphisms are isometries. Theorem 122 (Isometries). [↓] The mappings in Eq. (5.8) are Riemannian isome- tries. 5.2.3.2 Proper Velocity Riemannian Operators The Poincaré ball admits the closed-form Riemannian operators reviewed in Sec. 2.9.5. By Thm. 122, we can readily obtain the counterparts on PV space via the properties of Riemannian isometries reviewed in Thm. 33. 137 5.2. Proper Velocity Neural Networks Theorem 123 (PV Riemannian operators). [↓] Let π = π PV n K →P n K . Given x,y ∈ PV n K and v ∈ T x PV n K , the Riemannian operators on the PV space are Exp x (v) = x⊕ U 1 √ −K sinh √ −K(1 + β x ) β x ∥d x π(v)∥ d x π(v) ∥d x π(v)∥ , (5.12) Log x (y) = σ(x,y)z + τ (x,y)⟨x,z⟩x,(5.13) PT x→y (v) = 1 + β x β x ̃v− K (1 + β x )β y (1 + β y )β x ⟨y, ̃v⟩y,(5.14) d(x,y) = 2 √ −K tanh −1 √ −K∥π(−x⊕ U y)∥ ,(5.15) with z = (−x)⊕ U y. For the parallel transport, ̃v = gyr M [ ̄y,− ̄x] (d x π(v)) with gyr M as the Möbius gyration in Sec. 2.9.5, ̄x = β x 1+β x x and ̄y = β y 1+β y y. Here, the scalar coefficients in the logarithm are σ(x,y) = 2 √ −K tanh −1 √ −K∥π(z)∥ ∥z∥ , τ (x,y) = 2β x 1 + β x √ −K tanh −1 √ −K∥π(z)∥ ∥z∥ . (5.16) At the identity 0, the above operators can be further simplified: Exp 0 (v) = 1 √ −K sinh √ −K∥v∥ v ∥v∥ ,(5.17) Log 0 (y) = 1 √ −K sinh −1 √ −K∥y∥ y ∥y∥ ,(5.18) PT 0→y (v) = v− K β y 1 + β y ⟨y,v⟩y,(5.19) PT x→0 (v) = v + K β 2 x 1 + β x ⟨x,v⟩x,(5.20) d(0,y) = 1 √ −K sinh −1 √ −K∥y∥ .(5.21) The expressions containing normalized vectors or ∥z∥ −1 are understood by continu- ous extension in the zero cases. Thus, Exp x (0) = x and Log x (x) = 0. In particular, Exp 0 (0) = Log 0 (0) = 0. This implies that PV gyro operations can be expressed via Riemannian operations. 138 Chapter 5. Riemannian Neural Networks Theorem 124 (Gyro by Riemannian). [↓] The PV gyro operations can be rewritten as x⊕ U y = Exp x (PT 0→x (Log 0 (y))), ∀x,y ∈ PV n K , t⊗ U x = Exp 0 (t Log 0 (x)), ∀x∈ PV n K ,∀t∈ R. (5.22) 5.2.4 Proper Velocity Neural Networks Building on the above gyrovector and Riemannian tools, we introduce fundamental building blocks for PV neural networks, including MLR, FC, convolutional, activation, and BN layers, thereby enabling the construction of concrete deep architectures in this space. 5.2.4.1 Proper Velocity Multinomial Logistic Regression Following the point-to-hyperplane formulation in Sec. 4.2.2.1, we define the PV margin hyperplane and solve the corresponding point-to-hyperplane infimum under the PV geometry. We define the PV hyperplane as H a,p = n x∈ PV n K | Log p (x),a p = 0 o , p∈ PV n K ,a∈ T p PV n K ,(5.23) where p ∈ PV n K and a ∈ T p PV n K are the hyperplane parameters. As the Poincaré hyperplane can be expressed by the Möbius gyro operations [76, Eq. (22)], the PV hyperplane can also be expressed by the PV gyro operations. In addition, building PV MLR requires the PV point-to-hyperplane distance. The following theorem provides these results. Theorem 125. [↓] Let π = π PV n K →P n K . Given x,p∈ PV n K and a∈ T p PV n K , we have H a,p = n x∈ PV n K | Log p (x),a p = 0 o =x∈ PV n K |⟨−p⊕ U x,d p π(a)⟩ = 0, d(y,H a,p ) = inf w∈H a,p d(y,w) = 1 √ −K sinh −1 √ −K|⟨−p⊕ U y,d p π(a)⟩| ∥d p π(a)∥ . By Thm. 125, we define the C-class PV MLR as p(y = k | x)∝ exp (v k (x)), v k (x) = sign (⟨−p k ⊕ U x,d p k π(a k )⟩)∥a k ∥ p k d (x,H a k ,p k ), (5.24) where p k ∈ PV n K and a k ∈ T p k PV n K are the PV MLR parameters for class k. However, 139 5.2. Proper Velocity Neural Networks the above expression has three drawbacks: (i) the parameter p k is over-parameterized, as it corresponds to the scalar bias parameter in the Euclidean MLR; (i) the gyroad- dition in ⟨−p k ⊕ U x,d p k π(a k )⟩ complicates the computation; and (i) the parameters (p k ,a k ) are constrained, making optimization costly. To address these drawbacks, we follow Shimizu et al. [181] and adopt the parameterization p k = Exp 0 (r k z k /∥z k ∥), a k = PT 0→p k (z k ) with z k ∈ T 0 PV n K ∼ = R n and r k ∈ R. This parameterization avoids Riemannian optimization in PV MLR and further simplifies the formulation. Theorem 126 (PV MLR). [↓] For x∈ PV n K , the score v k (x) in Eq. (5.24) for each class k is v k (x) = ∥z k ∥ √ −K sinh −1 cosh( √ −Kr k ) √ −K ∥z k ∥ ⟨x,z k ⟩− sinh( √ −Kr k ) p 1− K∥x∥ 2 , (5.25) where z k ∈ R n and r k ∈ R are parameters for class k. In particular, as K → 0 − we have v k (x)→⟨x,z k ⟩ +b k with b k =−r k ∥z k ∥, which recovers the Euclidean MLR reviewed in Sec. 4.2.2.1. The parameterization (z k ,r k ) is essential for efficiency. In the original form Eq. (5.24), computing v k (x) for a batch x ∈ R b×n and C classes requires explicit gyroaddition −p k ⊕ U x for each class, producing an intermediate tensor of size b× C× n that could cause out-of-memory errors in high dimensions. One could instead loop over classes, but this is computationally inefficient. In contrast, Eq. (5.25) depends on inner products ⟨x,z k ⟩, which can be implemented as a matrix multiplication. 5.2.4.2 Proper Velocity Fully Connected Layer The Euclidean Fully Connected (FC) layer is defined as y = Ax + b with A ∈ R m×n and b ∈ R m . It can be expressed element-wise as y k = ⟨a k ,x⟩− b k = ⟨a k ,x− p k ⟩ with a k ,p k ∈ R n and ⟨p k ,a k ⟩ = b k . As shown by Shimizu et al. [181, Sec. 3.2] and Chen et al. [55, Sec. 3.1], the LHS y k is the signed distance from y to the hyperplane passing through the origin and orthogonal to the k-th axis of the output space, which can be formulated as sign (⟨e k ,y− 0⟩) d(y,H e k ,0 ) =⟨a k ,x− p k ⟩, ∀1≤ k ≤ m,(5.26) where e k denotes the vector whose k-th element is 1 and all others are 0. For the PV model, the LHS of Eq. (5.26) can be formulated by the signed point-to- hyperplane distance, while the RHS can be formulated by the v k in PV MLR. Specifi- 140 Chapter 5. Riemannian Neural Networks cally, the PV FC layer F : PV n K → PV m K from the n-dimensional to the m-dimensional PV spaces for the input x ∈ PV n K returns the output y ∈ PV m K by solving the m equations: sign (⟨d 0 π(e k ),−0⊕ U y⟩) d(y,H e k ,0 ) = v k (x), ∀1≤ k ≤ m,(5.27) where H e k ,0 and v k (x) are given by Thm. 125 and Eq. (5.25), respectively. This defini- tion has an explicit solution. Theorem 127 (PV FC layer). [↓] The output y = F (x) ∈ PV m K has the closed form y k = 1 √ −K sinh( √ −Kv k (x)),1≤ k ≤ m,(5.28) where v k (x) is defined in Eq. (5.25) with z k ∈ R n and r k ∈ R as the FC parameters. In particular, as K → 0 − we have y k → ⟨x,z k ⟩ + b k with b k = −r k ∥z k ∥, which recovers the Euclidean FC layer. Generalization. We can jointly express the Euclidean FC layer and activation σ, which yields the RHS of Eq. (5.26) with σ (⟨a k ,x− p k ⟩). Accordingly, we extend the PV FC by applying the activation to v k (x) in Eq. (5.27). Then, Eq. (5.28) becomes y k = 1 √ −K sinh( √ −Kσ(v k (x))),1≤ k ≤ m.(5.29) 5.2.4.3 Proper Velocity Convolution and Activation Convolution. As shown by Shimizu et al. [181], Bdeir et al. [15], Chen et al. [55], Euclidean convolution consists of linear maps between kernel weights and concatenated values in each receptive field. To define convolution on PV space, it therefore suffices to define PV concatenation, since we already have the PV FC layer. Because PV space is unconstrained, we define PV concatenation to coincide with Euclidean concatena- tion. For simplicity, we consider the 1D case. For PV inputs x i ∈ PV n K k i=1 in a 1D receptive field (where k is the kernel size), the PV convolution output y ∈ PV m K for this receptive field is y = F (Concat (x 1 ,...,x k )), where Concat(·) is standard Euclidean concatenation and F is the PV FC layer. Activation. A natural choice is to apply a Euclidean activation σ in the tangent space at the origin via the mapping x7→ Exp 0 (σ (Log 0 (x))), which has been shown to be effective in Poincaré networks [76]. Alternatively, since PV space is unconstrained, we can apply the activation directly in PV space as x 7→ σ(x). This direct PV-space 141 5.2. Proper Velocity Neural Networks activation avoids exponential and logarithmic maps and is therefore more efficient. 5.2.4.4 Proper Velocity Normalization We instantiate the GyroBN framework in Sec. 3.3.3 on the PV space and show that PV GyroBN can normalize sample statistics. Given activations x i ∈ PV n K N i=1 , the core operations of PV GyroBN are ∀i≤ N, ̃x i ← Biasing z | B⊕ U Scaling z | s √ v 2 + ε ⊗ U Centering z | −M ⊕ U x i ,(5.30) where M and v 2 denote the Fréchet mean and variance, and B ∈ PV n K and s ∈ R are parameters. Owing to the isometry between the PV space and Poincaré ball, the PV Fréchet mean can be computed via the Poincaré ball: map the data to the Poincaré ball, compute the Poincaré mean [142, Alg. 1], and map the result back. The following theorem guarantees that PV GyroBN can normalize sample statistics. Theorem 128 (Homogeneity). [↓] For N samples x i N i=1 ⊂ PV n K , we have Homogeneity of mean: FM B⊕ U x i N i=1 = B⊕ U FM x i N i=1 , ∀B ∈ PV n K , Homogeneity of dispersion from 0: 1 N X N i=1 d 2 (t⊗ U x i ,0) = t 2 · 1 N X N i=1 d 2 (x i ,0). Thm. 128 directly explains the PV GyroBN in Eq. (5.30). After the centering, the batch mean is shifted to the identity 0. After the scaling, the variance becomes s 2 . After the biasing, the batch mean is translated to B. 5.2.5 Connections to the Hyperboloid This subsection discusses the connections between the PV model and the hyperboloid model. We first show the isometry between the two models. Then, we show that several current hyperboloid network layers can be rewritten as PV layers. 142 Chapter 5. Riemannian Neural Networks Proposition 129 (PV–hyperboloid isometries). [↓] The following maps are Rie- mannian isometries between the hyperboloid model H n K and the PV model PV n K : π H n K →PV n K :H n K ∋ " x t x s # 7→ x s ∈ PV n K ,(5.31) π PV n K →H n K :PV n K ∋ x7→ q ∥x∥ 2 − 1 K x ∈ H n K .(5.32) The PV–hyperboloid isometries in Thm. 129 imply that several standard layers in hyperboloid networks can be rewritten as PV layers composed with π H n K →PV n K and π PV n K →H n K . The Lorentz activation [15, Eq. (13)], Lorentz FC layer [45, Sec. 3.1] and Lorentz concatenation [15, Eq. (32)] are LAct " x t x s #! = q ∥σ(x s )∥ 2 − 1 K σ(x s ) ,(5.33) LFC " x t x s #! = q ∥Wx s + b∥ 2 − 1 K Wx s + b ,(5.34) HCat(x i N i=1 ) = q P N i=1 x 2 i,t + N−1 K x 1,s . . . x N,s ∈ H nN K ,(5.35) where x = [x t ,x ⊤ s ] ⊤ ∈ H n K and x i = [x i,t ,x ⊤ i,s ] ⊤ ∈ H n K for 1 ≤ i ≤ N. Then Thm. 129 implies that the above Lorentz layers can be rewritten in terms of PV layers as follows: LAct(x) = π PV n K →H n K (σ(π H n K →PV n K (x))),(5.36) LFC(x) = π PV n K →H n K (Wπ H n K →PV n K (x) + b),(5.37) HCat(x i N i=1 ) = π PV n K →H n K Concat(π H n K →PV n K (x 1 ),...,π H n K →PV n K (x N )) .(5.38) These identities show that many hyperboloid constructions effectively operate by map- ping to PV space, applying Euclidean building blocks there, and mapping back through π PV n K →H n K . This perspective naturally motivates designing networks directly in PV space, instead of repeatedly switching between equivalent models. Moreover, even if one 143 5.2. Proper Velocity Neural Networks follows the pattern H n K → PV n K → PV m K → H m K to construct layers, the intermediate map should be the PV layers, such as Thm. 127, rather than Euclidean layers, since PV is a non-linear Riemannian manifold. 5.2.6 Experiments We evaluate PV embeddings and PVNNs on four representative tasks: • Sec. 5.2.6.1 evaluates the numerical advantage of the PV model against Poincaré and hyperboloid. • Sec. 5.2.6.2 compares PV, Poincaré, and hyperboloid MLRs on image classifica- tion. • Sec. 5.2.6.3 evaluates our PV MLR, FC, and GyroBN layers on graph learning. • Sec. 5.2.6.4 compares fully PV convolutional networks with fully hyperboloid con- volutional networks on genomic sequence learning. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.4. 5.2.6.1 Numerical Stability r Failure rateViolation rate PV n K P n K H n K PV n K P n K H n K 1000N/A032.50 5000N/A092.36 10000N/A099.76 20004.23N/A0100 500064.42N/A0100 750079.63N/A0100 1000088.26N/A0100 1500096.43N/A0100 20000100N/A0100 1000 00100N/A0100 Table 5.1: Failure and violation rates (%) of r⊗ H x in FP32. We study three aspects: gyro operator, Riemannian operator, and gradient be- havior. All experiments use curvature K = −1, dimension n = 16, and batch size 4096. Gyro Operator. We use scalar gy- romultiplication r⊗ H x as a probe of nu- merical stability across hyperbolic mod- els. Given random batches x and radii r, we evaluate two metrics. The failure rate is the fraction of outputs that contain NaN/Inf. The violation rate is defined only for models with manifold constraints: Poincaré ball requires∥x∥ 2 <−1/K, and hyperboloid requires⟨x,x⟩ L =∥x s ∥ 2 −x 2 t = 1 K for x = [x t ,x ⊤ s ] ⊤ . The tolerance is set to 10 −8 . As PV is unconstrained, its violation rate is reported as N/A. As shown in Tab. 5.1, PV maintains zero failures up to r = 1000 in 144 Chapter 5. Riemannian Neural Networks FP32. The Poincaré ball has zero failure and violation rates, whereas the hyperboloid model starts to fail around r = 20 and quickly accumulates both NaN/Inf outputs and off-manifold points under large scalar multipliers, revealing pronounced numerical instability. ModelFP32FP64 P n K 2.1× 10 −4 4.3× 10 −11 H n K 1.0× 10 0 1.0× 10 0 PV n K 2.1× 10 −7 6.7× 10 −16 Table 5.2: ∥Log 0 (Exp 0 (v))− v∥. Riemannian Operator. We evaluate the ex- ponential and logarithmic maps by measuring the round-trip error ∥Log 0 (Exp 0 (v))− v∥ for tangent vectors v with large norm ∥v∥ = 10. Since this quantity is theoretically zero, any non-zero value reflects numerical instability. We sample a batch of such vectors and report the average error in Tab. 5.2. PV achieves stable behavior in both FP32 and FP64, whereas the Poincaré ball already exhibits noticeable errors in FP32 and the hyperboloid model remains unstable in both precisions. Model ∥∇ x f r (x)∥ RangeGradient behavior P n K [7.6× 10 −13 , 1.1× 10 −11 ]Vanishing gradients H n K [0, NaN]Exploding gradients PV n K [2.1× 10 −6 , 1.1× 10 −4 ]Stable gradients Table 5.3: Gradient magnitude ∥∇ x f r (x)∥ across varying radii. Gradient. To compare gra- dient behavior, we study the gra- dient of f r (x) = ∥r⊗ H x− x∥ with respect to x. Specifi- cally, we sample 24 logarithmi- cally spaced radii r ∈ [1, 1000] and, for each radius, measure the ∥∇ x f r (x)∥ on a random batch. The range of ∥∇ x f r (x)∥ is summarized in Tab. 5.3. The Poincaré ball exhibits severe gradient vanishing near the boundary. In contrast, the hyperboloid model yields gradi- ents that vary from 0 to NaN, reflecting gradient explosion. PV maintains gradients in a safer band. 5.2.6.2 Image Classification We compare our PV MLR against previous Poincaré MLRs [76, 181] and Lorentz MLR [15]. Following Bdeir et al. [15], we train a ResNet-18 backbone [97] on CIFAR-10 and CIFAR-100 [126], replacing the final Euclidean MLR with a hyperbolic MLR. The backbone output is lifted to the target geometry via the exponential map at the identity. Since PV space is unconstrained, we also consider a direct variant that skips Exp 0 and treats the backbone output as PV coordinates. We denote these two PV heads as PV MLR (with Exp 0 ) and PV MLR (without Exp 0 ). Tab. 5.4 reports the 5-fold results. 145 5.2. Proper Velocity Neural Networks Model Method CIFAR-10 (δ = 0.26) CIFAR-100 (δ = 0.23) P n K Poincaré MLR [76, Eq. (25)]95.09± 1.5176.78± 0.67 Unidirectional MLR [181, Eq. (6)]95.12± 0.2077.19± 0.10 H n K Lorentz MLR [15, Thm. 2]95.02± 0.1277.96± 0.09 PV n K PV MLR (with Exp 0 )95.27± 0.1278.19± 0.59 PV MLR (without Exp 0 )95.30± 0.1878.20± 0.37 Table 5.4: Top-1 image classification accuracy (%) of hyperbolic MLRs on ResNet-18. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic), which comes from Bdeir et al. [15, Tab. 1]. Model Method Disease (δ = 0) Airport (δ = 1) PubMed (δ = 3.5) Cora (δ = 11) K n K KNN [147]79.41± 0.5592.10± 0.9769.36± 0.7652.26± 1.99 P n K HNN [76]79.90± 0.0182.16± 2.9569.28± 0.8549.68± 1.25 HNN++ [181]80.57± 0.2388.40± 0.1773.68± 0.3952.06± 0.90 H n K LNN [15]79.90± 0.0175.20± 1.0868.82± 0.8853.34± 1.65 PV n K PVNN81.15± 0.2397.96± 0.4274.33± 0.2251.42± 1.33 Table 5.5: Accuracies of hyperbolic networks on graph learning. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic). PV MLR matches or outperforms prior hyperbolic baselines, with the largest gains on CIFAR-100 where the decision boundaries are more complex. Both PV variants, with and without Exp 0 , achieve similar accuracies. 5.2.6.3 Graph Learning Data and Setup. We study node classification on four standard graph data sets: Dis- ease [4], Airport [229], Cora [177], and PubMed [155]. All models share the same archi- tecture consisting of two FC layers with nonlinear activations followed by an MLR clas- sifier; they differ only in the underlying hyperbolic model. Baselines include KNN [147] for the Klein ball, HNN/HNN++ [38, 181] for the Poincaré ball, and LNN [15] for the hyperboloid model. Our PVNN is built from PV FC, activation, and MLR layers. Main Results. For a fair comparison, we use a tangent activation in each model and set σ = id for the PV FC layer in Eq. (5.29). Tab. 5.5 summarizes the 5-fold results. On the three more hyperbolic data sets (Disease, Airport, and PubMed), PVNN consistently achieves the best performance, with especially large gains on Airport where it improves over the strongest baseline by 5.86%. On the weakly hyperbolic Cora data set, PVNN remains comparable to Poincaré- and Klein-based networks, and worse than 146 Chapter 5. Riemannian Neural Networks MethodDiseaseAirportPubMedCora PVNN+TFC80.86± 0.3086.99± 0.6174.40± 0.43 53.58± 0.81 PVNN81.24± 0.3697.93± 0.2974.16± 0.3252.26± 1.32 PVNN+TBN80.67± 0.3898.71± 0.3673.52± 0.1245.36± 2.44 PVNN+GyroBN81.24± 0.1999.03± 0.1874.34± 0.3146.64± 5.45 Table 5.6: Results of Tangent FC (TFC) vs PV FC, and Tangent BN (TBN) vs GyroBN. Method DiseaseAirportPubMedCora AccFit TimeAccFit TimeAccFit TimeAccFit Time Tangent81.15± 0.2326.0898.56± 0.3655.4861.50± 5.753.1033.10± 1.587.12 Euclidean81.15± 0.2325.8098.75± 0.3155.1969.82± 3.582.9932.62± 0.657.29 Fréchet 1 iter81.05± 0.2329.7988.93± 1.1765.1962.52± 8.443.3842.84± 6.157.67 Fréchet 2 iters81.05± 0.2330.1294.11± 0.4667.3773.78± 0.203.4945.68± 4.368.21 Fréchet 5 iters81.24± 0.3630.9098.50± 0.1682.2873.92± 0.444.0249.50± 1.839.15 Fréchet 10 iters81.24± 0.1930.4999.03± 0.18105.7974.34± 0.313.9646.64± 5.459.77 Fréchet ∞80.86± 0.0031.2998.46± 0.15122.3771.16± 3.934.4647.32± 4.739.27 Table 5.7: Comparison of methods in calculating mean and variance in PV GyroBN. Time is measured in milliseconds per training epoch. the hyperboloid-based one. Overall, these results suggest that PV geometry is more effective on strongly hyperbolic graphs. Tangent vs. Riemannian. A natural construction of hyperbolic layers is to work in the tangent space. To validate the benefits of our Riemannian PV layers, we com- pare our PV FC with TFC of the form Exp 0 (A Log 0 (x) + b), and our GyroBN with TBN given by Exp 0 (BN(Log 0 (x))) [109]. We denote these variants by PVNN+TFC and PVNN+TBN, respectively. As shown in Tab. 5.6, PVNN consistently outperforms PVNN+TFC on the more hyperbolic Disease and Airport data sets, while performance on the other two data sets is comparable and TFC can be slightly better. For normal- ization, PVNN+GyroBN improves over PVNN+TBN on all data sets. Overall, these ablations validate the effectiveness of our Riemannian PV constructions, especially in strongly hyperbolic settings. Ablations on Batch Statistics. PV GyroBN in Eq. (5.30) uses Fréchet mean and variance, which require iterative solvers. We also consider two efficient variants. A tangent variant computes batch statistics in the tangent space at the identity via M = Exp 0 1 N N X i=1 Log 0 (x i ) ! , v 2 = 1 N N X i=1 ∥Log 0 (x i )− Log 0 (M )∥ 2 , (5.39) and a Euclidean variant computes standard Euclidean mean and variance directly in 147 5.2. Proper Velocity Neural Networks Exp 0 DiseaseAirportPubMedCora ✗81.05± 0.2397.71± 0.3474.22± 0.2651.92± 2.01 ✓81.24± 0.3697.93± 0.2974.16± 0.3252.26± 1.32 Table 5.8: Ablations on PVNN with or without exponential map for the input PV feature. MethodDiseaseAirportPubMedCora Tangent Act.81.24± 0.3697.93± 0.2974.16± 0.3252.26± 1.32 FC σ81.34± 0.4399.40± 0.1574.02± 0.1751.34± 0.46 FC σ + Tangent Act.80.96± 0.1999.15± 0.3873.96± 0.2251.30± 1.65 Euc. Act.81.34± 0.4398.87± 0.3574.56± 0.5938.10± 3.30 Table 5.9: Ablations on PV activations. the unconstrained PV space. Tab. 5.7 shows that Tangent and Euclidean are up to 2× faster while achieving similar accuracies on Disease and Airport. Although Fréchet- based GyroBN attains the best accuracies, it is more computationally expensive. Ablations on PV Embedding. In the main experiments, the input features are first lifted to PV via Exp 0 and then processed by PVNN. Since PV space is uncon- strained, we also consider a variant that feeds the Euclidean features directly as PV coordinates. Tab. 5.8 compares these two settings. The two variants perform similarly, while using Exp 0 provides small improvements on Disease, Airport, and Cora. This dif- fers from image classification in Tab. 5.4, where the variant without Exp 0 is marginally better. This slight discrepancy may stem from the different nature of the inputs. In vision, the ResNet encoder can adapt its learned representation to the chosen lifting, whereas in graphs the raw node features benefit slightly from the explicit exponential map. Ablations on Activation. We ablate three types of activations in PVNN: the internal nonlinearity σ in the PV FC layer (fixed to tanh), and explicit activations applied either directly in PV (Euc. Act.) or in the tangent space (Tangent Act.). Tab. 5.9 reports the results. First, when comparing these three choices individually, the differences are small on Disease and PubMed, while FC σ performs best on Airport, Tangent Act. performs best on Cora, and Euc. Act. degrades substantially on Cora. Second, when comparing the composite variant FC σ + Tangent Act. against Tangent Act., the composite does not yield consistent gains, suggesting redundancy. 148 Chapter 5. Riemannian Neural Networks TaskData SetEuclidean CNNHCNN-SPVCNN Retrotransposons LINEs70.63± 1.2476.12± 2.16 81.83± 0.27 SINEs85.15± 1.6485.45± 1.1693.78± 0.54 DNA transposonshAT-Ac87.45± 0.9089.61± 1.3492.08± 0.80 Pseudogenes processed60.66± 0.8268.30± 0.9371.27± 0.78 unprocessed51.94± 2.6956.10± 0.56 62.31± 0.78 Table 5.10: Comparison in MCC of hyperbolic and Euclidean convolutional networks, including PVCNN, on TEB data sets. 5.2.6.4 Genomic Sequence Learning Khan et al. [119] recently proposed Hyperbolic Convolutional Neural Networks (HC- NNs) on the hyperboloid for genomic sequence learning, demonstrating that HCNNs outperform Euclidean CNNs on this task. Following Khan et al. [119], we evaluate on the Transposable Elements Benchmark (TEB) data set for DNA transposable element prediction. To ensure a fair comparison, all models share the same backbone network architecture, which consists of two convolutional blocks followed by an FC layer and a final MLR classifier [119]. We use a single curvature shared for all layers. Tab. 5.10 re- ports 5-fold Matthews Correlation Coefficient (MCC). The PV Convolutional Network (PVCNN) achieves the best performance on all TEB tasks, with particularly strong gains on SINEs, where it improves over HCNN-S by about 9 MCC points. These results demonstrate the benefits of PV convolutional networks. 5.3 Hyperbolic Busemann Neural Networks 5.3.1 Introduction The preceding section introduced the unconstrained PV representation and its neural layers, based on geodesic hyperplanes. We next use another powerful geometric tool, the Busemann function, to construct hyperbolic neural networks. For simplicity, we focus on the widely used Poincaré ball and Lorentz models. To support deep learning fully in hyperbolic spaces, several key building blocks in neural networks have recently been generalized to Poincaré or Lorentz spaces, including attention [90, 45, 221], BN [142, 15, 52, 54], linear feed-forward layers [76, 181, 45], acti- vation [76, 15], residual blocks [201, 118, 99], MLR [181, 15, 162], and graph convolution [38, 139, 13, 60]. Among these components, MLR classification and FC layers play a fundamental role in final decision-making and feature transformation. 149 5.3. Hyperbolic Busemann Neural Networks Recently, hyperplanes and point-to-hyperplane distances, which have been explored in Chapter 4 and Sec. 5.2, have been adopted to construct hyperbolic MLR in both Poincaré [76, 181] and Lorentz [15] models. Ganea et al. [76, Sec. 3.1] introduced the first Poincaré MLR based on Poincaré hyperplanes, but the formulation suffers from over-parameterization and lacks batch efficiency. Shimizu et al. [181, Sec. 3.1] alleviated these issues through a re-parameterization strategy. Building on these ideas, Bdeir et al. [15, Sec. 4.3] proposed a Lorentz MLR. However, its hyperplane is defined by the ambient Minkowski space, which is model-specific and may distort Lorentzian geometry. For hyperbolic FC layers, three main formulations exist. Ganea et al. [76, Sec. 3.2] in- troduced Möbius matrix-vector multiplication through the tangent space on the Poincaré ball. Shimizu et al. [181, Sec. 3.2] further proposed the Poincaré FC layer, defined in- trinsically but restricted to the Poincaré model. On the Lorentz model, Chen et al. [45, Sec. 2.2] constructed a Lorentz FC layer by applying linear transformations in the ambi- ent Minkowski space followed by projection onto the Lorentz model. Thus, Möbius and Lorentz FC rely on flat-space (tangent or ambient) approximations that could distort intrinsic geometry, whereas Poincaré FC is intrinsic but model-specific. On the other hand, the Busemann function and its level sets, horospheres, have emerged as powerful intrinsic tools for hyperbolic learning. They enjoy convenient met- ric properties [29, Ch. I.8] and admit closed-form expressions on both the Poincaré and Lorentz models [24, Prop. 9]. These operators have supported several hyperbolic algorithms, including SVM [73], PCA [39], Sliced Wasserstein distances [24], and proto- type learning [81]. We also note that Nguyen et al. [162, Cor. 4.3] proposed a Poincaré MLR based on the Busemann function. However, its induced point-to-hyperplane dis- tance is pseudo, coincides with the true distance only in Euclidean geometry, remains over-parameterized, and is not batch efficient. These observations motivate intrinsic and batch-efficient formulations for MLR and FC layers that can operate on both the Poincaré ball and the Lorentz model. To address this need, we propose Busemann Multinomial Logistic Regression (BMLR) and Buse- mann Fully Connected (BFC) layers, two Busemann-based components for hyperbolic networks. Our contributions are summarized as follows: • We introduce BMLR, deriving intrinsic logits directly from Busemann functions with a point-to-horosphere distance interpretation. BMLR uses a compact per- class parameterization, eliminates manifold-valued parameters in prior MLRs, remains batch-efficient, and recovers Euclidean MLR as curvature tends to zero. 150 Chapter 5. Riemannian Neural Networks • We develop BFC layers by generalizing the FC and activation layers through the Busemann function, providing intrinsic constructions on both the Poincaré and Lorentz models. BFC preserves comparable complexity and parameter counts, and recovers Euclidean FC in the zero curvature limit. • We provide empirical validation across image classification, genome sequence learning, node classification, and link prediction. BMLR and BFC generally out- perform existing hyperbolic layers. BMLR shows particularly large gains as the number of classes increases, and the Lorentz BMLR is the fastest among all hy- perbolic MLRs. 3 Outline. In Sec. 5.3.2, we recall Busemann functions and horospheres in hyperbolic space. In Sec. 5.3.3, we introduce BMLR and its point-to-horosphere interpretation. In Sec. 5.3.4, we develop BFC layers, and in Sec. 5.3.5, we evaluate both components. Proofs are deferred to Sec. B.7. 5.3.2 Preliminaries The metric-geometric notions of geodesic rays, asymptotic rays, Busemann functions, horoballs, horospheres, and Hadamard spaces have been reviewed in Sec. 2.5. The Poincaré and Lorentz models and their Riemannian operators have been reviewed in Sec. 2.9.5. Their gyro operators are presented in Tab. 2.13 and Sec. 3.3.4.5, respectively. Their gyrovector spaces are denoted by P n K ,⊕ M ,⊙ M and L n K ,⊕ L ,⊙ L , respectively. In the following, we review the Busemann function on the hyperbolic space. In Euclidean space, the Busemann function associated with the geodesic γ(t) = tv that starts at 0 with unit direction v ∈ S n−1 is B v (x) =−⟨x,v⟩, which coincides, up to a sign, with the inner product. Let H n K ∈ P n K , L n K be a hyperbolic space. We write B v (x) for the Busemann function associated with the ray that emanates from the origin e ∈ H n K in the direction v ∈ S n−1 ⊂ T e H n K . For K < 0, v ∈ S n−1 , and x ∈ H n K , closed forms of the Poincaré and Lorentz Busemann functions [24, Prop. 9] are P n K : B v (x) = 1 √ −K log v− √ −Kx 2 1 + K∥x∥ 2 ! ,(5.40) L n K : B v (x) = 1 √ −K log √ −K (x t −⟨x s ,v⟩) .(5.41) 3 The code is available at https://github.com/GitZH-Chen/HBNN. 151 5.3. Hyperbolic Busemann Neural Networks 101 x 1 1 0 1 x 2 Poincaré 4 0 4 ( x s ) 1 4 0 4 ( x s ) 2 1 3 6 x t Lorentz Figure 5.1: Illustration: red curves are different horospheres of B v . The level sets of a Busemann function are horospheres, the hyperbolic counterpart of Euclidean hyperplanes. In Euclidean space, for a unit direction v, the hyperplanes H v τ = x ∈ R n | ⟨x,v⟩ = τ with τ ∈ R are parallel. Analogously, for fixed v, the horospheres H v τ = x ∈ H n K | B v (x) = τ are equidistant, as established later in Thm. 132. Fig. 5.1 illustrates such horospheres, and Tab. 2.4 summarizes the above correspondence between Euclidean and hyperbolic notions. 5.3.3 Busemann Multinomial Logistic Regression We begin by reformulating the Euclidean MLR, then lift it to hyperbolic space via the Busemann function, introducing BMLR. We also present a point-to-horosphere interpretation. Finally, we compare BMLR with existing hyperbolic MLRs, highlighting our advantages in geometric fidelity, parameterization, and computational efficiency. 5.3.3.1 Formulation The Euclidean MLR softmax(Ax + b) computes the multinomial probability for each class k ∈1,...,C given an input x∈ R n . It admits the inner product form: ∀k, p(y = k | x) = exp (⟨a k ,x⟩ + b k ) P C j=1 exp (⟨a j ,x⟩ + b j ) ,(5.42) where a k ∈ R n and b k ∈ R are the weight and bias for class k. We write p(y = k | x)∝ exp (u k (x)) with u k (x) =⟨a k ,x⟩ +b k . Decomposing the weight vector into a magnitude 152 Chapter 5. Riemannian Neural Networks α k =∥a k ∥ > 0 and a unit direction v k = a k ∥a k ∥ ∈ S n−1 , each logit is u k (x) = α k ⟨v k ,x⟩ + b k .(5.43) As reviewed in Sec. 5.3.2, the Busemann function naturally generalizes the Eu- clidean inner product. Analogously to Eq. (5.43), we define the hyperbolic logits via the Busemann function, yielding BMLR: ∀k, p(y = k | x) = exp (u k (x)) P C j=1 exp (u j (x)) ,(5.44) u k (x) =−α k B v k (x) + b k ,(5.45) with α k > 0, v k ∈ S n−1 , and b k ∈ R as parameters. The following result shows that, as K → 0 − , both the Poincaré and Lorentz BMLRs reduce to the Euclidean MLR. Theorem 130 (Limits of BMLRs). [↓] As K → 0 − , the hyperbolic Busemann functions converge to the Euclidean inner product: Poincaré: B v (x) K→0 − −→−2⟨v,x⟩,(5.46) Lorentz: B v (x) K→0 − −→−⟨v,x s ⟩.(5.47) The hyperbolic BMLRs converge to the Euclidean MLR: Poincaré: u k (x) K→0 − −→ 2α k ⟨v k ,x⟩ + b k ,(5.48) Lorentz: u k (x) K→0 − −→ α k ⟨v k ,x s ⟩ + b k .(5.49) Remark 131 (Intuition). On the Poincaré ball, letting K → 0 − recovers Euclidean geometry [182, App. A.4.2]. For the Lorentz model, as K → 0 − , the temporal co- ordinate diverges while the spatial component approaches R n , making L n K converge to a Euclidean space. Consistently, the Poincaré and Lorentz Busemann functions and the associated BMLR logits reduce to their Euclidean counterparts, providing a natural generalization of Euclidean MLR. 153 5.3. Hyperbolic Busemann Neural Networks 5.3.3.2 Geometric Interpretation The point-to-hyperplane strategy underlying MLR has been established in Sec. 4.2.2.1. We now show that BMLR admits the corresponding interpretation through point-to- horosphere distances. Theorem 132 (Hadamard horosphere distance). [↓] Let (X, d) be a geodesically complete Hadamard space, and let B γ : X → R be the Busemann function associ- ated with a geodesic ray γ : [0,∞)→X. For any τ 1 ,τ 2 ∈ R, define the horospheres by H γ τ i =x∈X | B γ (x) = τ i , i = 1, 2.(5.50) The distance between these horospheres is constant: d H γ τ 1 ,H γ τ 2 = d H γ τ 2 ,H γ τ 1 =|τ 2 − τ 1 |.(5.51) In particular, the point-to-horosphere distance is d (x,H γ τ ) =|B γ (x)− τ|, ∀x∈X.(5.52) Corollary 133 (Point-to-horosphere distance). In a hyperbolic space H n K ∈P n K , L n K (5.53) with curvature K < 0, the point-to-horosphere distance is d (x,H v τ ) =|B v (x)− τ|,(5.54) where H v τ = x| B v (x) = τ denotes the horosphere with respect to direction v ∈ S n−1 . A Euclidean hyperplane can be parameterized by a unit direction, a positive mag- nitude, and a scalar bias. Similarly, we parameterize a hyperbolic horosphere as H v,α,b =x∈H n K |−αB v (x) + b = 0,(5.55) with v ∈ S n−1 , α > 0, and b ∈ R. With this parameterization, the signed point-to- horosphere logit is u k (x) = sign k α k d (x,H v k ,α k ,b k ),(5.56) 154 Chapter 5. Riemannian Neural Networks MethodLogit u k (x), ∀k ∈1,...,CSpaceDist#Params Compact params FLOPs Batch efficiency Euclidean MLR⟨a k ,x⟩ + b k , with a k ∈ R n ,b k ∈ R n RealC(n + 1)✓C(2n)✓ Poincaré MLR [76, Eq. (25)] λ K p k ∥a k ∥ √ −K sinh −1 2 √ −K⟨−p k ⊕ M x,a k ⟩ 1 + K∥−p k ⊕ M x∥ 2 ∥a k ∥ ! , with p k ∈ P n K ,a k ∈ T p k P n K P n K RealC(2n)✗C(19n + 29)✗ Poincaré MLR [181, Eq. (6)] 2 √ −K α k sinh −1 (α− β), α = λ K x √ −K⟨x,v k ⟩ cosh 2 √ −Kb k , β = λ K x − 1 sinh 2 √ −Kb k , with α k > 0,v k ∈ S n−1 ,b k ∈ R P n K RealC(n + 2)✓C(4n + 52)✓ Pseudo-Busemann MLR [162, Cor. 4.3] − d(x,p k ) B v k (−p k ⊕ M x) ∥−p k ⊕ M x∥ , with p k ∈ P n K ,v k ∈ S n−1 P n K PseudoC(2n)✗C(19n + 34)✗ Lorentz MLR [15, Eq. (12)] 1 √ −K sign(α)β sinh −1 √ −K α β , α = cosh √ −Kb k ⟨z k ,x s ⟩− sinh √ −Kb k , β = q cosh( √ −Kb k )z k 2 − (sinh( √ −Kb k )∥z k ∥) 2 , with z k ∈ R n , b k ∈ R L n K RealC(n + 1)✓C(4n + 52)✓ BMLR −α k B v k (x) + b k , with α k > 0,v k ∈ S n−1 ,b k ∈ R P n K L n K RealC(n + 2)✓ P n K : C(6n + 12) L n K : C(2n + 12) ✓ Table 5.11: Comparison of C-class MLR. In Dist, Real means the point-to-hyperplane distance is the real distance, obtained by inf y∈H d(x,y), where H is a hyperplane and d is the geodesic distance; Pseudo denotes a surrogate that coincides with the real distance only in Euclidean geometry. Compact params indicate whether each logit avoids an additional manifold-valued parameter. Batch efficiency indicates whether the MLR can avoid inefficient per-class loops in implementation (see Sec. A.4.2.1). In #Params, we highlight the heaviest in red. In FLOPs, we mark the slowest in red and the fastest in green. where sign k = sign (−α k B v k (x) + b k ). By Thm. 133, the exact point-to-horosphere distance is d (x,H v,α,b ) = |−αB v (x) + b| α .(5.57) Consequently, Eq. (5.56) equals the BMLR logit in Eq. (5.45). Remark 134 (Generality). Since B v (x) = −⟨v,x⟩ in Euclidean geometry, Eqs. (5.55) to (5.57) naturally generalize to their Euclidean counterparts. We also acknowledge Fan et al. [73, Eq. (2) and Prop. 3.1], who used horospheres and point-to-horosphere distances to construct a hyperbolic SVM. However, they con- sidered only the unit Poincaré ball with curvature K =−1, which is a special case of Eqs. (5.55) and (5.57). 5.3.3.3 Comparison with Existing Hyperbolic MLRs Based on the point-to-hyperplane reformulation, recent work extended MLR to the Poincaré [76, 181, 162] and Lorentz [15] models. Ganea et al. [76, Sec. 3.1] introduced the first Poincaré MLR by replacing the Euclidean point-to-hyperplane distance with its hyperbolic counterpart, where the hyperplane is defined by geodesics and the resulting 155 5.3. Hyperbolic Busemann Neural Networks distance is the real point-to-hyperplane distance, obtained as an infimum over the hyper- plane. However, the formulation is not batch efficient (see Sec. A.4.2.1). It also requires per-class parameters a k ∈ T p k P n K and p k ∈ P n K , which leads to over-parameterization. Shimizu et al. [181, Sec. 3.1] alleviated such issues via re-parameterization. Bdeir et al. [15, Sec. 4.3] further developed a Lorentz MLR, but its hyperplanes are defined by the ambient Minkowski space, which is tailored to the Lorentz model and does not fully respect intrinsic hyperbolic geometry. Moreover, Nguyen et al. [162, Cor. 4.3] proposed a Poincaré MLR based on the Busemann function. We refer to it as Pseudo-Busemann MLR, as the induced point-to-hyperplane distance is pseudo, coinciding with the real point-to-hyperplane distance only in Euclidean geometry. It also suffers from over- parameterization and is not batch efficient. As summarized in Tab. 5.11, 4 BMLR unifies advantages that prior hyperbolic MLRs offer only partially. In particular, BMLR respects the real point-to-horosphere dis- tance, uses compact parameters without an additional manifold-valued point, attains the lowest FLOPs on L n K and a competitive cost on P n K , and supports batch-efficient computation. On L n K , its FLOPs are even close to those of the Euclidean MLR. 5.3.4 Busemann Fully Connected Layer We first reformulate the Euclidean FC layer by the Busemann function, then present the manifestations in the Poincaré and Lorentz models. 5.3.4.1 Formulation As discussed in Sec. 5.2.4.2, the Euclidean FC layer can be written as ̄ d (y,H e k ,0 ) =⟨a k ,x⟩ + b k , ∀k ∈1,...,m,(5.58) where ̄ d (y,H e k ,0 ) = sign (⟨e k ,y− 0⟩) d(y,H e k ,0 ) is the signed distance. To extend Eq. (5.58) into hyperbolic space, the right-hand side can be replaced by Eq. (5.45), as it generalizes ⟨a k ,x⟩ + b k . For the left-hand side, a natural idea is to use the signed point-to-horosphere distance. However, as detailed in Sec. A.4.2.2, this may fail to admit a solution for y. We therefore follow the point-to-hyperplane distance in [76, Thm. 5] for the Poincaré model and the one in [15, Eq. (44)] for the Lorentz model. Given x∈H n K , the hyperbolic BFC layer F :H n K ∋ x7→ y ∈H m K is given by solving y 4 Relative to [162, Def. 4.2, Cor. 4.3, and App. B.1.2], the Pseudo-Busemann MLR written here includes an additional sign −; this is intentional and matches their official implementation. 156 Chapter 5. Riemannian Neural Networks via the following m equations: ̄ d (y,H e k ,e ) = u k (x), ∀k ∈1,...,m,(5.59) where u k (x) = −α k B v k (x) + b k with α k > 0,v k ∈ S n−1 ,b k ∈ R as parameters. Here, ̄ d (y,H e k ,e ) is the hyperbolic signed distance from y to the hyperplane passing through the output-space origin e ∈ H m K . Next, we show that the above implicit definition has an explicit solution for the output y. Theorem 135 (Poincaré BFC). [↓] Given an input x ∈ P n K , the Poincaré BFC layer F : P n K → P m K is given by y = ω 1 + q 1− K∥ω∥ 2 , ω = " sinh √ −Ku k (x) √ −K # m k=1 ,(5.60) where u k (x) =−α k B v k (x) + b k with α k > 0,v k ∈ S n−1 ,b k ∈ R as parameters for k = 1,...,m. Theorem 136 (Lorentz BFC). [↓] Given an input x∈ L n K , the Lorentz BFC layer F : L n K → L m K is given by y = " y t y s # = q 1 −K +∥y s ∥ 2 1 √ −K sinh √ −Ku(x) ,(5.61) where u(x) = (u 1 (x),...,u m (x)) ⊤ with u k (x) = −α k B v k (x) + b k . Here, α k > 0,v k ∈ S n−1 ,b k ∈ R are parameters for k = 1,...,m. Analogously to Thm. 130, our BFC layers converge to their Euclidean counterparts as K → 0 − . Theorem 137 (Limits of BFC layers). [↓] As K → 0 − , the hyperbolic BFC layer H n K ∋ x7→ y ∈H m K reduces to a Euclidean FC layer: Poincaré: y k K→0 − −→ α k ⟨v k ,x⟩ + 1 2 b k ,(5.62) Lorentz: (y s ) k K→0 − −→ α k ⟨v k ,x s ⟩ + b k .(5.63) 157 5.3. Hyperbolic Busemann Neural Networks MethodF :H n K ∋ x7→ y ∈H m K Space MethodologyParameters#ParamsFLOPs Möbius [76, Eq. (27)] 1 √ −K tanh ∥Wx∥ ∥x∥ tanh −1 √ −K∥x∥ Wx ∥Wx∥ P n K TangentW ∈ R m×n mn 2nm + 2n +2m + 24 Poincaré FC [181, Eq. (7)] y = ω 1 + q 1− K∥ω∥ 2 , ω k = sinh √ −Ku k (x) √ −K , with u k (x) in Tab. 5.11 P n K Poincaré geometry α k > 0,v k ∈ S n−1 , b k ∈ R, k = 1,...,m m(n + 2)4nm + 71m + 4 Lorentz FC [45, Eq. (3)] y = " q ∥ψ(Wx,v)∥ 2 − 1/K ψ(Wx,v) # , ψ(Wx,v) = λσ v ⊤ x + b ′ Wφ(x) + b ∥Wφ(x) + b∥ L n K Ambient Minkowski W ∈ R m×(n+1) , v ∈ R n+1 , b∈ R m , b ′ ∈ R, λ > 0 m(n + 1) + m +(n + 1) + 2 2nm + 8m +2n + 10 BFC y = ω 1 + q 1− K∥ω∥ 2 , ω = sinh √ −Ku(x) √ −K ; y s = 1 √ −K sinh √ −Ku(x) , y t = r 1 −K +∥y s ∥ 2 , with u k (x) = φ (−α k B v k (x) + b k ) P n K L n K Busemann α k > 0,v k ∈ S n−1 , b k ∈ R, k = 1,...,m m(n + 2) 6nm + 29m + 4 2nm + 30m + 2 Table 5.12: Comparison of hyperbolic FC layers. For simplicity, BFC layers do not involve the gyroaddition and assume φ is the identity map, which is in line with the Möbius and Lorentz FC layers. 5.3.4.2 Generalization Following Sec. 5.2.4.2, we extend the hyperbolic BFC by inserting the activation into Eq. (5.59): ̄ d (y,H e k ,e ) = φ (u k (x)), ∀k ∈1,...,m.(5.64) This is reflected in Thms. 135 and 136 by replacing every u k (x) with φ (−α k B v k (x) + b k ). Moreover, inspired by the Poincaré Möbius transformation [76, Sec. 3.2], a BFC trans- formation could be further followed by a gyroaddition⊕ H : H n K ∋ x7→F (x)⊕ H b∈H m K with b∈H m K as a gyro bias. For example, the Lorentz BFC layer is generalized as L n K ∋ x7→ y = q 1 −K +∥y s ∥ 2 1 √ −K sinh √ −Ku(x) ⊕ L b∈ L m K ,(5.65) where u k (x) = φ (−α k B v k (x) + b k ) with parametersα k > 0,v k ∈ S n−1 ,b k ∈ R m k=1 and b∈ L m K . 5.3.4.3 Comparison with Existing Hyperbolic FC Layers Tab. 5.12 compares BFC with prior hyperbolic FC layers. BFC faithfully respects hy- perbolic geometry, whereas the Möbius and Lorentz FC layers apply Euclidean trans- formations in the tangent or ambient Minkowski space, which can distort intrinsic geometry. BFC also offers flexibility across models, while Poincaré FC and Lorentz FC are tailored to their respective models. In addition, BFC uses a comparable parameter- ization and maintains O(nm) FLOPs. On L n K , its FLOPs are O(2mn), matching the fastest layers. 158 Chapter 5. Riemannian Neural Networks SpaceMethod CIFAR-10 (Num. classes: 10) CIFAR-100 (Num. classes: 100) Tiny-ImageNet (Num. classes: 200) ImageNet-1k (Num. classes: 1000) AccFit Time #ParamsAccFit Time #ParamsAccFit Time #ParamsAccFit Time #Params R n MLR95.14± 0.1210.665.13K77.72± 0.1510.6051.30K65.19± 0.1269.17102.60K71.872263.12513K P n K PMLR95.04± 0.1311.945.14K77.19± 0.5012.1151.40K64.93± 0.3871.90102.80K71.772300.11514K PBMLR-P95.23± 0.0821.9210.24K77.78± 0.1576.84102.40K65.43± 0.27336.58204.80K71.463907.121024K BMLR-P95.32± 0.1412.015.14K78.10± 0.3512.1351.40K66.16± 0.1971.98102.80K73.362300.77514K L n K LMLR94.98± 0.1211.555.13K78.03± 0.2111.7251.30K65.63± 0.1069.27102.60K72.462277.17513K BMLR-L95.25± 0.0211.085.14K78.07± 0.2611.2251.40K65.99± 0.1469.19102.80K73.242276.53514K Table 5.13: Top-1 image classification accuracy (%) of MLR methods on the ResNet-18 backbone. The best results within each hyperbolic model are bold. The slowest MLR and largest parameter count are shown in red. 020406080100 Epoch 40 50 60 70 Top-1 Accuracy Poincaré MLR PMLR PBMLR-P BMLR-P 020406080100 Epoch 40 50 60 70 Top-1 Accuracy Lorentz MLR LMLR BMLR-L Figure 5.2: Validation accuracy curves on ImageNet-1k. 5.3.5 Experiments We first compare BMLRs with prior hyperbolic MLRs on three architectures: ResNet-18 (image classification), CNN (genome sequences), and HGCN (node classification). We then compare BFC with prior hyperbolic FC layers on link prediction. All experiments use both the Poincaré and Lorentz models. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.5. 5.3.5.1 Image Classification Setup. Following Sec. 5.2.6.2, we use a hybrid architecture with a ResNet-18 [97] back- bone and an MLR head. We compare Euclidean MLR with hyperbolic variants in both models. In Poincaré, we evaluate Poincaré MLR (PMLR) re-parameterized by Shimizu et al. [181], Pseudo-Busemann MLR (PBMLR-P) [162], and our BMLR-P. In Lorentz, we evaluate Lorentz MLR (LMLR) [15] and our BMLR-L. For hyperbolic MLRs, we 159 5.3. Hyperbolic Busemann Neural Networks BenchmarkTaskData Set Num.P n K L n K classesPMLRPBMLR-PBMLR-PLMLRBMLR-L TEB Retrotransposons LTR Copia275.34± 1.0274.37± 1.4876.73± 1.0873.01± 1.0775.86± 1.52 LINEs285.54± 0.6185.92± 0.6586.05± 1.0883.14± 0.8086.72± 0.58 SINEs295.30± 0.8595.34± 1.5895.99± 0.7496.70± 0.8796.29± 0.59 DNA transposons CMC-EnSpm283.39± 0.5683.62± 1.0084.03± 0.7181.78± 1.0584.15± 1.00 hAT-Ac289.38± 0.9089.86± 0.5489.62± 0.7488.94± 0.6990.70± 0.51 Pseudogenes processed272.45± 1.4971.99± 2.0473.09± 1.6673.71± 1.7673.32± 1.65 unprocessed275.37± 2.2771.99± 1.4775.71± 1.8974.54± 1.9876.15± 1.61 GUE Core Promoter Detection tata280.95± 1.4779.32± 2.4480.29± 1.6380.90± 1.1581.76± 1.16 notata270.02± 0.5270.60± 0.7570.48± 0.3571.26± 0.5670.43± 0.39 all 267.64± 0.7768.02± 0.6368.50± 0.6167.63± 0.5668.36± 1.07 Promoter Detection tata280.30± 1.5980.27± 2.7182.83± 1.6983.27± 1.9582.55± 1.54 notata292.63± 0.3693.05± 0.3292.75± 0.5191.74± 0.5792.60± 0.49 all290.53± 0.5090.79± 0.7790.20± 0.6589.34± 0.4089.82± 0.45 Covid Variant ClassificationCovid974.09± 0.2570.84± 0.8073.40± 0.3064.07± 0.5172.45± 0.21 Species Classification Virus2067.24± 2.1059.17± 3.3277.12± 1.2371.34± 2.0577.21± 1.04 Fungi2515.06± 1.3218.75± 1.7730.01± 0.7615.07± 1.8830.14± 2.48 Table 5.14: Genomic MCC of MLR methods under the CNN backbone. The best results within each hyperbolic model are bold. map the ResNet-18 features to the target hyperbolic space before classification. We evaluate on CIFAR-10 [126], CIFAR-100 [126], Tiny-ImageNet [129], and ImageNet-1k [66]. On the first three data sets, we conduct five-fold experiments. Results. Tab. 5.13 reports top-1 validation accuracy, fit time per epoch, and classifier-head parameters. Fig. 5.2 presents the ImageNet-1k accuracy curves. Overall, BMLR-P and BMLR-L consistently outperform prior hyperbolic MLRs with compara- ble parameters. Within each hyperbolic model, the accuracy margin over prior hyper- bolic MLRs increases with the number of classes, from CIFAR-10 to CIFAR-100 and Tiny-ImageNet, with the largest gains on ImageNet-1k. This demonstrates the advan- tage of BMLR as task complexity increases. Besides, PBMLR-P uses approximately double the head parameters and is markedly slower due to complex batch-inefficient computation, whereas BMLR-L achieves the fastest fit time among all hyperbolic MLRs. 5.3.5.2 Genome Sequence Learning Setup. Similar to Sec. 5.2.6.4, we evaluate hyperbolic MLRs on genome sequence learning. Following Khan et al. [119], we adopt a CNN backbone, which consists of three convolutional blocks and an MLR head. Similar to Sec. 5.3.5.1, we compare our BMLR against previous hyperbolic MLR heads by replacing the final Euclidean MLR with a hyperbolic MLR. We validate on two benchmarks: TEB [119] and Genome Understanding Evaluation (GUE) [234], covering a total of 16 data sets. Results. Tab. 5.14 summarizes five-fold average MCC across TEB and GUE. Com- pared with other hyperbolic MLRs, our BMLR-P and BMLR-L achieve higher MCC in most tasks. Similar to Sec. 5.3.5.1, the gains are more pronounced on complex data sets 160 Chapter 5. Riemannian Neural Networks Data Set P n K L n K PMLR PBMLR-PBMLR-PLMLRBMLR-L LTR Copia3.965.114.063.923.77 LINEs5.806.905.735.505.32 SINEs 1.361.801.391.381.28 CMC-EnSpm3.114.383.042.892.82 hAT-Ac4.375.374.384.133.96 processed5.025.944.894.684.58 unprocessed3.303.903.293.052.95 CPD-tata0.721.350.700.670.62 CPD-notata6.1312.406.005.735.71 CPD-all6.8514.356.596.356.33 PD-tata0.981.200.971.040.94 PD-notata8.3310.538.298.117.84 PD-all 9.4012.079.199.068.83 Covid28.9645.5227.9727.5826.67 Virus25.2829.5725.6725.1224.85 Fungi 6.418.966.416.256.25 Table 5.15: Fit time (s/epoch) on genome sequence learning. The fastest times are bold and the slowest ones are red. SpaceMethodMethodology DiseaseAirportPubMedCora δ = 0δ = 1δ = 3.5δ = 11 P n K MöbiusTangent76.35± 1.8393.31± 0.4194.93± 0.0690.80± 0.56 Poincaré FCPoincaré geometry79.45± 1.0194.31± 0.1694.24± 0.2588.21± 0.72 BFC-PBusemann80.45± 0.9394.88± 0.3994.85± 0.0791.94± 0.32 L n K LTFCTangent71.32± 5.3692.68± 0.3594.85± 0.1789.37± 0.64 Lorentz FCAmbient Minkowski72.78± 2.0492.99± 0.3394.20± 0.1092.06± 0.62 BFC-LBusemann78.36± 0.5195.37± 0.1794.90± 0.0492.28± 0.12 Table 5.16: Comparison of hyperbolic FC layers on link prediction. The best results within each hyperbolic model are bold. with more classes, e.g.,, Virus (20 classes) and Fungi (25 classes), demonstrating the effectiveness of our approach. Tab. 5.15 reports fit time per epoch, where PBMLR-P is consistently the slowest due to batch inefficiency, and BMLR-L is the fastest. 5.3.5.3 Node Classification Setup. Following Nguyen et al. [162], we adopt the HGCN [38] backbone to evaluate our BMLR on graph data sets, including Disease [4], Airport [229], PubMed [155], and Cora [177]. The HGCN backbone consists of a hyperbolic Graph Convolutional Network (GCN) and an MLR as the final classification layer. Both the GCN and the MLR are built on the hyperbolic space. The vanilla HGCN uses a tangent MLR, which maps 161 5.3. Hyperbolic Busemann Neural Networks SpaceMethod DiseaseAirportPubMedCora δ = 0δ = 1δ = 3.5δ = 11 P n K HGCN86.87± 2.5885.34± 1.1676.29± 0.9876.56± 0.81 HGCN-PMLR88.98± 1.9684.78± 1.4876.02± 1.0977.47± 1.15 HGCN-PBMLR-P89.05± 0.7885.04± 0.9775.89± 0.7877.90± 1.00 HGCN-BMLR-P92.45± 0.9686.02± 0.5377.36± 0.7378.48± 1.52 L n K HGCN87.83± 0.7784.94± 1.4076.49± 0.8877.37± 1.72 HGCN-LMLR89.72± 1.5182.61± 1.0175.44± 1.1769.91± 3.61 HGCN-BMLR-L90.80± 1.1585.27± 1.1777.30± 0.4177.65± 2.10 Table 5.17: Node classification F1 scores of hyperbolic MLRs on the HGCN backbone, where δ denotes graph hyperbolicity (lower is more hyperbolic). The best results within each hyperbolic model are bold. features into the tangent space via Log e and applies a Euclidean MLR. We replace this with different hyperbolic MLRs. Results. Tab. 5.17 reports average F1 scores. Our BMLRs consistently outperform prior hyperbolic MLRs within each hyperbolic model. As graphs become less hyperbolic, that is, for larger δ, existing hyperbolic heads could underperform the vanilla tangent- based MLR, for example, PBMLR-P on PubMed, and LMLR on Airport, PubMed, and Cora. Especially on Cora, which has the largest δ, LMLR lags the tangent baseline by a large margin (69.91 vs. 77.37). In contrast, BMLR remains the top performer across all δ values, indicating that Busemann-based decoding robustly strengthens HGCN over a broader range of graph hyperbolicity. 5.3.5.4 Link Prediction Setup. We compare our BFC layers with prior hyperbolic FC layers, including the Möbius layer [76] that operates via the tangent space, the Lorentz FC layer [45] that op- erates through the ambient Minkowski space, and the Poincaré FC layer [181]. Mimick- ing the Möbius layer, we also implement a Lorentz tangent FC layer, Exp 0 (M Log 0 (x)), referred to as LTFC. Following Chami et al. [38], we evaluate on Disease, Airport, PubMed, and Cora. Following the HNN implementation [76, 38], all methods share the same backbone with two FC layers. For a fair comparison, all hyperbolic FC layers are followed by a gyroaddition biasing. For BFC, we use φ = tanh on Airport and Cora and the identity map on the other two data sets. Results. Tab. 5.16 reports five-fold test AUC. Our BFC layers generally outper- form prior hyperbolic FC layers. The gains are most pronounced on Disease, which is the most hyperbolic (δ = 0), where Busemann-based decoding is markedly more effec- tive than tangent or ambient methods, indicating better capture of intrinsic hyperbolic 162 Chapter 5. Riemannian Neural Networks SpaceMethod DiseaseAirportPubMedCora Fit Time #ParamsFit Time #ParamsFit Time #ParamsFit Time #Params P n K Möbius0.02004640.05354800.112082880.022923216 Poincaré FC0.01985280.05365440.117683520.024823280 BFC-P0.02015280.05125440.112383520.023123280 L n K LTFC0.03434640.08184800.163382880.037023216 Lorentz FC0.02325630.07155800.153788760.026124737 BFC-L0.02445280.07135440.152583520.028023280 Table 5.18: Efficiency comparison: fit time (s/epoch) and parameter count. Slowest results and largest parameter counts are in red. geometry. This observation aligns with geometric intuition, since tangent space or am- bient space approximations inherently struggle to represent curved manifolds in highly non-Euclidean cases. Training Time and Parameter Count. Tab. 5.18 summarizes fit time per epoch and parameter counts. Our BFC layers achieve training time and model size comparable to existing layers. In particular, LTFC is the slowest due to costly logarithmic and exponential maps, and LFC uses the largest number of parameters among Lorentz variants. 5.4 Full-Rank Correlation Networks 5.4.1 Introduction The preceding two sections studied manifold-specific designs for hyperbolic learning. We now turn to neural networks on full-rank correlation manifolds. Covariance matrices in the SPD manifold have achieved success in various applications, with many deep network architectures adapted to leverage their Riemannian geometries [106, 31, 37, 59, 167, 123, 206, 47, 118, 134, 172, 215, 115, 104]. In contrast, correlation matrices, despite serving as statistically compact alternatives to covariance matrices [7], remain unexplored in deep learning. As discussed in Sec. 2.9.2, Riemannian structures for correlation matrices have only recently been developed. David and Gu [62] identified full-rank correlation matrices as a quotient manifold of the SPD manifold, referred to as the correlation manifold. However, this quotient geometry does not guarantee uniqueness or closed forms of the Riemannian logarithm and Fréchet mean [195, Sec. 1.1]. To close this gap, Thanwerdas and Pennec [195] proposed three theoretically and computationally convenient geome- tries: ECM, LECM, and PHCM. Thanwerdas [191] further introduced two efficient permutation-invariant metrics: OLM and LSM. These Riemannian structures provide 163 5.4. Full-Rank Correlation Networks promising foundations for extending Euclidean deep learning to the correlation mani- fold. On the other hand, several fundamental layers in Euclidean deep learning, such as MLR, FC, and convolutional layers, have been extended to different manifolds by leveraging their rich Riemannian or algebraic structures [106, 107, 108, 76, 37, 45, 181, 15, 51, 161]. For the SPD manifold, these layers have been constructed using bilinear mapping [106], weighted Fréchet means [37], gyrovector spaces [159, 161], and Riemannian geometry [49, 51]. Inspired by these advancements, we develop MLR, FC, and convolutional layers for correlation manifolds in a geometrically intrinsic manner. We begin by systematically introducing four types of correlation-based MLR, FC, and convolutional layers, corre- sponding to ECM, LECM, OLM, and LSM, respectively. Besides, we discuss backprop- agation through Riemannian computations over the correlation manifold, with novel approaches for accurate backpropagation under OLM and LSM. As the above four metrics have zero curvature, our next focus is to build correlation layers under the geometry of non-zero curvature. We target PHCM, induced by the product of multi- ple hyperbolic spaces [195, Thm. 4.4]. By adapting existing Poincaré-based hyperbolic MLR, FC, and convolutional layers designed for a single Poincaré ball [76, 181], we con- struct their counterparts on the correlation manifold. Together with the corresponding backpropagation mechanisms, these layers constitute complete Correlation Networks (CorNets) under different geometries. The effectiveness is validated by experiments comparing our approach against existing SPD and Grassmannian baselines. Tab. 5.19 summarizes the correspondence between Euclidean and our correlation layers. In summary, our main contributions are as follows: (1) We systematically extend MLR, FC, and convolutional layers to the correlation manifold under five geometries: four with zero curvature and one with non-zero curvature. The developed layers enable flexible variation of the latent geometry under a consistent network architecture, allowing for straightforward comparisons across different correlation geometries. (2) We develop accurate backpropagation of Riemannian computations under OLM and LSM. (3) We conduct experiments against existing SPD and Grassmannian networks to demonstrate the effectiveness of correlation embeddings and networks. Outline. Sec. 5.4.2 constructs correlation MLR, FC, and convolutional layers under four flat geometries. Sec. 5.4.3 develops their counterparts under a non-zero-curvature 164 Chapter 5. Riemannian Neural Networks SpaceEuclidean R n Correlation Cor + (n) C-class MLRf : R n ∋ x7→ p = softmax(Ax + b)∈ R C f : Cor + (n)∋ X 7→ p∈ R C FC layerF : R n ∋ x7→ y = Ax + b∈ R m F : Cor + (n)∋ X 7→ Y ∈ Cor + (m) ConvolutionKernel-based FC in each receptive fieldKernel-based correlation FC in each receptive field GeometryEuclideanECM, LECM, OLM, LSM and PHCM Table 5.19: Correspondence between Euclidean and correlation-based layers. For con- volution, kernel-based FC refers to applying a convolution kernel to a receptive field, which is an FC transformation. geometry and establishes the order-invariance of the associated β-operations. Sec. 5.4.4 presents backpropagation over the correlation geometries, and Sec. 5.4.5 evaluates Cor- Nets under these five geometries. Proofs are deferred to Sec. B.8. 5.4.2 Log-Euclidean Correlation Layers Since ECM, LECM, OLM, and LSM are derived via diffeomorphisms from Euclidean spaces, they are collectively termed Log-Euclidean metrics [191]. This motivates the principled development of MLR, FC, and convolutional layers [191]. 5.4.2.1 Log-Euclidean Correlation MLRs As discussed in Sec. 4.3.2.1, the MLR can be rewritten by point-to-hyperplane for- mulations. We use the same margin-distance infimum, while the Euclidean isometries of ECM, LECM, OLM, and LSM make it possible to solve this infimum exactly un- der all four geometries. To avoid over-parameterization, we follow Sec. 5.2.4.1 and set P k = Exp E (γ k [Z k ]) and A k = PT E→P k (Z k ), with [Z k ] = Z k ∥Z k ∥ E as the unit direction vector of Z k . Here, E is the origin of M, while γ k ∈ R and Z k ∈ T E M ∼ = R m are the MLR parameters. This is a concrete instance of the trivialization strategy reviewed in Sec. 2.7. Under this trivialization, each hyperplane H A k ,P k is denoted as H Z k ,γ k . As all Log-Euclidean metrics are isometric to Euclidean spaces, the corresponding MLRs admit principled closed forms. Theorem 138. [↓] Let M,g M be an m-dimensional manifold that is isometric to the standard Euclidean space R m via the diffeomorphism φ :M→ R m . Denoting E = φ −1 (0) with 0 as the zero vector, each v k (X) and margin hyperplane H Z k ,γ k in the C-class Riemannian MLR are v k (X) = ⟨φ(X),φ ∗,E (Z k )⟩− γ k ∥φ ∗,E (Z k )∥ and H Z k ,γ k =X ∈M| v k (X) = 0, respectively. Here, Z k ∈ T E M ∼ = R m and γ k ∈ R for 1≤ k ≤ C are MLR parameters, while φ ∗ is the differential. 165 5.4. Full-Rank Correlation Networks Simple computations show that ECM: φ EC (I) = 0, LECM: log◦Θ(I) = 0, OLM: Log ◦ (I) = 0, LSM: Log ⋆ (I) = 0. (5.66) Therefore, we define the origin of the correlation manifold under four Log-Euclidean metrics as the identity matrix. Besides, Thm. 138 suggests that Log-Euclidean MLRs can be obtained modulo the calculation of diffeomorphisms and their differentials at the identity matrix I. Proposition 139 (Differentials). [↓] For any tangent vector V ∈ T I Cor + (n) ∼ = Hol(n), the differentials of φ EC , log◦Θ, Log ◦ , and Log ⋆ at the identity matrix I are φ EC ∗,I (V ) =⌊V⌋, (log◦Θ) ∗,I (V ) =⌊V⌋, Log ◦ ∗,I (V ) = V,Log ⋆ ∗,I (V ) = V − diag(V 1), (5.67) where diag : R n → Diag(n) returns a diagonal matrix, and 1 = (1,· , 1) ⊤ ∈ R n . Putting Thm. 139 into Thm. 138, we obtain correlation MLRs under four Log- Euclidean metrics. Theorem 140 (Log-Euclidean MLRs). Given C ∈ Cor + (n), the logits v k (C) for the k-th class in the correlation MLRs under four Log-Euclidean metrics are v EC k (C) =⟨⌊Θ(C)⌋,⌊Z k ⌋⟩− γ k ∥⌊Z k ⌋∥, v LEC k (C) =⟨log◦Θ(C),⌊Z k ⌋⟩− γ k ∥⌊Z k ⌋∥, v OL k (C) =⟨Log ◦ (C),Z k ⟩− γ k ∥Z k ∥, v LS k (C) = Log ⋆ (C), Log ⋆ ∗,I (Z k ) − γ k Log ⋆ ∗,I (Z k ) , (5.68) where Z k ∈ Hol(n) and γ k ∈ R are parameters. 5.4.2.2 Log-Euclidean FC and Convolutional Layers Following Sec. 5.2.4.2, we now generalize this point-to-hyperplane FC construction to the correlation manifold. Definition 141 (Correlation FC layers). Given a metric g, the correlation FC layer F : Cor + (n) ∋ X 7→ Y ∈ Cor + (m) returns the output Y by solving the following 166 Chapter 5. Riemannian Neural Networks d = m(m−1) /2 equations: s k d(Y,H O k ,I ) = v k (X;Z k ,γ k ),1≤ k ≤ d,(5.69) where s k = sign (⟨Log I (Y ),O k ⟩ I ), I is the identity matrix, d is the dimension of Cor + (m), O k d k=1 is an orthonormal basis over T I Cor + (m), d(·,·) is the margin distance to the hyperplane H O k ,I , and v k is defined by Sec. 4.3.2.1 for Cor + (n). The FC parameters are Z k ∈ Hol(n) d k=1 and γ k ∈ R d k=1 . Sec. A.4.3.1 details how Thm. 141 extends the existing SPD, Poincaré, and Eu- clidean FC layers. Although Thm. 141 is implicitly defined by d equations, the FC layers under four Log-Euclidean geometries admit explicit expressions in a principled manner. Analogous to Thm. 138, a corresponding result for the FC layer is presented in Thm. 201, which yields the Log-Euclidean FC layers. Theorem 142 (Log-Euclidean FC layers). [↓] Given an input correlation C ∈ Cor + (n), the correlation FC layers F (·) : Cor + (n) → Cor + (m) under different Log-Euclidean metrics are ECM: Y = Cor◦ Chol −1 V EC + I m ,(5.70) LECM: Y = Cor◦ Chol −1 ◦ exp V LEC ,(5.71) OLM: Y = Exp ◦ V OL ,(5.72) LSM: Y = Cor◦ exp V LS ,(5.73) where the (i,j)-th elements in V EC ∈ LT 0 (m), V LEC ∈ LT 0 (m), V OL ∈ Hol(m), and V LS ∈ Row 0 (m) are V EC ij = v EC ij (C), if i > j 0,otherwise (5.74) V LEC ij = v LEC ij (C), if i > j 0,otherwise (5.75) V OL ij = v OL ij (C) √ 2 , if i > j V OL ji ,if i < j 0,otherwise (5.76) 167 5.4. Full-Rank Correlation Networks Input 3-Channel Correlation Matrices C 1 ∈Cor + (n) C 2 ∈Cor + (n) C 3 ∈Cor + (n) FC Transformation on Each Receptive Field C 1 ,C 2 ∈ ( Cor + (n) ) 2 ̃ C 1 ∈Cor + (m) F 1 ̃ C 2 ∈Cor + (m) F 2 C 2 ,C 3 ∈ ( Cor + (n) ) 2 ̃ C 3 ∈Cor + (m) F 1 ̃ C 4 ∈Cor + (m) F 2 Split Figure 5.3: Illustration of the Log-Euclidean 1D convolution with two kernels. The 3-channel input is first split into two receptive fields along the channel dimension. In each receptive field, two kernels are applied to the product space. V LS ij = v LS ij (C) / √ 6,if m > i > j ≥ 1 v LS i (C) / √ 3,if m > i≥ 1 V LS ji ,if i < j − P m−1 k=1 V LS kj ,if i = m, 1≤ j < m P m−1 k=1 P m−1 l=1 V LS lk , if i = j = m (5.77) Each v g ij with g ∈EC, LEC, OL, LS is defined by Eq. (5.68) with parameters Z ij ∈ Hol(n) and γ ij ∈ R. For v EC ij , v LEC ij , and v OL ij , the indices satisfy i,j = 1,...,m and i > j. For v LS ij , they satisfy i,j = 1,...,m− 1 and i≥ j. Correlation Convolution. Following the convolution-as-FC construction in Sec. 5.2.4.3, we develop the correlation convolution. The c-channel correlation matrices C i ∈ Cor + (n) c i=1 within a receptive field are first concatenated into C ∈ (Cor + (n)) c . For each convolution kernel, C is then fed into a correlation FC layer. 5 Fig. 5.3 illustrates the above process. 5 Thm. 142 naturally supports product geometries, which are detailed in Sec. A.4.3.2. 168 Chapter 5. Riemannian Neural Networks 5.4.3 Poly-Hyperbolic-Cholesky Layers As detailed in Sec. 2.9.2, the space L n , consisting of the Cholesky factors of Cor + (n), can be identified with the product of n− 1 hyperbolic open hemispheres, PHS n−1 = Q n−1 i=1 HS i . We focus on the widely used hyperbolic Poincaré ball, whose MLR, FC, and β-concatenation components are reviewed in Sec. A.2.4. In the following, we focus on the canonical Poincaré ball (K = −1), namely the unit Poincaré ball P n . We first identify the correlation manifold with the poly-Poincaré space P n−1 = Q n−1 i=1 P i , the product of n− 1 unit Poincaré balls. Then, we develop correlation layers from the layers on a single Poincaré space. 5.4.3.1 Correlation Geometry via Poincaré Balls Proposition 143 (Isometries). [↓] The open hemisphere HS n is isometric to the unit Poincaré ball P n by ψ HS n →P n ((x ⊤ ,x n+1 ) ⊤ ) = x 1 + x n+1 , ψ P n →HS n (y) = 1 1 +∥y∥ 2 2y 1−∥y∥ 2 ! , (5.78) with (x ⊤ ,x n+1 ) ⊤ ∈ HS n ⊂ R n × R + and y ∈ P n ⊂ R n . Thm. 143 indicates that Cor + (n) can be identified with P n−1 = Q n−1 i=1 P i via the diffeomorphism Φ: C Chol 7−→ 10 ·0 L 21 L 22 ·0 . . . . . . . . . . . . L n1 L n2 · L n Q n−1 i=1 ψ i 7−→ ψ 1 (h 1 ) . . . ψ n−1 (h n−1 ) (5.79) with C ∈ Cor + (n), h i = (L i+1,1 ,· ,L i+1,i+1 ) ⊤ ∈ HS i , and ψ i = ψ HS i →P i . This identi- fication motivates us to construct the correlation layers using the corresponding layers over Poincaré spaces. 5.4.3.2 Revisiting Poincaré Layers The Poincaré MLR and FC layers are reviewed in Sec. A.2.4 and follow the point- to-hyperplane logic discussed in Chapter 4 and Secs. 5.2 and 5.3. The convolutional 169 5.4. Full-Rank Correlation Networks Input 푐-Channel Correlation Matrices ...... Chol Cholesky Factors ... C 1 ∈Cor + (n) L 1 =Chol(C 1 )∈LT + (n) L c =Chol(C c )∈LT + (n) 10·0 L 1 21 L 1 22 ·0 . . . . . . . . . . . . L 1 n1 L 1 n2 ·L 1 n 10·0 L c 21 L c 22 ·0 . . . . . . . . . . . . L c n1 L c n2 ·L c n Φ β-Concate Ψ 1 ( ( L 1 21 ,L 1 22 ) ! ) ∈P 1 . . . Ψ n−1 ( ( L 1 n1 ,·,L 1 n ) ! ) ∈P n−1 Ψ 1 ( (L c 21 ,L c 22 ) ! ) ∈P 1 . . . Ψ n−1 ( (L c n1 ,·,L c n ) ! ) ∈P n−1 x 1 ∈P n−1 = ∏ n−1 i=1 P i ⊂R n(n−1) 2 x∈P N N=c n(n−1) 2 C c ∈Cor + (n) Poincaré FC β-Split ̃x∈P M M= m(m−1) 2 Ψ −1 10·0 ̃ L 21 ̃ L 22 ·0 . . . . . . . . . . . . ̃ L m1 ̃ L m2 · ̃ L m x c ∈P n−1 = ∏ n−1 i=1 P i ⊂R n(n−1) 2 ̃ C∈Cor + (m) Poincaré MLR Classification Identifying the Correlation Manifold with the Poly-Poincaré Space FC Transformation MLR Classification ̃ L∈LT + (m) Poly-Poincar ́e Vectorsx i Figure 5.4: Illustration of the PHCM convolution and MLR. The multi-channel input correlation matrices are denoted asC i c i=1 . For the convolutional layer, the illustration focuses on the transformation within a receptive field and assumes a single-channel output. construction uses the Poincaré β-concatenation defined below. The Poincaré convolutional layer shares a logic similar to the correlation convolu- tion, except it uses β-concatenation to concatenate the Poincaré vectors in each re- ceptive field [181, Secs. 3.3–3.4], which can stabilize the norm of the Poincaré vec- tor. The Poincaré β-concatenation generalizes the Euclidean concatenation via the scaled concatenation in the tangent space. Given inputs x i ∈ P n i N i=1 , it is defined as Exp 0 β n β −1 n 1 v ⊤ 1 ,· ,β −1 n N v ⊤ N ⊤ ∈ P n , where v i = Log 0 (x i ) and n = P N i=1 n i . Here, β n i and β n are defined by the beta function β α = B ( α /2, 1 /2). The inverse is called the Poincaré β-split. The Poincaré convolution is: (1) β-concatenating the multi-channel feature in a given receptive field; and (2) performing the Poincaré FC transformation. 5.4.3.3 Building Poly-Hyperbolic-Cholesky Layers PHCM MLR. The input multi-channel correlation matrices, C =C i ∈ Cor + (n) c i=1 , are first mapped into poly-Poincaré spaces as x = x i = Φ(C i ) ∈ P n−1 c i=1 . The resulting Poincaré vectors are then β-concatenated into a single Poincaré vector x∈ P N , where N = c n(n−1) 2 . This concatenated vector is subsequently fed into the Poincaré MLR for classification. PHCM Convolutional and FC Layer. The convolutional layer follows a logic similar to Log-Euclidean convolution. The multi-channel correlation matrices within a receptive field C = C i ∈ Cor + (n) c i=1 are first mapped to a β-concatenated Poincaré vector x∈ P N as in the PHCM MLR, which is then fed into the Poincaré FC layer for dimensionality transformation. This produces a vector ex ∈ P M , with M = k m(m−1) 2 , which is then split using β-split. Subsequently, applying Φ −1 reconstructs new k×m×m 170 Chapter 5. Riemannian Neural Networks correlation matrices. When the input is a single correlation matrix, it is reduced to the correlation FC. Fig. 5.4 illustrates the PHCM layers. However, there is an underlying ambiguity in the above discussion. To clarify, we write each x i ∈ P n−1 in x as x i = p i 1 ∈ P 1 ,· ,p i n−1 ∈ P n−1 , which gives x = p i j ∈ P j i=c,j=n−1 i=1,j=1 . We can either concatenate twice by i → j or once along both i and j. A similar issue arises with β-split. The following theorem establishes this invariance. Theorem 144 (Order-invariance). [↓] Given multichannel data x i 1 ,...,i n ∈ P n i n with i j ∈ 1,...,N j , applying the β-concatenation sequentially n times in the order i n → · → i 1 is equivalent to a single β-concatenation along all indices simul- taneously. Similarly, β-splitting x ∈ P N into multichannel data x i 1 ,...,i n ∈ P n i n with i j ∈ 1,...,N j and N = Q n−1 j=1 N j P N n i n =1 n i n under the sequential order i 1 → · → i n is identical to the one under a single β-split to generate all indices simultaneously. Therefore, we always conduct the β-operation simultaneously along both i and j. 5.4.4 Backpropagation over Correlation Geometries Except for D and D ⋆ , all computations involved in the five metrics can be backprop- agated using existing techniques or PyTorch’s auto-differentiation. The matrix loga- rithm, matrix exponentiation, and Cholesky decomposition, together with their differ- entials and backpropagation, are reviewed in Sec. 2.8. It therefore remains to discuss D and D ⋆ . D and D ⋆ . Their gradients can be backpropagated either approximately through their iterative algorithms or accurately using the following two propositions. Proposition 145 (Gradients w.r.t. D). [↓] Let l(·) be the loss function and define F : Hol(n) → S n by F (H) = Y = D(H) + H for any symmetric hollow matrix H, where S n is the Euclidean space of n× n symmetric matrices. Let Y = U ∆U ⊤ be the eigendecomposition with (δ 1 ,· ,δ n ) as eigenvalues. Given the succeeding gradient ∂l ∂Y , the output gradient ∂l ∂H is ∂l ∂H = off ∂l ∂Y − exp ∗,Y D (H 0 ) −1 Dv ∂l ∂Y 1 ⊤ ,(5.80) with H 0 ∈ S n ++ having entries H 0 il = P j,k U ij U ik U lj U lk [L exp ] j,k , where L exp is the 171 5.4. Full-Rank Correlation Networks Loewner matrix in Eq. (2.91) specialized to f = exp and σ i = δ i . Here, D(·) : R n×n → Diag(n) extracts the diagonal matrix, while Dv(·) : R n×n → R n returns a vector of diagonal elements. Besides, off(·) subtracts the diagonal matrix from a matrix, and exp ∗,Y is the differential of the symmetric matrix exponential given by Eq. (2.90). Proposition 146 (Gradients w.r.t. D ⋆ ). [↓] Following the notation in Thm. 145, define F : Cor + (n) → Row + 1 (n) by F (C) = Σ = D ⋆ (C)CD ⋆ (C), where Row + 1 (n) is the manifold of n× n SPD matrices with unit row sum. Given the succeeding gradient ∂l ∂Σ , the output gradient ∂l ∂C is ∂l ∂C = ∆ ∂l ∂Σ − (I + Σ) −1 ev1 ⊤ sym ∆,(5.81) where ∆ = D(Σ) 1 /2 , ev = Dv Σ ∂l ∂Σ + ∂l ∂Σ Σ , I is the identity matrix, and 1 ∈ R n is the vector with all entries equal to 1. Here, (A) sym = A+A ⊤ 2 . 5.4.5 Experiments We construct Riemannian networks on the correlation manifold, termed CorNets, using the proposed convolutional and MLR layers. Following previous work [106, 31, 50], we evaluate our approach on the Radar data set [31] for radar signal classification, along with the HDM05 [153], FPHA [80] and NTU120 [138] data sets for human action recog- nition. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.6. Implementation. We denote CorNet-Metric as the CorNet composed of correla- tion convolution and MLR layers under a specified metric. In line with Nguyen et al. [161], each CorNet consists of one correlation convolutional layer followed by a correla- tion MLR layer, trained with cross-entropy loss. Following Wang et al. [210], Nguyen et al. [161], each raw feature is modeled as a multi-channel [c,n,n] SPD tensor. Since matrix power effectively activates SPD matrices by deforming their geometry, as de- tailed in Sec. 4.3.3.1 and prior work [194, 53], we first apply a matrix power, and then convert the result to correlation matrices as the input of CorNet. Due to trivializa- tion, all trainable manifold-valued parameters are represented by Euclidean parameters and optimized by standard Euclidean optimizers. We compare CorNets against rep- resentative Grassmannian and SPD networks, including GrNet [108], GyroGr [159], 172 Chapter 5. Riemannian Neural Networks ManifoldMethod RadarHDM05FPHANTU120 Mean±STD TimeMean±STD TimeMean±STD TimeMean±STD Time Grassmann GrNet [108]90.48± 0.761.3963.19± 0.701.6485.31± 0.900.7057.59± 0.2250.97 GyroGr ∗ [159]90.64± 0.571.3858.32± 1.232.4879.62± 0.490.7053.76± 0.18136.96 GyroGr-Scaling ∗ [159]88.88± 1.521.6339.75± 0.933.5258.62± 1.661.0343.90± 0.23338.01 SPD SPDNet [106]93.25± 1.100.6664.57± 0.610.5085.59± 0.720.2851.25± 0.3612.77 SPDNetBN [31]94.85± 0.991.2571.28± 0.790.9489.33± 0.490.5854.35± 0.4319.78 SPDResNet-AIM [118]95.71± 0.370.9664.95± 0.821.2386.63± 0.550.6957.33± 0.3523.84 SPDResNet-LEM [118]95.89± 0.860.7770.12± 2.450.5585.07± 0.990.3061.34± 2.0213.00 SPDNetLieBN-AIM [50]95.47± 0.901.2171.83± 0.691.1590.39± 0.660.9758.20± 0.4631.10 SPDNetLieBN-LCM [50]94.80± 0.711.1071.78± 0.441.1186.33± 0.430.5957.96± 0.4322.06 SPDNetMLR [51]94.59± 0.820.6665.90± 0.935.4685.60± 0.430.8858.59± 0.1322.48 GyroLE ∗ [159]96.24± 0.240.7973.17± 0.372.8690.73± 0.921.5959.29± 0.4222.08 GyroLC ∗ [159]93.60± 1.310.6667.53± 0.851.4976.10± 0.630.7859.29± 0.4214.14 GyroAI ∗ [159]96.29± 0.480.9972.34± 1.0622.8089.60± 0.3712.6262.21± 0.2998.31 GyroSPD++ ∗ [161]95.20± 0.885.0969.82± 1.79103.5789.50± 0.3766.3561.57± 0.30216.46 Correlation CorNet-ECM97.71± 0.611.0181.35± 1.270.6092.17± 0.490.5065.04± 0.1412.06 CorNet-LECM98.40± 0.701.1278.05± 1.140.6491.17± 0.320.5465.03± 0.1012.68 CorNet-OLM97.57± 0.761.3581.46± 0.610.9391.63± 0.120.7964.41± 0.2316.07 CorNet-LSM96.24± 1.481.5074.89± 1.070.9883.43± 0.650.8360.69± 0.8516.28 CorNet-PHCM96.56± 0.862.3782.26± 0.921.1090.03± 0.630.7760.01± 0.2216.92 Table 5.20: Five-fold results and training time per epoch on four data sets. The top 3 results are highlighted with red, blue, and cyan. ∗ denotes reproduced results due to missing official code. GyroGr-Scaling [159], SPDNet [106], SPDNetBN [31], RResNet [118], LieBN [50], SPD MLR [51], Gyro [159], and GyroSPD++[161]. 5.4.5.1 Main Results Tab. 5.20 reports the five-fold results comparing our CorNets against existing SPD and Grassmannian baselines. We summarize the key observations below. • Effectiveness. CorNets consistently outperform both SPD and Grassmannian networks. Specifically, CorNets surpass the classic SPDNet by 5.15%, 17.69%, 6.58%, and 13.79% on four data sets, respectively, and outperform the best Grassmannian networks by 7.76%, 19.07%, 6.86%, and 7.45%. Despite not us- ing BN or residual blocks, CorNets achieve superior performance compared to SPDNetBN, SPDNetLieBN, and RResNet. Notably, although CorNets share the same high-level architecture as GyroSPD++ (one manifold convolutional layer fol- lowed by one manifold MLR layer), CorNets exhibit better performance. These results highlight the effectiveness of correlation embedding and our method for constructing correlation networks. • Optimal Metric. The optimal metric for CorNets varies across data sets, in- dicating that the choice of geometry is a critical hyperparameter in Riemannian networks. Our framework enables seamless switching among five correlation ge- ometries in a consistent architecture, demonstrating the adaptability of our ap- proach to different tasks. 173 5.4. Full-Rank Correlation Networks Data SetHDM05FPHA Conv MLR ECMLECMOLMLSMPHCMECMLECMOLMLSMPHCM ECM81.35± 1.2773.38± 0.3480.11± 0.7778.54± 0.4380.80± 0.5492.17± 0.4991.50± 0.2191.67± 0.2887.37± 1.1491.97± 0.24 LECM66.49± 1.1378.05± 1.14 79.21± 1.2373.61± 0.9958.37± 2.2487.90± 0.5791.17± 0.3290.25± 0.2589.63± 0.3186.09± 0.98 OLM 77.82± 0.4876.56± 0.8981.46± 0.6180.77± 0.8177.39± 1.2992.17± 0.5892.27± 0.7891.63± 0.1289.90± 0.6791.83± 0.15 LSM68.83± 1.1970.41± 1.5767.56± 1.5274.89± 1.0772.69± 3.5678.97± 2.8075.10± 1.1582.25± 3.3883.43± 0.6578.97± 4.97 PHCM81.16± 0.4080.05± 0.4581.96± 0.5178.28± 0.6482.26± 0.9288.30± 0.8179.80± 0.6987.37± 0.7286.63± 0.2790.03± 0.63 Table 5.21: Ablations on mixed geometries. Each row shows the metric used for Convo- lution (Conv), and each column is the metric for MLR. Thediagonal entries indicate configurations where both layers use the same metric. The best result in each row is bold. • Efficiency. CorNets achieve efficiency comparable to or better than several baseline methods. The most efficient CorNet variant is based on ECM, owing to the simplest computations of ECM. Although GyroSPD++ uses the same architecture, CorNets achieve significantly greater efficiency, attributed to the heavy computational cost of the AIM-based computations in GyroSPD++ and the lightweight Riemannian computations on the correlation manifold. Particu- larly, on the largest NTU120 data set, CorNet-ECM and CorNet-LECM are the top two most efficient ones. 5.4.5.2 Ablations on Mixed Geometries Our main experiments use the same metric for convolution and MLR. To evaluate mixed geometries, we assign different metrics to the two layers. Tab. 5.21 reports five- fold results on HDM05 and FPHA. Overall, consistent metrics yield the best accuracy. 5.4.5.3 Visualization Figure 5.5: Illustration of the decision hyperplanes in the correlation MLRs under five different geometries. The 3 × 3 correlation manifold can be embedded as an open elliptope in R 3 , by visualizing the strictly lower triangular part of each C ∈ Cor + (3). The black dots denote the boundary. The PHCM hyperplane is defined by the one in the β-concatenated Poincaré space. Fig. 5.5 shows that different metrics induce visibly distinct curved hyperplanes. 174 Chapter 5. Riemannian Neural Networks 5.4.5.4 Potential and Necessity InputRadarHDM05FPHA SPD93.25± 1.1064.57± 0.6185.59± 0.72 Correlation89.49± 0.6766.81± 0.7383.37± 0.40 Table 5.22: SPDNet: SPD vs. correlation. Although correlation matrices are still SPD, naively treating them as SPD inputs and feeding them into existing SPD networks fails to leverage their intrinsic ge- ometric structures. To illustrate this, we use the classic SPDNet [106] but replace its covariance inputs with correla- tion matrices. The five-fold average results in Tab. 5.22 reveal two key insights: (1) on the HDM05 data set, correlation inputs lead to improved performance, suggesting that correlation embeddings can serve as compact and effective alternatives to covari- ance representations; and (2) on the other two data sets, the performance degrades, indicating that ignoring the specific geometry of correlation matrices can be detrimen- tal. These findings highlight both the promise and the necessity of designing networks respecting the unique geometry of the correlation manifold. 5.4.5.5 Ablations on Correlation Embeddings Data SetMeasurement SPDMLR-TrivlzCorMLR LEMLCMAIMECMLECMOLMLSMPHCM Radar Acc95.47± 0.66 95.55± 0.3594.87± 0.8789.47± 0.9387.41± 0.2385.79± 0.8391.63± 0.3283.33± 1.29 Fit Time (s/epoch)0.650.630.990.560.620.780.680.74 HDM05 Acc54.31± 1.6545.12± 1.0552.46± 2.4465.57± 0.6264.44± 0.6362.86± 0.6564.01± 0.9262.78± 0.85 Fit Time (s/epoch)3.245.38260.673.183.873.393.572.73 FPHA Acc84.13± 1.1476.62± 0.4383.25± 0.5985.37± 0.1685.24± 0.2284.67± 0.2780.17± 0.1573.67± 0.32 Fit Time (s/epoch)0.510.5218.960.510.640.80.810.45 Table 5.23: Comparison of SPDMLR-Trivlz on raw covariances against CorMLR on raw correlations on all three data sets. The input matrix dimensions are 93× 93, 63× 63, and 20× 20, respectively. To further evaluate the effectiveness of correlation embeddings, we compare the perfor- mance of directly classifying raw covariance matrices using the SPD MLRs in Thm. 117 with that of classifying corresponding raw correlation matrices using correlation MLR (CorMLR). The original SPDMLR involves an SPD matrix parameter for each class, which causes heavy Riemannian computations. For a fair comparison, we also imple- ment a similar trivialization as Sec. 5.4.2.1, denoted as SPDMLR-Trivlz. We implement SPDMLR-Trivlz under LEM, LCM, and AIM, respectively. Tab. 5.23 presents the 5-fold average results on all three data sets. CorMLR performs better than SPDMLR-Trivlz on HDM05 and FPHA. Although CorMLR performs worse on Radar, we emphasize that 175 5.4. Full-Rank Correlation Networks these comparisons are conducted on a single MLR layer, which fails to fully uncover the potential of correlation matrices. Besides, SPDMLR under AIM is much slower than others, especially on HDM05, due to its complex computations. In contrast, CorMLR, especially under ECM and PHCM, offers competitive or superior efficiency. 5.4.5.6 Analysis of Covariance versus Correlation In this section, we analyze when and why correlation matrices provide stronger repre- sentations than covariance matrices. The coefficient of variation of diagonal variances quantifies the variability of diagonal variances via per-sample coefficients of variation, and the ratio of diagonal to off-diagonal entries compares the magnitudes of diago- nal and off-diagonal entries via their ratios. These analyses lead to two insights: (1) large variability and magnitude of diagonal elements can act as nuisance noise for SPD networks by overshadowing informative off-diagonal correlations; (2) under such cases, correlation representations that normalize variances and emphasize pairwise correla- tions tend to be more effective, which is especially evident on HDM05. Coefficient of Variation of Diagonal Variances. 0.60.81.01.21.41.61.82.0 Coefficient of Variation 0 10 20 30 40 50 60 Count Channel 0 0.751.001.251.501.752.002.252.50 Coefficient of Variation Count Channel 1 1.01.52.02.5 Coefficient of Variation Count Channel 2 0.60.81.01.21.41.61.82.0 Coefficient of Variation 0 10 20 30 40 50 60 Count Channel 3 0.81.01.21.41.61.82.02.2 Coefficient of Variation Count Channel 4 0.751.001.251.501.752.002.252.50 Coefficient of Variation Count Channel 5 0.751.001.251.501.752.002.25 Coefficient of Variation 0 10 20 30 40 50 60 Count Channel 6 0.751.001.251.501.752.002.25 Coefficient of Variation Count Channel 7 1.01.52.02.5 Coefficient of Variation Count Channel 8 Figure 5.6: Distribution of per-sample coefficients of variation of diagonal variances on FPHA. Higher values indicate stronger diagonal variability, which could cause nuisance noise. 176 Chapter 5. Riemannian Neural Networks 1.01.52.02.53.0 Coefficient of Variation 0 20 40 60 80 Count Channel 0 1.01.52.02.53.0 Coefficient of Variation Count Channel 1 1.01.52.02.53.0 Coefficient of Variation Count Channel 2 Figure 5.7: Distribution of per-sample coefficients of variation of diagonal variances on HDM05. Higher values indicate stronger diagonal variability, which could cause nuisance noise. This section investigates why CorNets yield substantially larger gains over SPD networks on HDM05 compared to FPHA. Setup. For each covariance matrix Σ∈S n ++ we extract the diagonal vector v = (Σ 11 ,..., Σ n ).(5.82) We compute the coefficient of variation of v as CV = std(v) mean(v) + ε ,(5.83) where ε = 10 −8 ensures numerical stability. As shown in Sec. A.3.6.1, each sequence is modeled as a c-channel tensor of covariance matrices. The above procedure yields one coefficient of variation per channel for each sample. We visualize their empirical distributions per channel. Analysis. Figs. 5.6 and 5.7 show that the coefficients of variation w.r.t. diagonal variance are large on both data sets. On FPHA, most values fall between 0.8 and 2.0. On HDM05, they are even larger, typically between 1.0 and 3.0. Such large fluctuations indicate that diagonal variances change substantially and could bring nuisance noise for SPD networks. In contrast, correlation matrices allow CorNets to focus on pairwise relationships. This explains the consistent improvements over SPD networks and the larger gains on HDM05. Ratio of Diagonal to Off-Diagonal Entries in Covariance Features. 177 5.4. Full-Rank Correlation Networks 1.52.02.53.03.54.0 Ratio of diagonal to off-diagonal entries 0 10 20 30 40 50 60 70 Count Channel 0 1.52.02.53.03.54.0 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 1 1.52.02.53.03.5 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 2 1.52.02.53.03.54.0 Ratio of diagonal to off-diagonal entries 0 10 20 30 40 50 60 70 Count Channel 3 1.52.02.53.03.5 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 4 1.52.02.53.03.5 Ratio of diagonal to off-diagonal entries 0 10 20 30 40 50 60 70 Count Channel 5 1.52.02.53.03.54.04.5 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 6 1.52.02.53.03.54.0 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 7 1.52.02.53.03.54.0 Ratio of diagonal to off-diagonal entries 0 20 40 60 80 Count Channel 8 Figure 5.8: Distribution of ratios of diagonal to off-diagonal entries on FPHA. 234567 Ratio of diagonal to off-diagonal entries 0 25 50 75 100 125 150 Count Channel 0 234567 Ratio of diagonal to off-diagonal entries 0 50 100 150 Count Channel 1 23456 Ratio of diagonal to off-diagonal entries 0 25 50 75 100 125 150 Count Channel 2 Figure 5.9: Distribution of ratios of diagonal to off-diagonal entries on HDM05. This section further examines why CorNets achieve larger gains over SPD networks on HDM05 than on FPHA. We analyze the ratio of diagonal to off-diagonal entries in covariance matrices on FPHA and HDM05, to quantify how strongly variance terms overshadow pairwise correlations. Setup. For each covariance matrix Σ ∈ S n ++ we compute the mean magnitude of diagonal entries D = 1 n n X i=1 |Σ i |,(5.84) 178 Chapter 5. Riemannian Neural Networks and the mean magnitude of off-diagonal entries O = 1 n(n− 1) X i̸=j |Σ ij |.(5.85) We then form the sample-wise ratio R = D O ,(5.86) which measures how much larger the diagonal amplitudes are compared to the off- diagonal correlations. Each sample yields one ratio per channel, and we visualize the empirical distributions of these ratios on FPHA and HDM05. Analysis. Figs. 5.8 and 5.9 show that both data sets have ratios well above one. On FPHA, most ratios lie between 1.7 and 3.0, indicating that diagonal amplitudes are noticeably larger than off-diagonal correlations. HDM05 exhibits even larger ra- tios, typically between 2.0 and 6.0, with many above 3.0. These statistics indicate that covariance representations on both data sets are strongly dominated by diagonal en- tries, with more pronounced dominance on HDM05. When diagonal terms dominate, SPD networks trained on covariance inputs tend to overemphasize variances and un- derexploit informative pairwise correlations. Correlation matrices normalize variances and highlight off-diagonal interactions, which explains why CorNets outperform SPD baselines on both data sets and why the improvement is substantially larger on HDM05. 5.4.5.7 Normalized Covariance vs. Correlation Setup. We evaluate SPD-based baselines by covariance inputs normalized by their largest eigenvalue. Given a covariance matrix Σ, we get the normalized SPD input b Σ = Σ/λ max (Σ) and feed it into existing SPD networks. This variant is denoted by “-EigN”. We report results on the Radar, HDM05, and FPHA data sets for representative SPD models: SPDNet, SPDNetBN, SPDResNet, SPDNetLieBN, SPDNetMLR, GyroAI, and GyroSPD++. Here, SPDResNet is implemented under the LEM, while SPDNetLieBN follows the LCM. 179 5.4. Full-Rank Correlation Networks ManifoldMethodRadarHDM05FPHA S n ++ SPDNet93.25± 1.1064.57± 0.6185.59± 0.72 SPDNet-EigN86.91± 0.5766.62± 0.7384.90± 0.62 SPDNetBN94.85± 0.9971.28± 0.7989.33± 0.49 SPDNetBN-EigN89.25± 1.1971.59± 0.6888.47± 0.39 SPDResNet95.89± 0.8670.12± 2.4585.07± 0.99 SPDResNet-EigN92.61± 0.9671.02± 0.9184.53± 0.46 SPDNetLieBN94.80± 0.7171.78± 0.4486.33± 0.43 SPDNetLieBN-EigN88.91± 1.2170.61± 1.0483.73± 0.65 SPDNetMLR94.59± 0.8265.90± 0.9385.60± 0.43 SPDNetMLR-EigN89.41± 0.5866.89± 0.6383.63± 1.09 GyroAI96.29± 0.4872.34± 1.0689.60± 0.37 GyroAI-EigN91.36± 0.8072.64± 0.7089.90± 0.31 GyroSPD++95.20± 0.8869.82± 1.7989.50± 0.37 GyroSPD++-EigN90.83± 1.0966.92± 0.2884.29± 0.14 Cor + (n) CorNet-ECM97.71± 0.6181.35± 1.2792.17± 0.49 CorNet-LECM98.40± 0.7078.05± 1.1491.17± 0.32 CorNet-OLM97.57± 0.7681.46± 0.6191.63± 0.12 CorNet-LSM96.24± 1.4874.89± 1.0783.43± 0.65 CorNet-PHCM96.56± 0.8682.26± 0.9290.03± 0.63 Table 5.24: SPD networks with or without normalized SPD inputs. Results. Tab. 5.24 summarizes the results. On HDM05, eigenvalue normalization has only a marginal effect and the normalized variants achieve accuracy comparable to their unnormalized counterparts. On FPHA and, in particular, on Radar, normaliza- tion usually reduces accuracy. The behavior of GyroSPD++ is especially informative. GyroSPD++ and CorNet share a similar architecture, consisting of one convolution followed by an MLR layer. However, GyroSPD++-EigN performs worse than Gy- roSPD++ on all three data sets, while CorNet with correlation inputs achieves clear improvements over GyroSPD++. These phenomena can be explained by two factors. (1) Redundancy. The raw samples on HDM05 and FPHA have already undergone centering, scaling, and normalization before covariance modeling. Dividing by λ max (Σ) therefore introduces little additional control over scale, which explains the marginal effect on HDM05. (2) Scaled Covariance versus Correlation. Since EigN is equivalent to uniformly rescaling the raw samples before covariance computation, the normalized covari- ance matrices remain covariances and do not encode new statistical information. 180 Chapter 5. Riemannian Neural Networks Moreover, forcing the largest eigenvalue to 1 can remove potentially informative differences in overall energy across samples, which aligns with the degradation ob- served for EigN variants, especially GyroSPD++-EigN. In contrast, correlation normalization uses a different scaling factor for each pair of variables, Cor ij = Σ ij p Σ i Σ j ,(5.87) producing standardized correlation coefficients. Therefore, global eigenvalue scal- ing is statistically distinct from correlation normalization and fails to capture the benefits of explicit correlation modeling. 5.4.5.8 Ablations on Activations MetricActivationRadarHDM05FPHA ECM ReLU97.41± 0.2581.23± 0.4689.80± 0.58 None97.71± 0.6181.35± 1.2792.17± 0.49 LECM ReLU97.23± 0.6777.51± 1.0291.00± 0.15 None98.40± 0.7078.05± 1.1491.17± 0.32 OLM ReLU97.52± 0.4781.86± 0.6591.47± 0.19 None97.57± 0.7681.46± 0.6191.63± 0.12 LSM ReLU95.60± 0.97N/AN/A None96.24± 1.4874.89± 1.0783.43± 0.65 PHCM ReLU96.40± 0.2577.32± 1.5688.63± 0.22 None96.56± 0.8682.26± 0.9290.03± 0.63 Table 5.25: Comparison of CorNet with or without activations. In the main experiments, we follow HNN++ [181] and GyroSPD++ [161], and do not use explicit activations, as the manifold itself introduces nonlinearity. We further conduct an ablation on activations. Following Ganea et al. [76, Sec. 3.2], we define acti- vations in the tangent space at the identity, i.e., Exp I ◦δ◦ Log I for four Log-Euclidean metrics, and Exp 0 ◦δ◦ Log 0 for PHCM in the β-concatenated Poincaré vector, where δ is ReLU [82]. Specifically, we insert a ReLU after the correlation convolution. As shown in Tab. 5.25, adding activations generally yields no benefits and can even degrade performance. The variant without activation consistently achieves higher or compara- ble accuracy, except CorNet-OLM for HDM05. Moreover, CorNet-LSM with activation 181 5.4. Full-Rank Correlation Networks fails to converge on HDM05 and FPHA. These results suggest that CorNet already provides sufficient nonlinearity, rendering additional activations redundant. 5.4.5.9 Scalability of Correlation Metrics DimECMLECM OLM LSM PHCM 300.00040.00180.00120.00190.0131 500.00040.00270.03180.03340.0211 1000.00080.00540.07640.07810.0413 1500.00150.01000.12470.12670.2284 2000.00250.01970.19060.19380.3320 2500.00370.03450.23520.23790.4414 300 0.00530.07330.34340.34540.5732 4000.00920.17960.51630.52610.4807 5000.01430.30760.69070.69610.5693 6000.02060.59830.93310.94840.7923 700 0.02891.09611.24321.25751.0417 8000.0391.86891.66581.68151.3387 900 0.05352.98862.21562.23031.7324 10000.07063.72592.5392.57831.229 Table 5.26: Average runtime (s) of a single forward pass in CorNet under different metrics and input dimensions. The best results are bold. We evaluate the computational efficiency of correlation metrics across increasing input dimensions using CorNet with one correlation FC layer followed by one correlation MLR layer. Each input correlation matrix of size [n,n] is mapped to [20, 20] by the FC layer and then classified into 10 classes by the MLR layer. For each listed dimension 30 ≤ n ≤ 1000, we randomly generate 30 correlation matrices and record the average runtime of a single forward pass. As implied by Tab. 2.7, the runtime is governed by two factors: the codomain computation (Euclidean or hyperbolic) and the complexity of the diffeomorphism. The results are summarized in Tab. 5.26. We have the following findings. • ECM is consistently the most efficient metric, benefiting from both a Euclidean codomain and the simplest diffeomorphism. 182 Chapter 5. Riemannian Neural Networks • At very low dimensions (n≤ 100), the relative costs of the non-ECM metrics are not yet stable. At n = 50 and n = 100, the ordering is ECM < LECM < PHCM < OLM≈ LSM.(5.88) At the smallest tested dimension n = 30, all runtimes remain small and their rel- ative ordering differs. At intermediate dimensions (150≤ n≤ 300), the ordering becomes ECM < LECM < OLM≈ LSM < PHCM.(5.89) Here, the cost of PHCM’s hyperbolic computations dominates, while the dimension- dependent cost of LECM’s matrix functions is not yet pronounced. • From n = 400, PHCM becomes faster than OLM and LSM, and at n = 700, it also becomes faster than LECM. At high dimensions (n≥ 800), the ordering is ECM < PHCM < OLM≈ LSM < LECM.(5.90) Here, diffeomorphisms dominate: ECM and PHCM scale better thanks to rela- tively lightweight Cholesky decomposition, while OLM and LSM slow down due to matrix logarithm/exponentiation. LECM is the slowest, as its log◦Θ requires two nested matrix functions. 5.5 Conclusion This chapter developed manifold-specific neural components and architectures by ex- ploiting the additional structures of particular hyperbolic models and correlation man- ifolds. The first part addressed representation choice through PVNN. The unconstrained PV model is connected to the Poincaré and hyperboloid models by Riemannian isome- tries, but it avoids their explicit constraints. By deriving closed-form core Rieman- nian operators and relating them to the PV gyrovector structure, we constructed PV MLR, FC, convolutional, activation, and normalization layers. Together, these layers constitute a complete PVNN framework for constructing model-specific hyperbolic ar- chitectures. Experiments demonstrate the superior numerical stability and competitive performance of the PV representation. The second part shifted attention from the representation itself to the geometric principle for building neural layers. Using Busemann functions and horospheres, we developed a common construction for the Poincaré and Lorentz models. BMLR ad- 183 5.5. Conclusion mits an exact point-to-horosphere interpretation, avoids an additional manifold-valued point parameter, supports batch-efficient evaluation, and recovers Euclidean MLR in the zero-curvature limit. The same Busemann logits yield explicit BFC layers with practical O(nm) complexity and corresponding Euclidean limits. Experiments on im- age classification, genome sequence learning, node classification, and link prediction showed that these layers generally improve upon existing hyperbolic alternatives while retaining comparable computational cost. The third part extended manifold-specific network design from hyperbolic vectors to full-rank correlation matrices. The four flat geometries, ECM, LECM, OLM, and LSM, admit Euclidean isometries that yield closed-form correlation MLR, FC, and convolu- tional layers. For the non-flat PHCM geometry, the Cholesky representation identified correlation matrices with a product of Poincaré balls, enabling analogous layers through β-concatenation and β-splitting. Analytic gradients for the Riemannian computations enabled accurate end-to-end backpropagation under OLM and LSM. Across the four evaluated data sets, the best CorNet geometry outperformed the considered SPD and Grassmannian baselines. All three designs nevertheless operate under prescribed Riemannian geometries. The next chapter therefore turns from the design of modules and architectures to that of the underlying metrics themselves. 184 Chapter 6 Fast and Stable Geometries on SPD Manifolds 6.1 Introduction The preceding chapters developed Riemannian neural components and architectures under prescribed Riemannian metrics. Despite their different constructions, they share the underlying metric as a common starting point. As reviewed in Chapter 2, a Riemannian metric assigns inner products to tangent spaces and determines distance, geodesics, exponential and logarithmic maps, and par- allel transport. The induced distance also defines the Fréchet statistics used to summa- rize manifold-valued features. When combined with compatible Lie group or gyrovector structures, a Riemannian metric further supports manifold analogues of addition and scalar multiplication. These geometric primitives offer powerful toolkits for building Riemannian neural networks. Based on this observation, we shift our attention from designing neural modules and architectures under prescribed geometries to designing the underlying geometry itself. Our goal is to balance geometric flexibility, theoretical convenience, computational ef- ficiency, and numerical stability. Accordingly, we seek geometries whose Riemannian operators and compatible algebraic operations admit closed-form expressions and can be inserted directly into deep learning architectures. We focus on the SPD manifold, which encodes covariance or other second-order statistics and arises naturally in many applications [37, 31, 134, 61, 152, 141, 231, 106]. 1 1 Our focus is distinct from SPD metric-learning methods that learn distance functions induced by an existing Riemannian metric. Instead, we seek to design new Riemannian metrics themselves. 185 6.2. Adaptive Log-Euclidean Metrics We pursue two complementary routes. Sec. 6.2 starts from pullback Euclidean geom- etry and proposes Adaptive Log-Euclidean Metrics (ALEMs). The resulting adaptive geometry retains closed-form Riemannian operators and a convenient abelian Lie group structure. Sec. 6.3 instead exposes the product structure of the Cholesky manifold and transfers these geometries to the SPD manifold, yielding simple and stable closed-form Riemannian and gyro operators. 6.2 Adaptive Log-Euclidean Metrics 6.2.1 Introduction As reviewed in Sec. 2.9.1, most popular Riemannian metrics on the SPD manifold are fixed, which can limit the expressive capacity of the associated geometry. A com- mon way to construct SPD metrics is through pullbacks along diffeomorphisms, which transfer Riemannian structures from simpler source manifolds. For instance, Thanwer- das and Pennec [195] explained AIM as the pullback metric from a left-invariant metric on the Cholesky manifold. The matrix-power deformations in Secs. 3.2.5.1 and 4.3.3.1 are also representative examples. Inspired by the above observations, we leverage pullback techniques to introduce adaptive Riemannian metrics. In particular, we first show that several Riemannian metrics on SPD manifolds, including LEM, LCM, and their generalizations, can be explained as pullback metrics from the standard Euclidean space. We refer to these metrics as pullback Euclidean metrics. Then, we propose a general framework for char- acterizing the properties of pullback Euclidean metrics. Our framework can explain the widely used LEM [9] and LCM [137]. We focus on LEM on SPD manifolds and extend it into Adaptive Log-Euclidean Metrics (ALEMs). Besides, we present a com- plete theoretical discussion on the properties of ALEMs. Compared with the existing Riemannian metrics, our metrics are learnable, adapting to the characteristics of the data sets. The effectiveness of our metrics is demonstrated by experiments as well as the applications to recently developed Riemannian building blocks, including the RBN framework developed in Chapter 3, Riemannian residual blocks [117], and the Riemannian classifiers developed in Chapter 4. Drawing on this, our contributions are summarized as follows: (1) We reveal the connection of two popular Riemannian metrics (LEM and LCM) by the pullback technique and propose a general framework for pullback Euclidean metrics. 186 Chapter 6. Fast and Stable Geometries on SPD Manifolds (2) Based on our framework, we propose specific ALEMs on SPD manifolds and con- duct comprehensive analyses in terms of the algebraic, analytic, and geometric properties. (3) Extensive experiments on widely used SPD learning benchmarks demonstrate that our metrics exhibit consistent performance gain across data sets. 2 Outline. The necessary background on differential geometry, pullback metrics, and SPD geometry has been established in Chapter 2. Sec. 6.2.2 develops ALEM and its differentials, and Sec. 6.2.2.5 studies its geometric properties. Sec. 6.2.3 derives the gradients and parameter-update rules for the general matrix logarithm and exponential. Sec. 6.2.4 instantiates the general matrix logarithm as ALog in SPDNet and applies ALEM to other Riemannian building blocks. Proofs are deferred to Sec. B.9. 6.2.2 Adaptive Log-Euclidean Metrics In this section, we show that both (α,β)-LEM and LCM are pullback metrics from the Euclidean space. Inspired by this observation, we present a general framework for characterizing pullback Euclidean metrics. Then, we focus on generalizing LEM. 6.2.2.1 Rethinking (α,β)-LEM and LCM Among the existing Riemannian metrics on the SPD manifold, LEM is popular in many applications, given its closed form for the Fréchet mean and clear vector-space and Lie- group structures. In addition, the nascent LCM, which is gaining increasing attention, shares similar properties with LEM. LEM is derived from Lie group translation [9], while LCM is obtained as the pullback of a metric on L n ++ [137]. Besides, (α,β)-LEM is obtained as a pullback of LEM [196]. However, the same mathematical logic underlies their derivations. We denote the Euclidean space of n×n lower triangular matrices by LT n . We define ψ LC :S n ++ → LT n as ψ LC (P ) =⌊L⌋ + Dlog(D(L)),(6.1) where L is the Cholesky factor of the SPD matrix P, ⌊L⌋ is the strictly lower part of L, D(L) is a diagonal matrix with diagonal elements of L, and Dlog applies the natural logarithm element-wise to the diagonal. Then, we have the following theorem. 2 The code is available at https://github.com/GitZH-Chen/ALEM. 187 6.2. Adaptive Log-Euclidean Metrics Theorem 147. [↓] (α,β)-LEM is the pullback metric from the Euclidean space of S n with an O(n)-invariant inner product ⟨·,·⟩ (a,b) by matrix logarithm. Specifically, the standard LEM is the pullback metric from the Euclidean space of S n with the standard Frobenius inner product by matrix logarithm. LCM is the pullback metric from LT n with the Frobenius inner product by ψ LC . As Euclidean spaces of the same dimension are naturally isometric, it follows that both (α,β)-LEM and LCM are pulled back from the standard Euclidean space S n . Corollary 148. [↓] (α,β)-LEM and LCM are pullback metrics from S n with the standard Frobenius inner product. 6.2.2.2 Pullback Euclidean Metrics on SPD Manifolds Sec. 6.2.2.1 has shown how LEM is derived from matrix logarithm. Besides, as shown in Arsigny et al. [9], operations in Lie group and linear space on S n ++ are also induced from matrix logarithm. Now, let us explain the underlying mechanism in detail. A matrix logarithm is a diffeomorphism (a smooth bijection with a smooth inverse). The property of bijection offers the possibility of transferring algebraic structures from S n intoS n ++ . The smoothness of the matrix logarithm and its inverse suggests that smooth structures, such as a Lie group structure and a Riemannian metric, can be transferred to S n ++ . More generally, given an arbitrary diffeomorphism φ :S n ++ →S n , it suffices to pull various properties from the Euclidean space back to the SPD manifoldS n ++ by φ as well. Besides, the computation of the induced operators in S n ++ by φ is usually simple. Lemma 149. [↓] Let S 1 ,S 2 ∈ S n ++ , V ∈ T S 1 S n ++ , and k ∈ R, and let g E be the Frobenius inner product in S n . Let φ :S n ++ →S n be a diffeomorphism, and denote its differential at S ∈S n ++ by φ ∗,S . We define the following operations: Element Addition: S 1 ⊙ φ S 2 = φ −1 (φ(S 1 ) + φ(S 2 )),(6.2) Scalar Multiplication: k⊛ φ S 2 = φ −1 (kφ(S 2 )),(6.3) Inner Product: ⟨S 1 ,S 2 ⟩ φ =⟨φ(S 1 ),φ(S 2 )⟩,(6.4) Riemannian Metric: g φ = φ ∗ g E ,(6.5) Then, we have the following conclusions: (1) S n ++ ,⊙ φ ,⊛ φ ,⟨·,·⟩ φ is a Hilbert space over R. 188 Chapter 6. Fast and Stable Geometries on SPD Manifolds (2) S n ++ ,⊙ φ is an abelian Lie group. S n ++ ,g φ is a Riemannian manifold. The associated Riemannian operators are as follows: d φ (S 1 ,S 2 ) =∥φ(S 1 )− φ(S 2 )∥ F ,(6.6) Exp S 1 V = φ −1 (φ(S 1 ) + φ ∗,S 1 V ),(6.7) Log S 1 S 2 = φ −1 ∗,φ(S 1 ) (φ(S 2 )− φ(S 1 )),(6.8) PT S 1 →S 2 (V ) = φ −1 ∗,φ(S 2 ) ◦ φ ∗,S 1 (V ),(6.9) where ∥·∥ F is the Frobenius norm, V ∈ T S 1 S n ++ is a tangent vector, Exp S 1 , Log S 1 , and PT S 1 →S 2 are the Riemannian exponential map at S 1 , logarithmic map at S 1 , and parallel transport along the geodesic connecting S 1 and S 2 , respectively, and φ −1 ∗ denotes the differential of φ −1 . Then g φ is a bi-invariant metric, called a Pullback Euclidean Metric induced by φ. (3) φ is an isomorphism: (a) a linear isomorphism preserving the inner product; (b) a Lie group isomorphism; (c) a Riemannian isometry. In fact, (α,β)-LEM and LCM are special cases of Thm. 149, as are the linear-space and Lie-group structures in Arsigny et al. [9] and the Lie-group structure in Lin [137]. In addition, neither Arsigny et al. [9] nor Lin [137] reveals the Hilbert space structures in S n ++ . 6.2.2.3 Adaptive Log-Euclidean Metrics The key to Thm. 149 lies in the diffeomorphism φ. If we have a proper φ, Rieman- nian metrics on SPD manifolds can be induced. In the following, we will present our mappings and then discuss the induced metrics. As reviewed in Sec. 2.8.1, the matrix logarithm reduces to a scalar logarithm, which is a diffeomorphism between R ++ and R. Following this hint, the eigenvalue-based diffeomorphism between S n ++ and S n reduces to a scalar diffeomorphism between R ++ and R. A very natural idea is to substitute the natural logarithm with logarithms with arbitrary proper bases. Throughout this part, log(·) without a subscript denotes the natural scalar or matrix logarithm. Every scalar or matrix logarithm with a general or adaptive base is written explicitly as log α (·), without omitting the subscript. The base parameter α is interpreted as either a scalar or a vector according to its argument. For 189 6.2. Adaptive Log-Euclidean Metrics a scalar base α∈ R ++ \1 and x∈ R ++ , we define log α (x) = log(x) log(α) .(6.10) When α = e, this scalar logarithm reduces to the natural logarithm, which we write without a subscript as log(·). For a diagonal matrix X, the same notation is extended to a base vector as log α (X) = diag(log a 1 (x 11 ), log a 2 (x 22 ),· , log a n (x n )),(6.11) where α = (a 1 ,a 2 ,· ,a n ) ∈ (R ++ \1) n is the base vector, diag(·) is the diagonal- ization operator, and X is an n× n diagonal matrix. When α is scalar in a matrix expression, it denotes the constant base vector (α,...,α). Together with eigendecom- position, a general matrix logarithm is defined by log α (S) = U log α (Σ)U ⊤ ,(6.12) where S = U ΣU ⊤ is the eigendecomposition. As a special case, α = (e,e,· ,e)=⇒log α = log.(6.13) As with the scalar logarithm, we have the following proposition. Proposition 150 (Diffeomorphism). [↓] log α is a diffeomorphism, a smooth bijec- tion with a smooth inverse log −1 α :S n →S n ++ defined as log −1 α (X) = U diag(a Σ 11 1 ,a Σ 22 2 ,· ,a Σ n n )U ⊤ ,(6.14) where X = U ΣU ⊤ is the eigendecomposition. Remark 151. The general matrix logarithm log α is an arbitrary member of the following family log α | α = (a 1 ,· ,a n )∈ (R ++ \1) n .(6.15) Besides, there could be some ambiguity in Eq. (6.12) under different arrangements of eigenvalues and eigenvectors. In fact, there is a correspondence between scalar 190 Chapter 6. Fast and Stable Geometries on SPD Manifolds log a i and eigenvalues and eigenvectors. See Sec. A.4.4.1 for more details. Since log α is a diffeomorphism from S n ++ onto S n , all the results in Thm. 149 hold true. Theorem 152. [↓] Following the notation in Thm. 149, we define ⊕ ALE and ⊙ ALE as in Eqs. (6.2) and (6.3). We define ⟨·,·⟩ log α and g log α as in Eqs. (6.4) and (6.5). Then, we have the following conclusions: (1) S n ++ ,⊕ ALE ,⊙ ALE ,⟨·,·⟩ log α is a Hilbert space over R. (2) S n ++ ,⊕ ALE is an abelian Lie group. g log α is a Riemannian metric on S n ++ . We call this metric the Adaptive Log-Euclidean Metric (ALEM) and denote g log α by g ALE . The associated Riemannian operators are as follows: d ALE (S 1 ,S 2 ) =∥ log α (S 1 )− log α (S 2 )∥ F ,(6.16) Exp S 1 V = log −1 α log α (S 1 ) + (log α ) ∗,S 1 V ,(6.17) Log S 1 S 2 = log −1 α ∗,X 1 (log α (S 2 )− log α (S 1 )),(6.18) PT S 1 →S 2 (V ) = log −1 α ∗,X 2 ◦ (log α ) ∗,S 1 (V ),(6.19) where X i = log α (S i )∈S n for i = 1, 2. (3) log α is an isomorphism: (a) a linear isomorphism preserving the inner product; (b) a Lie group isomorphism; (c) a Riemannian isometry. Remark 153. Obviously, ALEM varies with different base parameters α in log α . We thus use the plural to describe our metrics. Besides, our metrics can be learned. This is why we call them adaptive metrics. Analogously to (α,β)-LEM, we can also define (a,b)-ALEM as the pullback of an O(n)-invariant inner product: g (a,b)-ALE = log ∗ α g (a,b)-E ,(6.20) where we denote the O(n)-invariant inner product ⟨·,·⟩ (a,b) by g (a,b)-E . g (a,b)-ALE also shares the properties presented in Thm. 152. Nevertheless, this part focuses on (a,b) = (1, 0). 191 6.2. Adaptive Log-Euclidean Metrics 6.2.2.4 Differentials of General Logarithms Eqs. (6.17) to (6.19) require the differential maps of log α and log −1 α . This subsection introduces the concrete formulae of the associated differential maps. Proposition 154 (Differentials). [↓] For a tangent vector V ∈ T S S n ++ , the differ- ential (log α ) ∗,S : T S S n ++ → T log α (S) S n of log α at S ∈S n ++ is given by (log α ) ∗,S (V ) = Q + Q ⊤ + W,(6.21) where Q = D U log α (Σ)U ⊤ , D U = ( (σ 1 I n − S) + V u 1 · (σ n I n − S) + V u n ), W = U diag u ⊤ 1 V u 1 σ 1 log(a 1 ) ,· , u ⊤ n V u n σ n log(a n ) U ⊤ , (·) + is the Moore–Penrose inverse, u 1 ,· ,u n are orthonormal eigenvectors of S, and the associated eigenvalues are σ 1 ,· ,σ n . Symmetrically, for a tangent vector e V ∈ T X S n , the differential log −1 α ∗,X : T X S n → T log −1 α (X) S n ++ of log −1 α at X ∈S n is given by log −1 α ∗,X ( e V ) = e Q + e Q ⊤ + f W,(6.22) where X = e U e Σ e U ⊤ is the eigendecomposition. Here, D e U is defined similarly and e Q = D e U diag a eσ 1 1 ,· ,a eσ n n e U ⊤ . Moreover, f W = e U diag log(a 1 )a eσ 1 1 eu ⊤ 1 e V eu 1 ,· , log(a n )a eσ n n eu ⊤ n e V eu n e U ⊤ . Arsigny et al. [9] write the differential of the matrix exponential as an infinite series. The differential of log −1 α can also be rewritten in this way. Proposition 155 (Differential as Infinite Series). [↓] Following the notation in Thm. 154, the differential of log −1 α can also be formulated as log −1 α ∗,X ( e V ) = ∞ X k=1 1 k! ( k−1 X l=0 ( e PX) k−l−1 (D e P X + e P e V )( e PX) l ), (6.23) 192 Chapter 6. Fast and Stable Geometries on SPD Manifolds where e P = e UB e U ⊤ , B = diag (log(a 1 ),· , log(a n )), D e P = D e U B e U ⊤ + e UBD ⊤ e U . When log −1 α is reduced to the matrix exponential, Eq. (6.23) coincides with the expression in Arsigny et al. [9, Eq. (8)], and our ALEM becomes exactly LEM. 6.2.2.5 Properties of ALEM Since our ALEMs are natural generalizations of LEM, they intuitively share many of its properties. This subsubsection introduces some useful properties of our ALEMs for machine learning. Fréchet means are important tools for SPD matrix learning [94, 36, 31, 34]. Like LEM, our ALEM also admits closed-form expressions for Fréchet means. We present a more general result, the weighted Fréchet mean. Proposition 156 (Weighted Fréchet Means). [↓] For m points S 1 ,...,S m on the SPD manifold with associated weights w 1 ,...,w m ∈ R + satisfying P m i=1 w i > 0, the weighted Fréchet mean M over the metric space S n ++ ,d ALE has a closed form M = log −1 α m X i=1 w i P m j=1 w j log α (S i ) ! .(6.24) Like LEM, although ALEM is not affine-invariant, it enjoys several other invariance properties. Proposition 157 (Bi-invariance). [↓] ALEM is a Lie group bi-invariant metric. Proposition 158 (Exponential Invariance). [↓] The Fréchet means under ALEM are exponential-invariant. In other words, for S 1 ,...,S m ∈S n ++ and β ∈ R, (FM(S 1 ,...,S m )) β = FM(S β 1 ,...,S β m ),(6.25) where FM(S 1 ,· ,S m ) denotes the Fréchet mean of S 1 ,· ,S m . In addition to exponential invariance, the Fréchet mean induced by our ALEM also satisfies various properties presented in Ando et al. [5]. Proposition 159. [↓] Let the following be SPD matrices: A,B,C,A 0 ,B 0 ,C 0 ∈S n ++ .(6.26) 193 6.2. Adaptive Log-Euclidean Metrics Let FM(A,B,C) denote the Fréchet mean of A,B,C under ALEM. Then the Fréchet mean satisfies the following properties. (1) U1: permutation invariance. For any permutation π of A,B,C, FM π(A,B,C) = FM(A,B,C).(6.27) (2) U2. FM(A,A,A) = A. The following properties hold if A,B,C,A 0 ,B 0 ,C 0 commute. (1) V1: joint homogeneity. FM(aA,bB,cC) = (abc) 1/3 FM(A,B,C), ∀a,b,c > 0.(6.28) (2) V2: monotonicity. The map (A,B,C) 7→ FM(A,B,C) is monotone. Specifi- cally, if A ≥ A 0 , B ≥ B 0 , and C ≥ C 0 , then FM(A,B,C) ≥ FM(A 0 ,B 0 ,C 0 ) in the positive semidefinite ordering. (3) V3: self-duality. FM(A,B,C) = FM(A −1 ,B −1 ,C −1 ) −1 . (4) V4: determinant identity. det FM(A,B,C) = (detA· detB· detC) 1/3 . In fact, Thm. 159 holds true for any finite number of SPD matrices. Besides, the geodesic distance induced by ALEMs has similarity invariance. Proposition 160 (Similarity Invariance). [↓] The geodesic distance under ALEM is similarity invariant. In other words, let R∈ SO(n) be a rotation matrix and let s∈ R ++ be a scale factor. Given any two SPD matrices S 1 and S 2 , we have d ALE (S 1 ,S 2 ) = d ALE (s 2 RS 1 R ⊤ ,s 2 RS 2 R ⊤ ).(6.29) Let us explain a bit more about the above three kinds of invariance. First, among metrics on Lie groups, bi-invariant metrics are the most convenient ones [187, Ch. V]. Second, exponential invariance offers a fast computation for Fréchet means under ex- ponential scaling. Finally, similarity invariance is significant for describing frequently encountered covariance matrices [9]. The above discussion focuses on the theoretical perspective. Now, let us reconsider Eq. (6.12) in a numerical way. 194 Chapter 6. Fast and Stable Geometries on SPD Manifolds Proposition 161. [↓] log α can be rewritten as log α (S) = U log α (Σ)U ⊤ ,(6.30) = UA log(Σ)U ⊤ ,(6.31) = U log(Σ) B U ⊤ ,(6.32) where X Y is the diagonal division, B = diag (log(a 1 ),· , log(a n )), and A = I n B . Based on the above proposition, more analyses could be carried out from a numerical point of view. First, log α (·) can balance the eigenvalues of an input SPD matrix S by exploiting different bases for different eigenvalues. In Riemannian algorithms, manifold- valued features usually contain vibrant information. We expect that by the above adaptation, manifold-valued data could be better fitted and the learning ability of algorithms could be further promoted. Remark 162. Note that the discussion in Sec. 6.2.2.3 and Sec. 6.2.2.5 can also be readily transferred to LCM, generating an adaptive version of LCM. 6.2.3 Parameter Learning As shown in Thm. 152, Riemannian computations under ALEM are built upon the general matrix logarithm log α and its inverse. Accordingly, this subsection studies the backpropagation and parameter optimization of these two maps. 6.2.3.1 Gradient Computation For both the general matrix logarithm and exponential, gradients are required with respect to their parameters and inputs. Since log α involves a structured matrix decom- position, the following derivations rely heavily on structured-matrix backpropagation (BP) [110], whose key idea is invariance of the first-order differential form. The general matrix logarithm is a special case of eigenvalue functions. Based on the formula given by Bhatia [20] and the matrix BP techniques presented by Ionescu et al. [110], we can obtain all the gradients in the following propositions. Proposition 163. [↓] Let us denote X = log α (S), where S ∈ S d ++ is the input SPD matrix. The input gradient ∇ S L is obtained by specializing the Daleckii– Krein expression in Eqs. (2.90) and (2.91) with V =∇ X L and f (σ i ) = A i log(σ i ). 195 6.2. Adaptive Log-Euclidean Metrics NameDetailConstraintMethod RELUOptimizing base vector α (Eq. (6.30))Positiveshift-ReLU max(ε,α) MULOptimizing diagonal elements of A (Eq. (6.31))UnconstrainedStandard BP DIVOptimizing diagonal elements of B (Eq. (6.32))UnconstrainedStandard BP Table 6.1: Parameter learning for the general matrix logarithm and exponential. The parameter gradient is ∇ A L = [U ⊤ (∇ X L)U ]⊛ log(Σ),(6.33) where S = U ΣU ⊤ is the eigendecomposition of an SPD matrix and σ 1 ,...,σ d are the diagonal entries of Σ. As the inverse map corresponding to Eq. (6.31), the general matrix exponential log −1 α can be rewritten as log −1 α (X) = U diag a Σ 11 1 ,· ,a Σ n n U ⊤ = U exp Σ A U ⊤ , (6.34) where X = U ΣU ⊤ ∈S n is an eigendecomposition of X. Following Thm. 163, we obtain the backpropagation of log −1 α . Proposition 164. [↓] Let us denote X = log −1 α (S) with S ∈ S d . The input gra- dient ∇ S L is obtained by specializing the Daleckii–Krein expression in Eqs. (2.90) and (2.91) with V =∇ X L and f (σ i ) = e σ i A i . The parameter gradient is ∇ A L = [U ⊤ (∇ X L)U ]⊛ diag a Σ 11 1 ,· ,a Σ d d −Σ A 2 ,(6.35) where S = U ΣU ⊤ is the eigendecomposition of a symmetric matrix and σ 1 ,...,σ d are the diagonal entries of Σ. 6.2.3.2 Parameter Updates The general matrix logarithm and exponential share the same base parameters. Let us focus on the former. Let the input SPD matrix S have dimension d× d. Recalling Eqs. (6.30) to (6.32), there are three ways to implement parameter learning. We could learn the base vector α in Eq. (6.30), diagonal matrix A in Eq. (6.31), or diagonal 196 Chapter 6. Fast and Stable Geometries on SPD Manifolds matrix B in Eq. (6.32), respectively. For learning A in Eq. (6.31) or B in Eq. (6.32), since the parameters (diagonal elements) lie in a Euclidean space R d , the optimization can be easily integrated into the BP algorithm. We call learning A MUL and learning B DIV. For learning α in Eq. (6.30), each element a of α satisfies a > 0 and a ̸= 1. Since the equality case can be avoided by setting a = 1 +ε whenever a = 1, where ε∈ R ++ , it remains to enforce positivity during optimization. We consider two strategies for doing so. The first strategy applies the shift-ReLU max(ε,a) to an unconstrained parameter. We call this strategy RELU. Other transformations, such as squaring the parameter, are also feasible, but we focus on RELU. The second strategy takes a geometric approach by viewing a as a point on a one-dimensional SPD manifold and optimizing it using the Riemannian optimization strategy reviewed in Sec. 2.7. We call this strategy GEOM. Its RSGD update is given in the following proposition. Proposition 165. [↓] Viewing a positive scalar a as a point in a one-dimensional SPD manifold, we have the following RSGD update formula. a (t+1) = a (t) e −γ (t) a (t) ∇ a (t) L ,(6.36) where ∇ a (t) L is the Euclidean gradient of L with respect to a at a (t) , γ (t) is the learning rate, and e (·) is the natural exponential function. Moreover, the following proposition shows that GEOM is equivalent to DIV. Proposition 166. [↓] For parameter learning in log α , optimizing the base vector α by RSGD is equivalent to optimizing the divisor matrix B by Euclidean stochastic gradient descent (ESGD). Consequently, the three distinct update schemes are RELU, DIV, and MUL, as summarized in Tab. 6.1. 6.2.4 Experiments In this section, we validate the efficacy of our approaches on multiple data sets. Rieman- nian metrics are foundational to Riemannian neural networks. Therefore, our ALEM can redesign basic blocks in Riemannian neural networks. Beyond the main SPDNet experiments, we apply our ALEM to other Riemannian building blocks, including the LieBN framework in Chapter 3, Riemannian residual blocks [117], and Riemannian classifiers [159]. More details on data sets and experimental settings are provided in 197 6.2. Adaptive Log-Euclidean Metrics Learning rate1e −2 5e −2 Architecture 93, 30 93, 70, 30 93, 70, 50, 30 93, 30 93, 70, 30 93, 70, 50, 30 SPDNet62.92±0.8162.87±0.6063.03±0.6763.89±0.7364.00±0.6563.72±0.61 SPDNetBN63.03±0.7558.27±1.752.02±2.3463.75±0.6948.78±5.1537.84±6.10 ALog-MUL63.52±0.7563.86±0.5863.94±0.4464.4±0.6864.60±0.6964.36±0.49 ALog-DIV63.60±0.7963.93±0.5263.81±0.764.81±0.6464.84±0.6564.80±0.36 ALog-RELU63.02±0.7963.94±0.6463.14±0.6563.97±0.7564.10±0.6363.78±0.46 Table 6.2: Results of ALog on the HDM05 data set. The best results are bold. Secs. A.1 and A.3.7. 6.2.4.1 Applications in SPDNet In existing SPD neural networks, activation and classification layers commonly map SPD features into the logarithmic domain through the matrix logarithm [106, 232, 37, 156, 47]. This mapping is an isomorphism that identifies the SPD manifold under LEM with the Euclidean space S n . Replacing the natural matrix logarithm log with the learnable general logarithm log α allows the resulting layer to adapt the underlying geometry to the learned SPD features. We focus on the classic SPDNet [106], whose BiMap, ReEig, and LogEig layers are reviewed in Sec. A.2.1. Specifically, replacing the matrix logarithm in its LogEig layer with the learnable general matrix logarithm log α yields the Adaptive Logarithm (ALog) layer. We compare SPDNet, SPDNetBN, and the ALog-RELU/MUL/DIV variants on HDM05, FPHA, and AFEW. On the three data sets, the numbers of training epochs are set to 200, 500, and 100. We verify our ALog on SPDNet with various architectures. In addition, we test the robustness of the proposed layer against different learning rates on the HDM05 and FPHA data sets. Generally speaking, among all three implementations, ALog-MUL shows the most robust performance gain and achieves consistent improvement over the vanilla matrix logarithm. We also observe that ALog-MUL is comparable to or even better than SPDNetBN, which, however, introduces substantially greater complexity than our approach. The main reason for the superiority of our ALog against the vanilla matrix logarithm is that our ALog can adaptively respect the vibrant geometry of SPD manifolds, depending on the characteristics of data sets, while only LEM can be respected by the matrix logarithm. The following are detailed observations and analyses. Results on the HDM05 data set. The 10-fold results are presented in Tab. 6.2, where the data split and weight initialization are randomized. Following Huang and 198 Chapter 6. Fast and Stable Geometries on SPD Manifolds 0100200300400500 Training epoch 0 20 40 60 80 Acc SPDNet SPDNet-ALog-MUL Figure 6.1: Accuracy curves on the FPHA data set. SPDNetSPDNetBN ALog MULDIVRELU 85.73±0.8086.83±0.7487.8±0.7188.07±1.1386.65±0.68 Table 6.3: Results of ALog on the FPHA data set. Van Gool [106], three architectures are implemented on this data set, i.e., 93, 30, 93, 70, 30, and 93, 70, 50, 30. Generally speaking, endowed with ALog, SPDNet achieves consistent improvement. Among all three implementations, RELU only brings limited improvement. The reason might be that RELU fails to respect the innate geometry of the positive constraint. There is another interesting observation worth mentioning. In Brooks et al. [31], only the result of SPDNetBN under the architecture of 93, 30 is reported on this data set. Our experiments show that with the network going deeper, SPDNetBN tends to collapse, while our ALog layer performs robustly in all settings. Results on the FPHA data set. We validate our approach on this data set, with a learning rate of 1e −2 , over 10-fold cross-validation on random initialization. Since our experiments indicate that the vanilla SPDNet is already saturated with 1 BiMap layer, we just report the results on the architecture of 63, 33, which are presented in Tab. 6.3. Although DIV performs best on this data set, it presents the largest variance. There is an underlying nonlinear scaling mechanism in the update of DIV, which might undermine its robustness. Without loss of generality, let us focus on a single scalar parameter b in Eq. (6.32). The ultimate factor multiplied by the plain logarithm is 1/b. 199 6.2. Adaptive Log-Euclidean Metrics Therefore, the change of the multiplier after the update would be 1/(b− ∆)− 1/b = ∆/[(b− ∆)b].(6.37) Eq. (6.37) will scale the original ∆ to some extent. This scaling mechanism might undermine the robustness of the ALog layer. However, ALog-MUL achieves robust improvement and even surpasses SPDNetBN. This again demonstrates the significance of our adaptive mechanism for Riemannian deep networks. Finally, in terms of conver- gence analysis, accuracy curves with and without ALog are also reported in Fig. 6.1. Depth1234 SPDNet48.5346.8948.2447.22 SPDNetBN46.8946.6547.6248.35 ALog-MUL48.5748.1349.4550.62 ALog-DIV48.4248.0248.1349.89 ALog-RELU48.0647.2548.8648.1 Table 6.4: Results of ALog on the AFEW data set. Results on the AFEW data set. On this data set, the learning rate is 5e −2 , and we validate our method under four network architectures, i.e., 512, 100, 512, 200, 100, 512, 400, 200, 100, and 512, 400, 300, 200, 100. Note that, on this data set, SPDNetBN tends to present relatively large fluctuations in performance, so we compute the median of the last ten epochs. On various architectures, consistent improvement can be observed when SPDNet is endowed with our ALog. In addition, MUL performs best among all three implementations. Another interesting observation is that SPDNetBN seems ineffective on these deep features, while our methods show consistently superior perfor- mance, most notably for ALog-MUL. This indicates that our adaptive layer maintains effectiveness when applied to covariance matrices from deep features. Model complexity. Our ALog manifests the same complexity, no matter how it is optimized. Without loss of generality, the discussion below focuses on ALog-MUL. The extra computation and memory costs caused by the ALog layer are minor. It only depends on the final dimension of the network. Let us take the deepest one on the AFEW data set as an example. Our ALog only brings 100 unconstrained scalar parameters, while SPDNetBN needs an SPD matrix parameter for each RBN layer. The total number of parameters in RBN layers is 400 2 + 300 2 + 200 2 , which is much larger than ours. In addition, SPDNetBN needs to store the running mean of SPD matrices in every RBN layer, while our ALog only needs to store a vector. In terms of computation, the extra cost of our ALog is secondary as well. The forward and backward computation of our ALog is generally the same as the plain matrix logarithm, while computation in 200 Chapter 6. Fast and Stable Geometries on SPD Manifolds 051015202530 Diagonal elements of A 1.00 1.05 1.10 1.15 1.20 1.25 Acc SPDNet-ALog-MUL-[93, 30] SPDNet-ALog-MUL-[93, 70, 30] SPDNet-ALog-MUL-[93, 70, 50, 30] (a) HDM05. 051015202530 Diagonal elements of A 0.5 0.6 0.7 0.8 0.9 1.0 1.1 1.2 Acc SPDNet-ALog-MUL-[63, 33] SPDNet-ALog-MUL-[63, 53, 33] SPDNet-ALog-MUL-[63, 53, 43, 33] (b) FPHA. Figure 6.2: Visualization of parameters in the ALog layer on the HDM05 and FPHA data sets. Data SetHDM05FPHA Architecture93, 3093, 70, 3093, 70, 50, 3063, 33 SPDNet-Log263.93±0.8163.54±0.5063.98±0.6386.65±0.67 SPDNet63.89±0.7364.00±0.6563.72±0.6185.73±0.80 SPDNet-Log1063.45±0.3363.8±0.7163.64±0.6478.42±0.77 SPDNet-ALog-MUL64.4±0.6864.60±0.6964.36±0.4987.8±0.71 Table 6.5: Results of fixed bases on the HDM05 and FPHA data sets. the RBN layer is much more complex. All in all, our ALog can consistently improve the performance of SPDNet and achieve comparable or better results than SPDNetBN with much lower computation and memory costs. Visualization. We visualize the final learned parameters of the ALog layer. Since ALog-MUL is the most robust strategy, we visualize the parameters of ALog-MUL. Specifically, we plot the final values of the diagonal elements of A in Eq. (6.31) and visualize the results in Figs. 6.2a and 6.2b. We observe that the distribution of the parameters is consistent within the same data set but varies between data sets. This indicates that our approach can capture vibrant patterns in different data sets, respect- ing their specific geometry. Ablation studies. To further demonstrate the utility of the adaptive mechanisms in our approach, we validate the ALog layer with fixed bases. As decimal and binary are the two most common systems, we use log 10 and log 2 as examples of shrinking and expanding the natural logarithm log. Specifically, we set the scalar base α in log α to 10 and 2 in Eq. (6.30), respectively. We refer to the network with binary/decimal base as SPDNet-Log2/SPDNet-Log10. Note that when α = e, log α = log, and Eq. (6.30) 201 6.2. Adaptive Log-Euclidean Metrics MethodGeometry[93, 30][93, 70, 30][93, 70, 50, 30] NoneN/A63.89±0.7364.00±0.6563.72±0.61 SPDNetBNAIM63.75±0.6948.78±5.1537.84±6.10 SPDBNAIM64.33±0.8964.31±0.9263.62±1.21 LieBN-LEMLEM63.67±0.8565.77±0.8965.34±0.83 LieBN-ALEMALEM65.24±0.7170.11±0.9668.86±0.72 Table 6.6: Comparison of RBN methods on the HDM05 data set. reduces to the vanilla matrix logarithm. The network is then our baseline, i.e., SPDNet. We conduct 10-fold experiments on the HDM05 and FPHA data sets and set the learning rate to 5e −2 and 1e −2 , respectively, while keeping the other settings consistent with previous experiments. The results are presented in Tab. 6.5. We observe that the fixed logarithms show similar or slightly worse results than the vanilla log, while our ALog shows consistent improvement. Besides, log 10 does not converge on the FPHA data set. In fact, log 10 could shrink the gradient, slowing down convergence, especially under a small learning rate. In contrast, our ALog maintains consistent effectiveness. In summary, our ALog can respect vibrant geometry induced by log α and thus benefit SPD network learning. 6.2.4.2 Riemannian Batch Normalization The LieBN framework, its normalization guarantee, and its SPD manifestations are de- veloped in Chapter 3. As shown in Thm. 152,S n ++ ,⊕ ALE forms a Lie group. Besides, Thm. 157 demonstrates that ALEM is bi-invariant with respect to this group struc- ture. Therefore, LieBN under ALEM can also normalize Riemannian sample statistics. Following the LieBN algorithm and its SPD specialization, we implement LieBN under ALEM, denoted LieBN-ALEM. We compare it against AIM-based SPDNetBN [31] and SPDBN [124], and against LieBN under LEM. Following previous work [124, 31], we adopt the SPDNet backbone. Tab. 6.6 presents the 10-fold average results on the HDM05 data set under different network architec- tures. Our LieBN-ALEM achieves the best performance among the RBN methods. In particular, the AIM-based SPDNetBN brings worse performance under deeper architec- tures. In contrast, our LieBN-ALEM can consistently improve the performance across different architectures. Besides, compared with LieBN-LEM, our LieBN-ALEM shows better performance, demonstrating the effectiveness of our ALEM. 202 Chapter 6. Fast and Stable Geometries on SPD Manifolds 6.2.4.3 Riemannian Residual Blocks MethodHDM05NTU60 SPDNet63.89±0.7345.90±1.11 RResNet-AIM63.82±0.5845.22± 1.23 RResNet-LEM66.51±0.9348.73±0.60 RResNet-ALEM69.03±1.0657.09±0.59 Table 6.7: Experiments on RResNet under dif- ferent geometries. The general RResNet construction and SPD residual block are reviewed in Sec. A.2.5. Since its Rieman- nian exponential is metric-dependent, the Riemannian residual block un- der ALEM is obtained by substitut- ing Eq. (6.17) into that construc- tion. The required backpropagation of log −1 α is derived in Thm. 164. Following Katsman et al. [117], we compare RResNet under different geometries on the HDM05 and NTU60 data sets. Tab. 6.7 reports the 10-fold and 5-fold average results on these data sets. Compared with the vanilla SPDNet, RResNet-AIM brings little improvement, while LEM and ALEM show much better performance. In particular, the ALEM-based RResNet can bring a clear performance improvement, underscoring the effectiveness of our ALEM. 6.2.4.4 Riemannian Classifiers Learning rate1e −2 5e −2 GyroMLR-AIM54.28±0.4741.41±0.71 GyroMLR-LCM42.68±0.8842.06±0.49 GyroMLR-LEM53.22±0.4739.62±1.30 GyroMLR-ALEM56.21±0.3951.65±0.44 Table 6.8: Comparison of Gyro MLRs on the NTU60 data set. Euclidean MLR, which consists of FC and softmax, has become a stan- dard classification block in Euclidean neural networks. Inspired by this, Nguyen and Yang [159] extended Eu- clidean MLR to the SPD manifold using gyrostructures [157] for intrin- sic classification, referred to as gyro MLR. Three gyro MLRs under LCM, AIM, and LEM were introduced by Nguyen and Yang [159]. Following the logic in Nguyen and Yang [159, Sec. 2.4.2], we can obtain the gyro MLR under ALEM. Theorem 167 (Gyro MLR). [↓] Given an SPD feature S ∈S n ++ and C classes, the SPD gyro MLR under ALEM computes the multinomial probability of each class: p(y = k | S)∝ exp hD log α (S)− log α (P k ), (log α ) ∗,P k ( ̃ A k ) Ei ,(6.38) 203 6.3. Product Cholesky Metrics where k ∈1,...,C, P k ∈S n ++ , and ̃ A k ∈ T P k S n ++ . Following Chapter 4, we set ̃ A k = PT I n →P k (A k ) with A k ∈ T I n S n ++ . Therefore, the RHS of Eq. (6.38) becomes exp hD log α (S)− log α (P k ), (log α ) ∗,I n (A k ) Ei .(6.39) As (log α ) ∗,I n (A k )∈ T 0 S n ∼ = S n , we view (log α ) ∗,I n (A k ) as the parameter. We use SPDNet as the backbone. We compare Gyro MLR under our ALEM with those under LEM, LCM, and AIM on the NTU60 data set. Tab. 6.8 presents the 5-fold average results under different learning rates. Our ALEM outperforms the other met- rics within the gyro MLR framework. When the learning rate is 5e −2 , our GyroMLR- ALEM shows a larger performance advantage, especially compared with GyroMLR- LEM. These results demonstrate that Riemannian networks can benefit from the adap- tivity of our ALEM. 6.3 Product Cholesky Metrics 6.3.1 Introduction Whereas Sec. 6.2 introduces adaptive flexibility through pullback Euclidean geometry, this part pursues a complementary route centered on computational efficiency and nu- merical stability. LCM provides the natural bridge between these routes. It combines simple closed-form Riemannian operators with the fast and stable computation of the Cholesky decomposition [137]. LCM is induced by the Cholesky decomposition from a Riemannian metric on the Cholesky manifold, which is the space of lower triangular ma- trices with positive diagonal entries. We refer to this source metric as the diagonal log metric. Its interpretation as a pullback Euclidean metric was established in Thm. 147. We reveal a simple product structure underlying the diagonal log metric: a Euclidean metric on the strictly lower triangular part together with n copies of a Riemannian metric on R ++ for the diagonal part. This observation opens up a principled design space, as any metric on R ++ can induce a metric on the Cholesky manifold, and further yield a corresponding metric on the SPD manifold via the Cholesky decomposition. Building on this product structure, we introduce two Cholesky metrics, the di- agonal power metric and the diagonal Bures–Wasserstein metric, which induce two SPD metrics via the Cholesky decomposition: the Power-Cholesky Metric (PCM) and 204 Chapter 6. Fast and Stable Geometries on SPD Manifolds Bures–Wasserstein–Cholesky Metric (BWCM). Unlike the diagonal logarithm in LCM, our metrics rely on diagonal powers, improving numerical stability by avoiding expo- nentiation and logarithms. We further define in the Cholesky factors of SPD matrices a diagonal power deformation, which continuously connects existing and new metrics. As power θ → 0, the deformed metric converges to LCM, while at θ = 1 it recovers our proposed metrics, thereby offering a tunable trade-off. All proposed SPD metrics admit closed-form Riemannian operators, including geodesics, logarithmic and exponen- tial maps, parallel transport, as well as gyrovector operators [200], which extend vector addition and scalar multiplication into manifolds. These operators make our metrics directly applicable to SPD neural networks. In particular, by substituting these opera- tors into the Riemannian MLR formulation developed in Chapter 4 and residual blocks [118], we directly obtain SPD MLR classifiers and residual blocks under our metrics. We validate our metrics with experiments on SPD neural networks, numerical stability analyses, and tensor interpolation, showing the effectiveness, efficiency, and robustness of our metrics. In summary, our main contributions are: (1) Revealing the underlying simple product structure in the Cholesky manifold; (2) Proposing two Cholesky metrics and their SPD counterparts, PCM and BWCM, which admit fast and stable closed-form operators; (3) Developing SPD classifiers and residual blocks based on our metrics for SPD neural networks. 3 Outline. Sec. 6.3.2 recalls the Cholesky geometry used in this part. Sec. 6.3.3 develops the Cholesky product geometries, and Sec. 6.3.4 derives their SPD counter- parts. Their applications to SPD neural networks and the experimental evaluation are presented in Secs. 6.3.5 and 6.3.6. Proofs are deferred to Sec. B.10. 6.3.2 Preliminaries Pullback metrics are defined in Thm. 33, while the commonly used SPD geometries are summarized in Tabs. 2.5 and 2.6. We further recall the Cholesky geometry. The Euclidean space of n× n lower triangular matrices is denoted LT n . Its open subset, whose diagonal elements are all positive, is denoted by L n ++ . The Cholesky space L n ++ forms a submanifold of LT n [137]. For a Cholesky matrix L∈L n ++ and tangent vectors X,Y ∈ T L L n ++ , the Riemannian metric on the Cholesky manifold, referred to as the 3 The code is available at https://github.com/GitZH-Chen/PCM_BWCM. 205 6.3. Product Cholesky Metrics diagonal log metric, is g DL L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ +⟨L −1 X, L −1 Y⟩.(6.40) Here, ⌊X⌋ and ⌊Y⌋ are the strictly lower triangular parts of X and Y , while L, X, and Y are the diagonal matrices formed from their diagonal entries. LCM is the pullback metric of g DL by the Cholesky decomposition. As shown in Thm. 147, the diagonal log metric is the pullback metric, by the diagonal log map, of the Euclidean metric over LT n , which rationalizes our nomenclature. 6.3.3 Product Geometries on the Cholesky We first unveil the product structure beneath the existing diagonal log metric on the Cholesky manifold. Based on this, we propose two novel Cholesky metrics. 6.3.3.1 Disentangling the Cholesky Geometry We denote the space of n× n diagonal matrices with positive diagonal elements by Diag + (n) and the space of n× n strictly lower triangular matrices by LT 0 (n). Then, LT 0 (n) is a Euclidean space, and Diag + (n) ∼ = (R ++ ) n is an open submanifold of R n . Recalling Eq. (6.40), it is defined separately on LT 0 (n) and Diag + (n). Besides, Diag + (n) can be identified as the product of n copies of R ++ . The above discussion implies a product structure. We denote the standard Euclidean metric over LT 0 (n) by g E and define the Riemannian metric on R ++ as g R ++ p (v,w) = p −2 vw, ∀p∈ R ++ and v,w ∈ T p R ++ .(6.41) Then, L n ++ is the product manifold of LT 0 (n) and n copies of R ++ : L n ++ ,g DL =LT 0 (n),g E × n z | R ++ ,g R ++ ×·×R ++ ,g R ++ .(6.42) 6.3.3.2 Product Geometries on the Cholesky The following definition characterizes the underlying product structure in Eq. (6.42). Definition 168 (Product geometries). Suppose g LT 0 is a Euclidean inner product on LT 0 (n) andg i n i=1 are Riemannian metrics on R ++ . Then, the weighted product 206 Chapter 6. Fast and Stable Geometries on SPD Manifolds metric g on L n ++ is defined as g L (X,Y ) = g LT 0 (⌊X⌋,⌊Y⌋) + P n i=1 α i g i L i (X i ,Y i ), with L ∈ L n ++ , X,Y ∈ T L L n ++ , and α i > 0. Here, ⌊X⌋ and ⌊Y⌋ are the strictly lower triangular parts of X and Y , while L i , X i , and Y i are the i-th diagonal elements. For simplicity, we focus on the case where g LT 0 = g E is the standard Euclidean metric, all α i are equal to 1, and all g i are identical. Since R ++ can be viewed as a one-dimensional SPD manifold S 1 ++ , the Riemannian metrics reviewed in Sec. 2.9.1 can be immediately used to build Riemannian metrics on the Cholesky manifold. We additionally consider the Generalized Bures–Wasserstein Metric (GBWM), which is the pullback of BWM by S 7→ M − 1 2 SM − 1 2 for S,M ∈S n ++ [93]. For clarity, we denote PEM and GBWM by θ-EM and M-BWM, respectively. Simple computations show that AIM, LEM, and LCM coincide with Eq. (6.41) on S 1 ++ . Therefore, these metrics reduce to three classes on S 1 ++ : (1) LEM, LCM, or AIM; (2) θ-EM; (3) M-BWM or BWM. When the metric on R ++ is AIM (LEM or LCM), the resulting product metric on the Cholesky manifold is the diagonal log metric, and the pullback SPD metric via the Cholesky decomposition is exactly LCM. Inspired by the above analysis, we obtain two new metrics on the Cholesky manifold by setting each g i in Thm. 168 to θ-EM and M-BWM, termed the Diagonal Power Met- ric (θ-DPM) and Diagonal Bures–Wasserstein Metric (M-DBWM with M∈ Diag + (n)), respectively. By product geometries [131], we can obtain closed-form expressions for their Riemannian operators, such as the geodesic, logarithmic map, parallel transport, and weighted Fréchet mean. These operators are particularly important for building concrete learning algorithms [227, 132, 31, 141]. Theorem 169 (θ-DPM). [↓] Let L,K ∈ L n ++ and X,Y ∈ T L L n ++ , and let L i ∈ L n ++ N i=1 have weights w i N i=1 satisfying w i > 0 for all i and P N i=1 w i = 1. Then, the Riemannian operators under θ-DPM with θ ̸= 0 are g θ-DE L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ +⟨L θ−1 X, L θ−1 Y⟩,(6.43) γ (L,X) (t) =⌊L⌋ + t⌊X⌋ + L I n + tθL −1 X 1 θ ,(6.44) Log L (K) =⌊K⌋−⌊L⌋ + 1 θ L h L −1 K θ − I n i ,(6.45) PT L→K (X) =⌊X⌋ + L −1 K 1−θ X,(6.46) d 2 (L,K) =∥⌊K⌋−⌊L⌋∥ 2 F + 1 θ 2 ∥K θ − L θ ∥ 2 F ,(6.47) 207 6.3. Product Cholesky Metrics WFM(w i ,L i ) = X i w i ⌊L i ⌋ + X i w i L θ i 1 θ ,(6.48) where ∥·∥ F is the Frobenius norm. X, Y, L, K, and L i are diagonal matrices with diagonal elements from X, Y , L, K, and L i . γ (L,X) (t) denotes the geodesic starting at L with initial velocity X. PT L→K (·) is the parallel transport along the geodesic connecting L and K. Log, d, and WFM are the Riemannian logarithm, geodesic distance, and weighted Fréchet mean, respectively. Note that γ (L,X) (t) is locally defined in t∈ R|L + tθX∈ Diag + (n). Theorem 170 (M-DBWM). [↓] Following the notation in Thm. 169, the Rieman- nian operators under M-DBWM with M∈ Diag + (n) are g M-DBW L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ + 1 4 ⟨L −1 X, M −1 Y⟩,(6.49) γ (L,X) (t) =⌊L⌋ + t⌊X⌋ + L I n + t 1 2 L −1 X 2 ,(6.50) Log L (K) =⌊K⌋−⌊L⌋ + 2L h L −1 K 1 2 − I n i ,(6.51) PT L→K (X) =⌊X⌋ + L −1 K 1 2 X,(6.52) d 2 (L,K) =∥⌊K⌋−⌊L⌋∥ 2 F +∥M − 1 2 K 1 2 − L 1 2 ∥ 2 F ,(6.53) WFM(w i ,L i ) = X i w i ⌊L i ⌋ + X i w i L 1 2 i ! 2 ,(6.54) where the geodesic γ (L,X) (t) is locally defined in t ∈ R | L + t 2 X ∈ Diag + (n). When M = I n in M-DBWM, the resulting metric is denoted by DBWM. On the SPD manifold, GBWM is locally AIM [93]. Similarly, on the Cholesky manifold, our M-DBWM is locally the diagonal log metric at L∈L n ++ : g L-DBW L (X,Y ) = ⟨⌊X⌋,⌊Y⌋⟩ + 1 4 ⟨L −1 X, L −1 Y⟩. GBWM on the SPD manifold generally has no closed- form expression for the Fréchet mean [22]. Moreover, the closed-form expression of parallel transport under BWM is known only if two SPD matrices commute [196]. In contrast, all these operators have closed-form expressions under M-DBWM on L n ++ . 6.3.3.3 Deformed Cholesky Metrics On SPD manifolds, metrics deformed by the matrix power can interpolate between a given metric and an LEM-like metric [194, Sec. 3.1]. Inspired by this, we define 208 Chapter 6. Fast and Stable Geometries on SPD Manifolds a diagonal power deformation on the Cholesky manifold. For θ ̸= 0, we denote the diagonal power by DPow θ : Diag + (n) ∋ P 7−→ P θ ∈ Diag + (n). We will show how our proposed metric is connected to the existing diagonal log metric by DPow θ . Definition 171. LetL n ++ ,g =LT 0 (n),g E ×Diag + (n), ̃g be a product metric and θ ̸= 0. We define the diagonal-power-deformed metric of g as L n ++ ,g θ = LT 0 (n),g E ×Diag + (n), 1 θ 2 DPow ∗ θ ̃g. The following lemma shows that g θ in Thm. 171 converges to a diagonal-log-like metric as θ → 0. Lemma 172. [↓] Given L∈L n ++ and X,Y ∈ T L L n ++ , g θ in Thm. 171 satisfies g θ L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ + ̃g L θ L θ−1 X, L θ−1 Y −→ θ→0 ⟨⌊X⌋,⌊Y⌋⟩ + ̃g I n (L −1 X, L −1 Y). (6.55) Now, we discuss the deformation of the diagonal log metric, θ-DPM, and M-DBWM. First, Eq. (6.55) indicates that the diagonal-power-deformed metric of the diagonal log metric is itself. Second, θ-DPM is the diagonal-power-deformed metric of the Euclidean metric on the Cholesky manifold. Besides, θ-DPM interpolates between the diagonal log metric (θ → 0) and the Euclidean metric (θ = 1). Third, the diagonal-power-deformed metric of M-DBWM, referred to as (θ, M)-DBWM, is g (θ,M)-DBW L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ + 1 4 ⟨L θ−2 X, M −1 Y⟩.(6.56) When M = I n , the deformed metric of DBWM, i.e., θ-DBWM, tends to be a scaled diagonal log metric as θ → 0: lim θ→0 g θ-DBW L (X,Y ) =⟨⌊X⌋,⌊Y⌋⟩ + 1 4 ⟨L −1 X, L −1 Y⟩.(6.57) As (θ, M)-DBWM is the pullback metric by diagonal power and scaled by a constant, the Riemannian operators also have closed-form expressions, which are discussed in Sec. A.4.5.1. 6.3.3.4 Algebraic Structures As reviewed in Sec. 2.6, Eqs. (2.77) and (2.78) define gyroaddition and scalar gyro- multiplication on a Riemannian manifold. The gyro operations under the diagonal log metric reduce to vector-space operations, as the metric is induced by the Euclidean 209 6.3. Product Cholesky Metrics metric over the lower triangular matrices. This subsection studies the gyro-structures over θ-DPM and (θ, M)-DBWM. Let the identity matrix be the origin andC be θ-DPM or (θ, M)-DBWM. We have the following. Lemma 173 (Gyro-structures). [↓] For L,K ∈L n ++ and t∈ R, the gyro operations are L⊕ C K =⌊L⌋ +⌊K⌋ + L β + K β − I n 1 β ,(6.58) t⊙ C L = t⌊L⌋ + tL β + (1− t)I n 1 β ,(6.59) where β = θ for θ-DPM, and β = θ /2 for (θ, M)-DBWM. ⊕ C requires L and K to satisfy L β + K β − I n ∈ Diag + (n), while ⊙ C requires (1− t)I n + tL β ∈ Diag + (n). These gyro operations are defined only under the assumptions in Thm. 173, which arise from the locally defined Riemannian exponential map. Throughout, we impose these assumptions implicitly. Theorem 174. [↓] When the gyro operations are well defined under the conditions in Thm. 173, L n ++ ,⊕ C satisfies all the axioms of gyrocommutative gyrogroups (Thms. 51 and 52), and L n ++ ,⊕ C ,⊙ C satisfies all the axioms of gyrovector spaces (Thm. 53). Corollary 175. The identity element of L n ++ ,⊕ C is the identity matrix, i.e., ∀L∈L n ++ ,I n ⊕ C L = L. The inverses are ⊖ C L =−1⊙ C L =−⌊L⌋ + 2I n − L β 1 β , for L∈L∈L n ++ | 2I n − L β ∈ Diag + (n). Remark 176. The gyrostructures on the Grassmannian have shown success in build- ing Riemannian algorithms [157, 159]. Like our gyrostructure, the gyrostructures of the Grassmannian also require some assumptions for well-definedness [157, Sec. 3.2]. We therefore examine the well-definedness of the gyrostructures under our metrics. In practice, such positivity constraints can be remedied by numerical techniques. Taking 2I n − L β ∈ Diag + (n) as an example, one can use d i ← max(d i ,ε) for each diagonal element d i with a small constant ε > 0. 6.3.3.5 Numerical Advantages over Diagonal Log Metric Tab. 6.9 summarizes all the Riemannian and gyro operators. The Riemannian opera- tors under (θ, M)-DBWM and θ-DPM are mostly computed using the diagonal power 210 Chapter 6. Fast and Stable Geometries on SPD Manifolds OperatorsDiagonal Log Metricθ-DPM(θ, M)-DBWM g L (X,Y )⟨⌊X⌋,⌊Y⌋⟩ +⟨L −1 X, L −1 Y⟩⟨⌊X⌋,⌊Y⌋⟩ +⟨L θ−1 X, L θ−1 Y⟩⟨⌊X⌋,⌊Y⌋⟩ + 1 4 ⟨L θ−2 X, M −1 Y⟩ γ (L,X) (t)⌊L⌋ + t⌊X⌋ + L exp(tL −1 X) ⌊L⌋ + t⌊X⌋ + L (I n + tθL −1 X) 1 θ ⌊L⌋ + t⌊X⌋ + L I n + t θ 2 L −1 X 2 θ Log L (K)⌊K⌋−⌊L⌋ + L log(L −1 K)⌊K⌋−⌊L⌋ + 1 θ L h (L −1 K) θ − I n i ⌊K⌋−⌊L⌋ + 2 θ L h (L −1 K) θ 2 − I n i PT L→K (X)⌊X⌋ + (L −1 K)X⌊X⌋ + (L −1 K) 1−θ X⌊X⌋ + (L −1 K) 1− θ 2 X d 2 (L,K) ∥⌊K⌋−⌊L⌋∥ 2 F +∥log(K)− log(L)∥ 2 F ∥⌊K⌋−⌊L⌋∥ 2 F + 1 θ 2 ∥K θ − L θ ∥ 2 F ∥⌊K⌋−⌊L⌋∥ 2 F + 1 θ 2 ∥M − 1 2 K θ 2 − L θ 2 ∥ 2 F WFM(w i ,L i ) P i w i ⌊L i ⌋ + exp ( P i w i log(L i )) P i w i ⌊L i ⌋ + P i w i L θ i 1 θ P i w i ⌊L i ⌋ + P i w i L θ 2 i 2 θ L⊕ K⌊L⌋ +⌊K⌋ + LK⌊L⌋ +⌊K⌋ + L θ + K θ − I n 1 θ ⌊L⌋ +⌊K⌋ + L θ 2 + K θ 2 − I n 2 θ t⊙ Lt⌊L⌋ + L t t⌊L⌋ + tL θ + (1− t)I n 1 θ t⌊L⌋ + tL θ 2 + (1− t)I n 2 θ Table 6.9: Riemannian and gyro operators of different metrics on the Cholesky manifold. For the diagonal log metric, log(·) and exp(·) are diagonal logarithm and exponentiation. function, while those under the diagonal log metric are computed using the diagonal logarithm or exponentiation. This indicates that our θ-DPM and (θ, M)-DBWM may have better numerical stability than the existing diagonal log metric, as logarithm or exponentiation might overly stretch the diagonal elements compared with the power function. The gyro operations under our θ-DPM and (θ, M)-DBWM also have numeri- cal advantages over those under the diagonal log metric. The former are based on linear operations combined with power and its inverse, causing relatively minor changes to the input magnitude, while the gyro operations under the diagonal log metric are based on products or powers, resulting in more noticeable alterations to the input magnitude. 6.3.4 Geometries on the SPD Manifold This section discusses the Riemannian metrics on the SPD manifold via the Cholesky decomposition. We first review some basic properties of the Cholesky decomposition, followed by the SPD metrics. The Cholesky decomposition, denoted by Chol(·) : S n ++ → L n ++ , is a diffeomor- phism [137]. Therefore, it can pull back the Riemannian and gyro-structures from the Cholesky manifold L n ++ to the SPD manifold S n ++ . We call the pullbacks of θ-DPM and (θ, M)-DBWM through the Cholesky decomposition the Power-Cholesky Met- ric (θ-PCM) and Bures–Wasserstein–Cholesky Metric ((θ, M)-BWCM), respectively. Then, the Riemannian operators under θ-PCM and (θ, M)-BWCM can be obtained from the properties of Riemannian isometries (Thm. 33). Operators. Let C ∈ θ-DPM, (θ, M)-DBWM and S ∈ θ-PCM, (θ, M)-BWCM. We denote the Riemannian logarithm, exponential map, geodesic, parallel transport along the geodesic, geodesic distance, weighted Fréchet mean, gyroaddition, and scalar 211 6.3. Product Cholesky Metrics gyromultiplication on S n ++ ,g S by Log S , Exp S , γ S , PT S , d S (·,·), WFM S , ⊕ S , and ⊙ S , respectively, while Log C , Exp C , γ C , PT C , d C (·,·), WFM C , ⊕ C , and ⊙ C are their counterparts on L n ++ ,g C . For P,Q ∈ S n ++ , V,W ∈ T P S n ++ , and P i ∈ S n ++ N i=1 with weights w i N i=1 satisfying w i > 0 for all i and P N i=1 w i = 1, we have the following Riemannian and gyro operators: γ S (P,V ) (t) = Chol −1 γ C (L, e V ) (t) ,(6.60) Log S P (Q) = (Chol ∗,P ) −1 Log C L (K) ,(6.61) Exp S P (V ) = Chol −1 Exp C L e V ,(6.62) PT S P→Q (V ) = (Chol ∗,Q ) −1 PT C L→K e V ,(6.63) d S (P,Q) = d C (L,K),(6.64) WFM S (w i ,P i ) = Chol −1 WFM C (w i ,L i ) ,(6.65) P ⊕ S Q = Chol −1 (L⊕ C K),(6.66) t⊙ S P = Chol −1 (t⊙ C L),(6.67) where P = L ⊤ , Q = K ⊤ , and P i = L i L ⊤ i are Cholesky decompositions. Here, e V = Chol ∗,P (V ) is the Cholesky differential as defined in Sec. 2.8.1. Besides, when C is the diagonal log metric, the above recovers the Riemannian and gyro-structures under LCM. The above gyro operations, when well-defined, also satisfy the gyrovector-space axioms. Theorem 177. [↓] S n ++ ,⊕ S ,⊙ S satisfies all the axioms of gyrovector spaces. Remark 178. As discussed in Sec. 6.3.3.5, the Cholesky metrics θ-DPM and (θ, M)-DBWM are more numerically stable than the existing diagonal log met- ric. As pullback metrics through the Cholesky decomposition, our θ-PCM and (θ, M)-BWCM therefore preserve the advantage of numerical stability over the ex- isting LCM. Besides, all the Riemannian operators have closed-form expressions and are easy to use, as the differential maps of the Cholesky decomposition can be easily calculated. 212 Chapter 6. Fast and Stable Geometries on SPD Manifolds 6.3.5 Applications to SPD Neural Networks The closed-form operators derived in Sec. 6.3.4 make the proposed metrics directly applicable to SPD neural networks. In this section, we apply the proposed SPD metrics θ-PCM and (θ, M)-BWCM to build MLR classifiers and residual blocks on the SPD manifold. MLR. The Euclidean point-to-hyperplane formulation and its Riemannian exten- sion are developed in Chapter 4. Substituting the operators of the proposed metrics into that formulation gives the following SPD MLRs. Theorem 179. [↓] Given an input SPD matrix S ∈S n ++ , the C-class SPD MLRs under θ-PCM and (θ, M)-BWCM are θ-PCM : p(y = k | S ∈S n ++ )∝ exp ⟨⌊K⌋−⌊L k ⌋,⌊A k ⌋⟩ + 1 2θ ⟨K θ − L θ k , A k ⟩ ,(6.68) (θ, M)-BWCM : p(y = k | S ∈S n ++ )∝ exp ⟨⌊K⌋−⌊L k ⌋,⌊A k ⌋⟩ + 1 4θ ⟨K θ 2 − L θ 2 k , M −1 A k ⟩ , (6.69) where S = K ⊤ and P k = L k L ⊤ k are Cholesky decompositions. The parameters are P k ∈S n ++ and A k ∈ LT n for each class k = 1,· ,C. Residual blocks. The general construction and the specialized SPD residual-block expression are reviewed in Sec. A.2.5. The only component that varies across metrics is the Riemannian exponential map; substituting the operators derived above therefore gives residual blocks under the proposed metrics. 6.3.6 Experiments We first compare our metrics against the popular AIM, LEM, and LCM when building SPD MLR classifiers and residual blocks. Then, we evaluate the proposed metrics through numerical experiments. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.8. 6.3.6.1 Riemannian Classifiers Following the SPD learning setup in Sec. 4.4 and Huang and Van Gool [106], we adopt the Radar data set [31] for radar signal classification, and the HDM05 [153] and FPHA [80] data sets for human action recognition. We compare the SPD MLRs under our 213 6.3. Product Cholesky Metrics (a) Radar MetricAccTime AIM94.53± 0.950.80 LEM93.55± 1.210.76 LCM93.49± 1.250.72 θ-PCM95.79± 0.380.72 θ-BWCM93.93± 0.790.71 (b) HDM05 Metric 1-Block2-Block3-Block AccTimeAccTimeAccTime AIM58.07± 0.6417.3260.72± 0.6218.7561.14± 0.9419.23 LEM56.97± 0.612.2160.69± 1.022.9260.28± 0.913.50 LCM60.69± 1.891.8362.61± 1.462.4062.33± 2.152.90 θ-PCM62.51± 1.651.5863.66± 1.302.2965.75± 2.862.76 θ-BWCM62.71± 0.881.6464.52± 0.562.2767.40± 0.902.87 (c) FPHA MetricAccTime AIM85.57± 0.507.14 LEM85.90± 0.470.98 LCM86.37± 0.590.74 θ-PCM89.40± 0.130.69 θ-BWCM86.27± 0.600.70 Table 6.10: SPD MLRs under different metrics on the SPDNet backbone. The best results are bold. metrics with those under AIM, LEM, and LCM in Thm. 117. Following Sec. 4.4 and Nguyen and Yang [159], we adopt SPDNet [106] and GyroSPD [159] as two backbones, both mimicking feedforward neural networks. For simplicity, we set M in (θ, M)-BWCM to the identity matrix. SPDNet. On HDM05, we further evaluate architectures with up to three transfor- mation blocks. Tab. 6.10 reports the five-fold accuracy and training time per epoch, from which we draw the following observations. • Effectiveness. Our metrics generally yield higher accuracy than their counter- parts. Notably, they outperform LCM, although both originate from the Cholesky product structure. This improvement is attributed to the fact that the diagonal logarithm and exponentiation in LCM tend to overly stretch the diagonal entries, i.e., eigenvalues of the Cholesky factors, whereas our diagonal power transforma- tion achieves a more balanced scaling. • Efficiency. Our metrics substantially reduce computational cost compared to AIM, remain faster than LEM, and achieve efficiency comparable to LCM. To- gether with their superior accuracy, these results highlight the dual advantages of our approach in both effectiveness and efficiency. Metric RadarHDM05FPHA AccTimeAccTimeAccTime AIM96.80± 0.591.2366.05± 1.8021.6585.77± 0.5211.48 LEM96.58± 0.271.1866.42± 0.472.0285.87± 0.791.22 LCM96.29± 0.531.1268.37± 0.661.6689.83± 0.280.98 θ-PCM97.04± 0.641.1871.93± 1.211.5191.17± 0.301.00 θ-BWCM96.21± 0.251.0572.74± 0.431.5891.00± 0.110.96 Table 6.11: SPD MLRs on the GyroSPD backbone. GyroSPD. Tab. 6.11 reports the results on the GyroSPD backbone, which consists of a single gyrotranslation layer fol- lowed by an SPD MLR. Similar to the SPDNet results, our met- rics achieve comparable or supe- rior performance to LCM across all data sets while maintaining comparable efficiency. On HDM05 and FPHA, both proposed metrics deliver higher accuracy with lower runtime than AIM and LEM. On Radar, PCM achieves the highest accuracy, while BWCM has the lowest runtime. 214 Chapter 6. Fast and Stable Geometries on SPD Manifolds 6.3.6.2 Riemannian Residual Blocks Metric RadarHDM05FPHA AccTimeAccTimeAccTime AIM96.41.0257.011.1487.330.72 LEM97.070.8167.520.5286.170.32 LCM97.070.8566.270.6386.830.49 θ-PCM97.870.8568.050.6388.330.48 Table 6.12: Results on residual blocks. Riemannian ResNet (RResNet) was intro- duced by Katsman et al. [118]. Its back- bone architecture largely follows SPDNet. The key difference lies in the head: while SPDNet directly applies a classification layer, RResNet inserts a residual block be- fore the classification head. Accordingly, we adopt the following classification head under each metric: Log I n +FC + softmax. Since the Riemannian exponential is similar for θ-PCM and (θ, M)-BWCM, we focus on θ-PCM and compare it against AIM, LEM, and LCM in constructing RResNet. Tab. 6.12 reports the best results across three tri- als, showing that our metric consistently achieves superior accuracy while maintaining comparable efficiency. 6.3.6.3 Numerical Stability As discussed by Lin [137, p. 16], LCM is more stable than AIM and LEM owing to the numerical advantage of Cholesky decomposition over SVD. Moreover, as highlighted in Thm. 178, the essential distinction between our metrics and LCM lies in the diagonal operations: ours rely on diagonal power, while LCM employs diagonal exponentiation and logarithm. This structural difference grants our metrics stronger numerical stability and robustness compared with LCM, as well as AIM and LEM. To validate this, we evaluate geodesics on the Cholesky manifold. We generate 100,000 synthetic n × n Cholesky matrices L and tangent vectors X ∈ LT n , where each entry is uniformly sampled from [0, 1], and we set the smallest eigenvalue (diagonal entry) of L to ε. We test two representative sizes: 3× 3 matrices, commonly used in diffusion tensor imaging [10], and 256× 256 matrices, typical in computer vision [133, 206]. The deformation parameter θ is set to 1.5, 0.5, and 0.15. For (θ, M)-DBWM, we set M = I n . As shown in Tab. 6.13, our θ-DPM and θ-DBWM remain highly stable across a wide range of ε. For 3× 3 matrices, the diagonal log metric already deteriorates at ε = 1e −3 , with failure rates increasing rapidly as ε decreases. For 256 × 256 matrices, instability emerges even earlier at ε = 1e −1 . In contrast, our metrics remain stable down to ε = 1e −30 in most cases. The only exception occurs when θ = 0.15, where failures appear around ε = 1e −20 . This behavior is expected since both θ-DPM and θ-DBWM converge to the diagonal log metric as θ → 0, thereby inheriting its instability in this limit. Overall, 215 6.3. Product Cholesky Metrics ε 3× 3 for small matrices256× 256 for large matrices DLM θ = 1.5θ = 0.5θ = 0.15 DLM θ = 1.5θ = 0.5θ = 0.15 DPMDBWMDPMDBWMDPMDBWMDPMDBWMDPMDBWMDPMDBWM 1e −1 0.6200000014.29000000 1e −2 5.7000000018.48000000 1e −3 51.3200000058.35000000 1e −4 94.3400000095.02000000 1e −5 99.3900000099.47000000 1e −10 100000000100000000 1e −15 100000000100000000 1e −20 100000000.002100000000.02 1e −21 100000000.03100000000.01 1e −22 100000000.25100000000.23 1e −23 100000002.26100000002.42 1e −24 1000000022.981000000023.13 1e −25 1000000086.341000000086.58 1e −30 100 0000010010000000100 Table 6.13: Failure probabilities (%) of geodesics under different metrics with small eigenvalues in L ∈ L n ++ . An output matrix containing any Inf or NaN is considered a failure. Here, DLM denotes the diagonal log metric, while DPM and DBWM denote θ-DPM and θ-DBWM, respectively. these results demonstrate the superior numerical robustness of our metrics. 6.3.6.4 Tensor Interpolation As shown by Arsigny et al. [9, 10], geodesic interpolation of SPD matrices is important in diffusion tensor imaging. This experiment illustrates geodesic interpolation under different SPD metrics. Let P = L ⊤ and Q = K ⊤ be the Cholesky decompositions of P,Q ∈ S n ++ . The geodesics connecting P and Q under θ-PCM and (θ, M)-BWCM are θ-PCM: Chol −1 h ⌊L⌋ + t(⌊K⌋−⌊L⌋) + L θ + t(K θ − L θ ) 1 θ i ,(6.70) (θ, M)-BWCM: Chol −1 ⌊L⌋ + t(⌊K⌋−⌊L⌋) + L θ 2 + t(K θ 2 − L θ 2 ) 2 θ .(6.71) As the geodesics under θ-PCM and (θ, M)-BWCM have similar expressions, we focus on θ-PCM. Fig. 6.3 visualizes the geodesic interpolations on S 3 ++ under different metrics, including θ-EM, LEM, AIM, BWM, LCM, and θ-PCM. Tab. 6.14 presents the associated determinants of the interpolated SPD matrices. We can make the following observations. (1) The standard Euclidean metric (1-EM) exhibits a significant swelling effect, where the maximal determinant of interpolation is much larger than the determinants of the starting and end points. Although matrix power can mitigate the swelling effect, θ-EM still suffers from swelling. (2) We find that BWM also demonstrates a clear swelling effect. In contrast, LCM, 216 Chapter 6. Fast and Stable Geometries on SPD Manifolds Figure 6.3: Geodesic interpolation of SPD matrices under different Riemannian metrics. Each 3× 3 SPD matrix can be visualized as an ellipsoid [9]. The two endpoints are fixed across all metrics. AIM, and LEM show no swelling effect. (3) The trivial PCM (θ = 1) considerably mitigates the swelling effect compared to the Euclidean metric, but it still exhibits some level of swelling. However, by introducing Cholesky power deformation, θ-PCM effectively reduces the swelling effect. Notably, the swelling effect of θ-PCM is significantly weaker than that of θ-EM under the same θ. (4) Our PCM shows an interpolation visually similar to that under LCM. This suggests that our metric retains some practical potential of LCM but with better numerical stability (as demonstrated in Sec. 6.3.6.3). 6.3.6.5 Asymptotic Complexity We investigate the scalability of our PCM and BWCM in comparison with five existing SPD metrics, namely AIM, LEM, LCM, PEM, and BWM. We use SPD MLR as a representative application. We first analyze the asymptotic complexity of each SPD MLR. We then complement this analysis with synthetic experiments that measure the actual runtime of a single SPD MLR training step across different matrix dimensions. 217 6.3. Product Cholesky Metrics Metric The determinant of the i-th interpolation 0123456789 1.0-EM3.07104.86182.09234.38261.35262.64237.86186.64108.613.38 0.5-EM3.0718.6739.9359.5371.9673.7964.1445.0121.733.38 0.1-EM3.074.255.426.386.967.056.625.754.63.38 LEM3.073.13.143.173.23.243.273.313.343.38 AIM3.073.13.143.173.23.243.273.313.343.38 BWM3.0715.3232.0448.1459.2262.0955.3339.9319.983.38 LCM3.073.13.143.173.23.243.273.313.343.38 0.1-PCM3.073.153.233.293.343.373.393.43.43.38 0.5-PCM3.073.353.593.793.913.973.943.833.643.38 1.0-PCM3.073.64.074.464.724.834.764.494.033.38 Table 6.14: Swelling effects of geodesic SPD interpolations. Deeper greens indicate greater swelling. Metric Num. spectral matrix functions Num. Cholesky decompositions AIM1 + 2C0 LEM1 + C0 LCM01 + C PEM1 + C0 BWM1 + 3C PCM01 + C BWCM01 + C Table 6.15: Number of matrix functions required per sample for a C-class SPD MLR. Spectral matrix functions include matrix logarithm, matrix power, and the Lyapunov operator. As shown in Thms. 117 and 179, the C-class SPD MLRs under different metrics for the input S ∈S n ++ are AIM: p(y = k | S)∝ exp hD log P − 1 2 k SP − 1 2 k ,A k Ei , LEM: p(y = k | S)∝ exp [⟨log(S)− log(P k ),A k ⟩], LCM: p(y = k | S)∝ exp "* ⌊K⌋−⌊L k ⌋ + log(K) − log(L k ) ,⌊A k ⌋ + 1 2 A k +# , PEM: p(y = k | S)∝ exp 1 θ S θ − P θ k ,A k , BWM: p(y = k | S)∝ exp 1 2 D (P k S) 1 2 + (SP k ) 1 2 − 2P k ,L P k (L k A k L ⊤ k ) E , θ-PCM : p(y = k | S)∝ exp ⟨⌊K⌋−⌊L k ⌋,⌊A k ⌋⟩ + 1 2θ K θ − L θ k , A k , 218 Chapter 6. Fast and Stable Geometries on SPD Manifolds (θ, M)-BWCM : p(y = k | S)∝ exp ⟨⌊K⌋−⌊L k ⌋,⌊A k ⌋⟩ + 1 4θ D K θ 2 − L θ 2 k , M −1 A k E , where P k ∈ S n ++ and A k ∈ S n are MLR weights, log(·) is the matrix logarithm, and L P [V ] is the solution to the matrix linear systemL P [V ]P +PL P [V ] = V , known as the Lyapunov operator. Analysis. Tab. 6.15 summarizes the number of spectral and Cholesky matrix func- tions required by each SPD MLR. Cholesky decomposition requires O( 1 /3n 3 ) flops, while eigendecomposition costs O(9n 3 ) flops [83, Algs. 4.2.3 and 8.3.3]. Combining these counts, Tab. 6.16 reports the resulting asymptotic per-sample complexity for each met- ric. Cholesky-based metrics (LCM, PCM, BWCM) are asymptotically more efficient than the eigen-based metrics (LEM, PEM, AIM, and BWM), with AIM and BWM being the slowest among the considered methods. In addition, PCM and BWCM can be practically more efficient than LCM, since diagonal powers are cheaper to compute than diagonal logarithms. Metric Asymptotic complexity AIMO 9(1 + 2C)n 3 LEMO 9(1 + C)n 3 LCMO 1+C 3 n 3 PEMO 9(1 + C)n 3 BWMO (9(1 + 3C) + C 3 )n 3 PCMO 1+C 3 n 3 BWCMO 1+C 3 n 3 Table 6.16: Asymptotic per- sample complexity of a C-class SPD MLR for an n×n input SPD matrix. Setup. To validate the asymptotic complex- ity in Tab. 6.16, we measure the average wall- clock time of a single forward–backward training step of an SPD MLR classifier as the matrix di- mension increases. The model consists of a sin- gle SPD MLR layer with 50 output classes fol- lowed by a cross-entropy loss. For each dimension n ∈ 32, 64, 128, 256, 512, we randomly generate a batch of 30 n×n SPD matrices. In each run, we perform one forward and one backward pass and record the total runtime of this step. For PEM, we set the matrix power to 0.5. Results. As reported in Tab. 6.17, our PCM and BWCM are the fastest metrics across all tested dimensions, and the gap becomes particularly pronounced in the high-dimensional case. For small and medium scales (32 and 64), LCM, PCM, and BWCM have very similar runtimes and all are clearly faster than AIM, LEM, PEM, and BWM. When the dimension increases to 512, AIM and BWM require about 60 and 70 seconds per training step, whereas PCM and BWCM remain within roughly 1.7 seconds. In this setting, PCM and BWCM are even faster than LCM. 219 6.4. Conclusion DimAIMLEM LCM PEM BWMPCMBWCM 320.23800.00770.00460.00760.23770.00400.0040 641.01390.03950.03030.04731.12050.02510.0225 1283.62560.18320.14900.18444.0674 0.10130.1019 25614.51420.77930.58330.785316.59180.38480.4077 51260.19183.29482.50303.435770.86471.75531.7526 Table 6.17: Average runtime (in seconds) of one SPD MLR training step across different matrix dimensions. 6.4 Conclusion This chapter developed SPD metric design along two complementary routes. The first route proposed a pullback Euclidean framework and used a learnable general matrix logarithm to construct ALEM. Through this pullback construction, ALEM inherits compatible Hilbert-space, abelian Lie-group, and Riemannian structures, together with closed-form Riemannian operators and weighted Fréchet means. We further derived the differentials, gradients, and parameter-update schemes required to learn the logarithm bases. Experiments with applications to SPDNet, LieBN, Riemannian residual blocks, and gyro MLR demonstrate the effectiveness of our metrics. The second route revealed the product structure of the Cholesky manifold. This structure yielded diagonal power and diagonal Bures–Wasserstein geometries, whose deformed versions converge to the diagonal log metric as the power parameter ap- proaches zero. Transferring these geometries through the Cholesky decomposition pro- duced PCM and BWCM on the SPD manifold. The resulting metrics admit closed-form Riemannian and gyro operators, which directly yield SPD MLR classifiers and residual blocks. Experiments on SPD classifiers and residual networks, together with numerical experiments, supported their effectiveness, efficiency, and numerical robustness. Together, these routes show that the underlying Riemannian geometry can itself be designed to provide the flexibility, tractability, efficiency, and numerical stability required by Riemannian deep learning. 220 Chapter 7 Conclusion and Future Work 7.1 Conclusion This thesis studied Riemannian deep learning from three connected perspectives: uni- fied Riemannian module design across manifolds, manifold-specific Riemannian network design, and the design of the underlying Riemannian geometries. The main contribu- tions are summarized as follows: • Chapters 3 and 4 developed unified network modules from geometric structures shared across different manifolds. The normalization chapter first developed LieBN on Lie groups, then introduced pseudo-reductive gyrogroups, a new al- gebraic structure that generalizes classical gyrogroups and Lie groups, and finally developed GyroBN on this foundation, with LieBN recovered as a special case. The classification development progressed from SPD MLRs obtained through ex- act evaluation of the point-to-hyperplane infimum under flat pullback Euclidean metrics to a general RMLR based on a Riemannian-trigonometric formulation requiring only a well-defined Riemannian logarithm. • Chapter 5 developed manifold-specific Riemannian network designs by exploiting additional structures. PVNN developed the geometry and core neural layers of the stable PV model, Hyperbolic Busemann Neural Networks (HBNN) derived intrinsic and efficient BMLR and BFC layers from Busemann functions and horo- spheres on the Poincaré and Lorentz models, and CorNet developed correlation MLR, FC, and convolutional layers under five geometries together with accurate Riemannian backpropagation under OLM and LSM. • Chapter 6 designed the underlying Riemannian geometries. ALEM learned pull- 221 7.2. Future Work back Euclidean metrics through general matrix logarithms, whereas PCM and BWCM used the product structure of the Cholesky manifold to obtain efficient and numerically stable SPD geometries. Taken together, these contributions address Riemannian deep learning at the levels of unified network modules, manifold-specific network designs, and underlying Rieman- nian geometries while balancing intrinsic structure, generality, computational tractabil- ity, and numerical stability. 7.2 Future Work Two directions are particularly promising: representation learning with novel geometries and geometry-aware generative modeling. • Novel geometries for complex relational structure. Hyperbolic spaces provide an effective inductive bias for hierarchical data [164]. Although mixed- curvature products [87, 182] and matrix manifolds [59] have demonstrated the benefit of matching geometry to heterogeneous structure, real systems may also contain more complex relation types and structures. Future work could develop novel geometric structures to encode such complex relationships while retaining efficient and tractable Riemannian computations. • Geometry-aware generative modeling. Geometry arises both in latent repre- sentations and in spaces of probability distributions. At the latent level, Rieman- nian variational autoencoders [114], continuous normalizing flows [148], score- based models [64], and flow matching [43] have shown how manifold geometry can shape generative dynamics, while other work has analyzed the geometry of learned latent spaces [11, 168]. At the distribution level, a growing body of work has explored the geometry of probability distributions through information geom- etry [56, 63, 57] and optimal transport [96, 58]. A central problem is therefore to balance the geometry of the latent space with the geometry of the distributions evolving on it. 222 Bibliography [1] P-A Absil, Robert Mahony, and Rodolphe Sepulchre. Optimization Algorithms on Matrix Manifolds. Princeton University Press, 2008. [2] Bijan Afsari. Riemannian Lp center of mass: Existence, uniqueness, and convex- ity. In Proceedings of the American Mathematical Society, 2011. [3] Shun-ichi Amari. Information geometry and its applications, volume 194. Springer, 2016. URL https://doi.org/10.1007/978-4-431-55978-8. [4] Roy M Anderson and Robert M May. Infectious Diseases of Humans: Dynamics and Control. Oxford University Press, 1991. [5] Tsuyoshi Ando, Chi-Kwong Li, and Roy Mathias. Geometric means. Linear al- gebra and its applications, 385:305–334, 2004. URL https://doi.org/10.1016/ j.laa.2003.11.019. [6] Ilya Archakov and Peter Reinhard Hansen. A new parametrization of correlation matrices. Econometrica, 89(4):1699–1715, 2021. [7] Ilya Archakov and Peter Reinhard Hansen. A canonical representation of block matrices with applications to covariance and correlation matrices. Review of Economics and Statistics, 106(4):1099–1113, 2024. [8] Marc Arnaudon, Frédéric Barbaresco, and Le Yang. Riemannian medians and means with applications to radar signal processing. IEEE Journal of Selected Topics in Signal Processing, 7(4):595–604, 2013. URL https://doi.org/10. 1109/JSTSP.2013.2261798. [9] Vincent Arsigny, Pierre Fillard, Xavier Pennec, and Nicholas Ayache. Fast and simple computations on tensors with log-Euclidean metrics. PhD thesis, INRIA, 2005. URL https://doi.org/10.1007/11566465_15. 223 Bibliography [10] Vincent Arsigny, Pierre Fillard, Xavier Pennec, and Nicholas Ayache. Geometric means in a novel vector space structure on symmetric positive-definite matrices. SIAM journal on matrix analysis and applications, 29(1):328–347, 2007. [11] Georgios Arvanitidis, Lars Kai Hansen, and Søren Hauberg. Latent space oddity: on the curvature of deep generative models. ICLR, 2018. [12] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. [13] Gregor Bachmann, Gary Bécigneul, and Octavian Ganea. Constant curvature graph convolutional networks. In ICML, 2020. [14] Frédéric Barbaresco. Gaussian distributions on the space of symmetric positive definite matrices from Souriau’s Gibbs state for Siegel domains by coadjoint or- bit and moment map. In Geometric Science of Information: 5th International Conference, 2021. [15] Ahmad Bdeir, Kristian Schwethelm, and Niels Landwehr. Fully hyperbolic con- volutional neural networks for computer vision. In ICLR, 2024. [16] Ahmad Bdeir, Johannes Burchert, Lars Schmidt-Thieme, and Niels Landwehr. Robust hyperbolic learning with curvature-aware optimization. In NeurIPS, 2025. [17] Gary Becigneul and Octavian-Eugen Ganea. Riemannian adaptive optimization methods. In ICLR, 2019. [18] Gary Bécigneul and Octavian-Eugen Ganea. Riemannian adaptive optimization methods. In ICLR, 2019. [19] Thomas Bendokat, Ralf Zimmermann, and P-A Absil. A Grassmann manifold handbook: Basic geometry and computational aspects. Advances in Computa- tional Mathematics, 50(1):1–51, 2024. [20] Rajendra Bhatia. Positive Definite Matrices. Princeton University Press, 2007. [21] Rajendra Bhatia. Matrix analysis, volume 169 of Graduate Texts in Mathematics. Springer New York, 2013. doi: 10.1007/978-1-4612-0653-8. [22] Rajendra Bhatia, Tanvi Jain, and Yongdo Lim. On the Bures-Wasserstein dis- tance between positive definite matrices. Expositiones Mathematicae, 37(2):165– 191, 2019. 224 Bibliography [23] Victoria Bloom, Dimitrios Makris, and Vasileios Argyriou. G3D: A gaming ac- tion dataset and real time action recognition evaluation framework. In CVPR Workshops, 2012. [24] Clément Bonet, Lucas Drumetz, and Nicolas Courty. Sliced-Wasserstein distances and flows on Cartan-Hadamard manifolds. JMLR, 2025. [25] Silvere Bonnabel. Stochastic gradient descent on Riemannian manifolds. IEEE Transactions on Automatic Control, 58(9):2217–2229, 2013. URL https:// arxiv.org/abs/1111.5280. [26] Silvere Bonnabel, Anne Collard, and Rodolphe Sepulchre. Rank-preserving geo- metric means of positive semi-definite matrices. Linear Algebra and its Applica- tions, 438(8):3202–3216, 2013. [27] Nicolas Boumal. An introduction to optimization on smooth manifolds. Cambridge University Press, 2023. [28] Nicolas Boumal and P-A Absil. A discrete regression method on manifolds and its application to data on so (n). IFAC Proceedings Volumes, 44(1):2284–2289, 2011. [29] Martin R Bridson and André Haefliger. Metric spaces of non-positive curvature, volume 319. Springer Science & Business Media, 1999. [30] Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Van- dergheynst. Geometric deep learning: going beyond Euclidean data. IEEE Signal Processing Magazine, 34(4):18–42, 2017. [31] Daniel Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, and Matthieu Cord. Riemannian batch normalization for SPD neural networks. In NeurIPS, 2019. [32] Daniel A Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, and Matthieu Cord. Exploring complex time-series representations for Rieman- nian machine learning of radar data. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3672– 3676. IEEE, 2019. URL https://doi.org/10.1109/ICASSP.2019.8683056. 225 Bibliography [33] James W Cannon, William J Floyd, Richard Kenyon, and Walter R Parry. Hy- perbolic geometry. In Silvio Levy, editor, Flavors of Geometry, volume 31, pages 59–115. Cambridge University Press, 1997. [34] Rudrasis Chakraborty. ManifoldNorm: Extending normalizations on Riemannian manifolds. arXiv preprint arXiv:2003.13869, 2020. [35] Rudrasis Chakraborty and Baba C Vemuri. Statistics on the Stiefel manifold: theory and applications. The Annals of Statistics, 47(1):415–438, 2019. [36] Rudrasis Chakraborty, Chun-Hao Yang, Xingjian Zhen, Monami Banerjee, Derek Archer, David Vaillancourt, Vikas Singh, and Baba Vemuri. A statistical recurrent model on the manifold of symmetric positive definite matrices. In NeurIPS, 2018. [37] Rudrasis Chakraborty, Jose Bouza, Jonathan H Manton, and Baba C Vemuri. Manifoldnet: A deep neural network for manifold-valued data with applications. IEEE TPAMI, 2020. [38] Ines Chami, Zhitao Ying, Christopher Ré, and Jure Leskovec. Hyperbolic graph convolutional neural networks. NeurIPS, 2019. [39] Ines Chami, Albert Gu, Dat P Nguyen, and Christopher Ré. HoroPCA: Hyper- bolic dimensionality reduction via horospherical projections. In ICML, 2021. [40] Woong-Gi Chang, Tackgeun You, Seonguk Seo, Suha Kwak, and Bohyung Han. Domain-specific batch normalization for unsupervised domain adaptation. In CVPR, 2019. [41] Kai-Xuan Chen, Jie-Yi Ren, Xiao-Jun Wu, and Josef Kittler. Covariance de- scriptors on a Gaussian manifold and their application to image set classification. Pattern Recognition, 107:107463, 2020. doi: 10.1016/j.patcog.2020.107463. [42] Kaixuan Chen, Jie Song, Shunyu Liu, Na Yu, Zunlei Feng, Gengshi Han, and Mingli Song. Distribution knowledge embedding for graph pooling. IEEE TKDE, 2023. [43] Ricky TQ Chen and Yaron Lipman. Flow matching on general geometries. In ICLR, 2024. [44] Tianyu Chen, Xingcheng Fu, Yisen Gao, Haodong Qian, Yuecen Wei, Kun Yan, Haoyi Zhou, and Jianxin Li. Galaxy walker: Geometry-aware VLMs for galaxy- scale understanding. In CVPR, 2025. 226 Bibliography [45] Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. Fully hyperbolic neural networks. In ACL, 2022. [46] Yuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li, Ying Deng, and Weiming Hu. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In ICCV, 2021. [47] Ziheng Chen, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Zhiwu Huang, and Josef Kittler. Riemannian local mechanism for SPD neural networks. In AAAI, 2023. [48] Ziheng Chen, Tianyang Xu, Xiao-Jun Wu, Rui Wang, and Josef Kittler. Hy- brid Riemannian graph-embedding metric learning for image set classification. IEEE Transactions on Big Data, 9(1):75–92, 2023. doi: 10.1109/TBDATA.2021. 3113084. [49] Ziheng Chen, Yue Song, Gaowen Liu, Ramana Rao Kompella, Xiaojun Wu, and Nicu Sebe. Riemannian multinomial logistics regression for SPD neural networks. In CVPR, 2024. [50] Ziheng Chen, Yue Song, Yunmei Liu, and Nicu Sebe. A Lie group approach to Riemannian batch normalization. In ICLR, 2024. [51] Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, and Nicu Sebe. RMLR: Ex- tending multinomial logistic regression into general geometries. In NeurIPS, 2024. [52] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. Gyrogroup batch normal- ization. In ICLR, 2025. [53] Ziheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu, and Nicu Sebe. Understanding matrix function normalizations in covariance pooling through the lens of Rieman- nian geometry. In ICLR, 2025. [54] Ziheng Chen, Xiao-Jun Wu, Bernhard Schölkopf, and Nicu Sebe. Riemannian batch normalization: A gyro approach. arXiv preprint arXiv:2509.07115, 2025. [55] Ziheng Chen, Xiaojun Wu, Bernhard Schölkopf, and Nicu Sebe. Building transfor- mation layers for Riemannian neural networks, 2025. URL https://openreview. net/forum?id=1tJVBCpVD0. [56] Chaoran Cheng, Jiahan Li, Jian Peng, and Ge Liu. Categorical flow matching on statistical manifolds. In NeurIPS, 2024. 227 Bibliography [57] Chaoran Cheng, Jiahan Li, Jiajun Fan, and Ge Liu. α-flow: A unified framework for continuous-state discrete flow matching models, 2025. URL https://arxiv. org/abs/2504.10283. [58] Jaemoo Choi, Jaewoong Choi, and Myungjoo Kang. Scalable Wasserstein gradient flow for generative modeling through unbalanced optimal transport. In ICML, 2024. [59] Calin Cruceru, Gary Bécigneul, and Octavian-Eugen Ganea. Computationally tractable Riemannian manifolds for graph embeddings. In AAAI, 2021. [60] Jindou Dai, Yuwei Wu, Zhi Gao, and Yunde Jia. A hyperbolic-to-hyperbolic graph convolutional network. In CVPR, 2021. [61] Tingting Dan, Ziquan Wei, Won Hwa Kim, and Guorong Wu. Exploring the enigma of neural dynamics through a scattering-transform mixer landscape for Riemannian manifold. In ICML, 2024. [62] Paul David and Weiqing Gu. A Riemannian structure for correlation matrices. Operators and Matrices, 13(3):607–627, 2019. [63] Oscar Davis, Samuel Kessler, Mircea Petrache, Ismail Ilkan Ceylan, Michael Bron- stein, and Avishek Joey Bose. Fisher flow matching for generative modeling over discrete data. In NeurIPS, 2024. [64] Valentin De Bortoli, Emile Mathieu, Michael Hutchinson, James Thornton, Yee Whye Teh, and Arnaud Doucet. Riemannian score-based generative mod- elling. In NeurIPS, 2022. [65] Thibault de Surrel, Sylvain Chevallier, Fabien Lotte, and Florian Yger. Geometry- aware visualization of high dimensional symmetric positive definite matrices. TMLR, 2025. [66] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. [67] Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shan- mukha Ramakrishna Vedantam. Hyperbolic image-text representations. In ICML, 2023. 228 Bibliography [68] Abhinav Dhall, Amanjot Kaur, Roland Goecke, and Tom Gedeon. Emotiw 2018: Audio-video, student engagement and group-level affect prediction. In Proceedings of the 20th ACM International Conference on Multimodal Interaction, pages 653– 656, 2018. URL https://doi.org/10.1145/3242969.3264993. [69] Manfredo Perdigao Do Carmo and J Flaherty Francis. Riemannian Geometry, volume 6. Springer, 1992. [70] Ian L Dryden, Xavier Pennec, and Jean-Marc Peyrat. Power Euclidean metrics for covariance matrices with application to diffusion tensor imaging. arXiv preprint arXiv:1009.3045, 2010. URL https://arxiv.org/abs/1009.3045. [71] Alan Edelman, Tomás A Arias, and Steven T Smith. The geometry of algorithms with orthogonality constraints. SIAM journal on Matrix Analysis and Applica- tions, 20(2):303–353, 1998. [72] Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, and Ivan Oseledets. Hyperbolic vision transformers: Combining improvements in metric learning. In CVPR, 2022. [73] Xiran Fan, Chun-Hao Yang, and Baba Vemuri. Horospherical decision boundaries for large margin classification in hyperbolic space. In NeurIPS, 2023. [74] Maurice Fréchet. Les éléments aléatoires de nature quelconque dans un espace distancié. Annales de l’institut Henri Poincaré, 10(4):215–310, 1948. [75] Xingcheng Fu, Yisen Gao, Yuecen Wei, Qingyun Sun, Hao Peng, Jianxin Li, and Xianxian Li. Hyperbolic geometric latent diffusion model for graph generation. In ICML, 2024. [76] Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. Hyperbolic neural net- works. NeurIPS, 2018. [77] Zhi Gao, Yuwei Wu, Mehrtash Harandi, and Yunde Jia. A robust distance mea- sure for similarity-based classification on the spd manifold. IEEE TNNLS, 2019. [78] Zhi Gao, Yuwei Wu, Yunde Jia, and Mehrtash Harandi. Curvature generation in curved spaces for few-shot learning. In ICCV, 2021. [79] Zhi Gao, Chen Xu, Feng Li, Yunde Jia, Mehrtash Harandi, and Yuwei Wu. Ex- ploring data geometry for continual learning. In CVPR, 2023. 229 Bibliography [80] Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with RGB-D videos and 3D hand pose an- notations. In CVPR, 2018. [81] Mina Ghadimi Atigh, Martin Keller-Ressel, and Pascal Mettes. Hyperbolic buse- mann learning with ideal prototypes. In NeurIPS, 2021. [82] Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In AISTATS, 2011. [83] Gene H. Golub and Charles F. Van Loan. Matrix Computations. JHU press, 2013. [84] Alexandre Gramfort. MEG and EEG data analysis with MNE-Python. Frontiers in Neuroscience, 7, 2013. [85] Karish Grover, Geoffrey J Gordon, and Christos Faloutsos. CurvGAD: Leveraging curvature for enhanced graph anomaly detection. In ICML, 2025. [86] Karish Grover, Haiyang Yu, Xiang Song, Qi Zhu, Han Xie, Vassilis N Ioannidis, and Christos Faloutsos. Spectro-Riemannian graph neural networks. In ICLR, 2025. [87] Albert Gu, Frederic Sala, Beliz Gunel, and Christopher Ré. Learning mixed- curvature representations in product spaces. In ICLR, 2019. [88] Nicolas Guigui, Nina Miolane, and Xavier Pennec. Introduction to Riemannian geometry and geometric statistics: From basic theory to implementation with Geomstats. Foundations and Trends in Machine Learning, 16(3):329–493, 2023. doi: 10.1561/2200000098. [89] Johann Guilleminot and Christian Soize. Generalized stochastic approach for constitutive equation in linear elasticity: a random matrix model. International Journal for Numerical Methods in Engineering, 90(5):613–635, 2012. URL https: //doi.org/10.1002/nme.3338. [90] Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pascanu, Karl Moritz Hermann, Peter Battaglia, Victor Bapst, David Raposo, Adam San- toro, et al. Hyperbolic attention networks. In ICLR, 2019. [91] Yunhui Guo, Xudong Wang, Yubei Chen, and Stella X Yu. Clipped hyperbolic classifiers are super-hyperbolic classifiers. In CVPR, 2022. 230 Bibliography [92] Andi Han, Bamdev Mishra, Pratik Kumar Jawanpuria, and Junbin Gao. On Rie- mannian optimization over positive definite matrices with the Bures-Wasserstein geometry. NeurIPS, 2021. [93] Andi Han, Bamdev Mishra, Pratik Jawanpuria, and Junbin Gao. Learning with symmetric positive definite matrices via generalized Bures-Wasserstein geometry. In International Conference on Geometric Science of Information, pages 405–415. Springer, 2023. [94] Mehrtash Harandi, Mathieu Salzmann, and Richard Hartley. Dimensionality re- duction on SPD manifolds: The emergence of geometry-aware methods. IEEE TPAMI, 2018. [95] Richard Hartley, Jochen Trumpf, Yuchao Dai, and Hongdong Li. Rotation aver- aging. IJCV, 2013. [96] Doron Haviv, Aram-Alexandre Pooladian, Dana Pe’Er, and Brandon Amos. Wasserstein flow matching: Generative modeling over families of distributions. In ICML, 2025. [97] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. [98] Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Le- andros Tassiulas, Menglin Yang, and Rex Ying. HELM: Hyperbolic large language models via mixture-of-curvature experts. In NeurIPS, 2025. [99] Neil He, Menglin Yang, and Rex Ying. Lorentzian residual neural networks. In KDD, 2025. [100] Uwe Helmke and John B Moore. Optimization and Dynamical Systems. Springer Science & Business Media, 2012. [101] Marcel F. Hinss, Ludovic Darmet, Bertille Somon, Emilie Jahanpour, Fabien Lotte, Simon Ladouce, and Raphaëlle N. Roy. An EEG dataset for cross-session mental workload estimation: Passive BCI competition of the Neuroergonomics Conference 2021, 2021. [102] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997. 231 Bibliography [103] Chen Hu, Rui Wang, Xiaoning Song, Tao Zhou, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. A correlation manifold self-attention network for EEG decoding. In IJCAI, 2025. [104] Chen Hu, Ziheng Chen, Rui Wang, Yefeng Zheng, and Nicu Sebe. Riemannian high-order pooling for brain foundation models. In ICLR, 2026. [105] Xiaoqiang Hua, Yongqiang Cheng, Hongqiang Wang, Yuliang Qin, Yubo Li, and Wenpeng Zhang. Matrix CFAR detectors based on symmetrized Kullback–Leibler and total Kullback–Leibler divergences. Digital Signal Processing, 69:106–116, 2017. URL https://doi.org/10.1016/j.dsp.2017.06.019. [106] Zhiwu Huang and Luc Van Gool. A Riemannian network for SPD matrix learning. In AAAI, 2017. [107] Zhiwu Huang, Chengde Wan, Thomas Probst, and Luc Van Gool. Deep learning on Lie groups for skeleton-based action recognition. In CVPR, 2017. [108] Zhiwu Huang, Jiqing Wu, and Luc Van Gool. Building deep networks on Grass- mann manifolds. In AAAI, 2018. [109] Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep net- work training by reducing internal covariate shift. In ICML, 2015. [110] Catalin Ionescu, Orestis Vantzos, and Cristian Sminchisescu. Matrix backpropa- gation for deep networks with structured layers. In ICCV, 2015. [111] Catalin Ionescu, Orestis Vantzos, and Cristian Sminchisescu. Training deep networks with structured layers by matrix backpropagation. arXiv preprint arXiv:1509.07838, 2015. [112] Vinay Jayaram and Alexandre Barachant. MOABB: trustworthy algorithm benchmarking for BCIs. Journal of Neural Engineering, 15(6):066011, 2018. [113] Shaocheng Jin, Tao Zhou, Rui Wang, Ziheng Chen, Xiaoqing Luo, Xiao-Jun Wu, and Josef Kittler. Towards robust EEG decoding based on Riemannian self- attention. In KDD, 2026. [114] Dimitris Kalatzis, David Eklund, Georgios Arvanitidis, and Søren Hauberg. Vari- ational autoencoders with riemannian brownian motion priors. In ICML, 2020. 232 Bibliography [115] Huan Kang, Hui Li, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Chunyang Cheng, and Josef Kittler. SMLNet: A SPD manifold learning network for infrared and visible image fusion. IJCV, 2025. [116] Hermann Karcher. Riemannian center of mass and mollifier smoothing. Com- munications on Pure and Applied Mathematics, 30(5):509–541, 1977. URL https://doi.org/10.1002/cpa.3160300502. [117] Isay Katsman, Eric Ming Chen, Sidhanth Holalkere, Anna Asch, Aaron Lou, Ser- Nam Lim, and Christopher De Sa. Riemannian residual neural networks. In NeurIPS, 2023. [118] Isay Katsman, Eric Chen, Sidhanth Holalkere, Anna Asch, Aaron Lou, Ser Nam Lim, and Christopher M De Sa. Riemannian residual neural networks. NeurIPS, 2024. [119] Raiyan R Khan, Philippe Chlenski, and Itsik Pe’er. Hyperbolic genome embed- dings. In ICLR, 2025. [120] Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova, Ivan Oseledets, and Victor Lempitsky. Hyperbolic image embeddings. In CVPR, 2020. [121] Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [122] Diederik P Kingma and Jimmy Lei Ba. Adam: A method for stochastic opti- mization. In ICLR, 2015. [123] Reinmar Kobler, Jun-ichiro Hirayama, Qibin Zhao, and Motoaki Kawanabe. SPD domain-specific batch normalization to crack interpretable unsupervised domain adaptation in EEG. In NeurIPS, 2022. [124] Reinmar J Kobler, Jun-ichiro Hirayama, and Motoaki Kawanabe. Controlling the Fréchet variance improves batch normalization on the symmetric positive definite manifold. In ICASSP, 2022. [125] Max Kochurov, Rasul Karimov, and Serge Kozlukov. Geoopt: Riemannian opti- mization in pytorch. arXiv preprint arXiv:2005.02819, 2020. [126] A Krizhevsky. Learning multiple layers of features from tiny images. Master’s thesis, University of Tront, 2009. 233 Bibliography [127] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012. [128] Serge Lang. Algebra. Springer Science & Business Media, 2012. [129] Yann Le and Xuan Yang. Tiny ImageNet visual recognition challenge, 2015. [130] Guy Lebanon and John Lafferty. Hyperplane margin classifiers on the multinomial manifold. In ICML, 2004. [131] John M Lee. Introduction to Riemannian Manifolds, volume 2. Springer, 2018. [132] Mario Lezcano Casado. Trivializations for gradient-based optimization on mani- folds. In NeurIPS, 2019. [133] Peihua Li, Jiangtao Xie, Qilong Wang, and Zilin Gao. Towards faster training of global covariance pooling networks by iterative matrix square root normalization. In CVPR, 2018. [134] Shanglin Li, Motoaki Kawanabe, and Reinmar J Kobler. SPDIM: Source-free unsupervised conditional and label shift adaptation in EEG. In ICLR, 2025. [135] Shanglin Li, Shiwen Chu, Okan Koç, Yi Ding, Qibin Zhao, Motoaki Kawanabe, and Ziheng Chen. HEEGNet: Hyperbolic embeddings for EEG. ICLR, 2026. [136] Yancong Li, Xiaoming Zhang, Ying Cui, and Shuai Ma. Hyperbolic graph neu- ral network for temporal knowledge graph completion. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Re- sources and Evaluation (LREC-COLING 2024), 2024. [137] Zhenhua Lin. Riemannian geometry of symmetric positive definite matrices via Cholesky decomposition. SIAM Journal on Matrix Analysis and Applications, 40 (4):1353–1370, 2019. [138] Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. NTU RGB+D 120: A large-scale benchmark for 3D human activity understanding. IEEE T-PAMI, 2019. [139] Qi Liu, Maximilian Nickel, and Douwe Kiela. Hyperbolic graph neural networks. In NeurIPS, 2019. 234 Bibliography [140] Yuanpei Liu, Zhenqi He, and Kai Han. Hyperbolic category discovery. In CVPR, 2025. [141] Federico López, Beatrice Pozzetti, Steve Trettel, Michael Strube, and Anna Wien- hard. Vector-valued distance and Gyrocalculus on the space of symmetric positive definite matrices. In NeurIPS, 2021. [142] Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. Differentiating through the Fréchet mean. In ICML, 2020. [143] Miroslav Lovrić, Maung Min-Oo, and Ernst A Ruh. Multivariate normal distri- butions parametrized as a riemannian symmetric space. Journal of Multivariate Analysis, 74(1):36–48, 2000. [144] Jan R Magnus and Heinz Neudecker. Matrix differential calculus with applications in statistics and econometrics. John Wiley & Sons, 2019. URL https://doi. org/10.1002/9781119541219. [145] Luigi Malagò, Luigi Montrucchio, and Giovanni Pistone. Wasserstein Riemannian geometry of gaussian densities. Information Geometry, 1:137–179, 2018. [146] Jonathan H Manton. A globally convergent numerical algorithm for computing the centre of mass on compact Lie groups. In The 8th Control, Automation, Robotics and Vision Conference, 2004., volume 3, pages 2211–2216. IEEE, 2004. [147] Yidan Mao, Jing Gu, Marcus C Werner, and Dongmian Zou. Klein model for hyperbolic neural networks. arXiv preprint arXiv:2410.16813, 2024. [148] Emile Mathieu and Maximilian Nickel. Riemannian continuous normalizing flows. In NeurIPS, 2020. [149] Debin Meng, Xiaojiang Peng, Kai Wang, and Yu Qiao. Frame attention networks for facial expression recognition in videos. In 2019 IEEE International Conference on Image Processing (ICIP), pages 3866–3870. IEEE, 2019. URL https://doi. org/10.1109/ICIP.2019.8803603. [150] Aaron Meurer, Christopher P Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K Moore, Sartaj Singh, et al. Sympy: symbolic computing in python. PeerJ Computer Science, 3: e103, 2017. 235 Bibliography [151] Hà Quang Minh. Alpha Procrustes metrics between positive definite operators: a unifying formulation for the Bures-Wasserstein and Log-Euclidean/Log-Hilbert- Schmidt metrics. Linear Algebra and its Applications, 636:25–68, 2022. [152] Maher Moakher. On the averaging of symmetric positive-definite tensors. Journal of Elasticity, 82(3):273–296, 2006. URL https://doi.org/10.1007/ s10659-005-9035-z. [153] Meinard Müller, Tido Röder, Michael Clausen, Bernhard Eberhardt, Björn Krüger, and Andreas Weber. Documentation mocap database HDM05. Tech- nical report, Universität Bonn, 2007. [154] Iain Murray. Differentiation of the Cholesky decomposition. arXiv preprint arXiv:1602.07527, 2016. [155] Galileo Namata, Ben London, Lise Getoor, Bert Huang, and U Edu. Query-driven active surveying for collective classification. In 10th International Workshop on Mining and Learning with Graphs, 2012. [156] Xuan Son Nguyen. Geomnet: A neural network based on Riemannian geome- tries of SPD matrix space and Cholesky space for 3D skeleton-based interaction recognition. In ICCV, 2021. [157] Xuan Son Nguyen. The Gyro-structure of some matrix manifolds. In NeurIPS, 2022. [158] Xuan Son Nguyen. A Gyrovector space approach for symmetric positive semi- definite matrix learning. In ECCV, 2022. [159] Xuan Son Nguyen and Shuo Yang. Building neural networks on matrix manifolds: A Gyrovector space approach. In ICML, 2023. [160] Xuan Son Nguyen, Luc Brun, Olivier Lézoray, and Sébastien Bougleux. A neural network based on SPD manifold learning for skeleton-based hand gesture recog- nition. In CVPR, 2019. [161] Xuan Son Nguyen, Shuo Yang, and Aymeric Histace. Matrix manifold neural networks++. In ICLR, 2024. [162] Xuan Son Nguyen, Shuo Yang, and Aymeric Histace. Neural networks on sym- metric spaces of noncompact type. In ICLR, 2025. 236 Bibliography [163] Maximillian Nickel and Douwe Kiela. Poincaré embeddings for learning hierar- chical representations. In NeurIPS, 2017. [164] Maximillian Nickel and Douwe Kiela. Learning continuous hierarchies in the Lorentz model of hyperbolic geometry. In ICML, 2018. [165] Barrett O’Neill. Semi-Riemannian Geometry with Applications to Relativity, vol- ume 103 of Pure and Applied Mathematics. Academic Press, New York, 1983. [166] Avik Pal, Max van Spengler, Guido Maria D’Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. Compositional entailment learning for hyperbolic vision-language models. In ICLR, 2025. [167] Yue-Ting Pan, Jing-Lun Chou, and Chun-Shu Wei. MAtt: A manifold attention network for EEG decoding. In NeurIPS, 2022. [168] Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, and Youngjung Uh. Understanding the latent space of diffusion models through the lens of Riemannian geometry. In NeurIPS, 2023. [169] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019. [170] Xavier Pennec. Probabilities and statistics on Riemannian manifolds: A geometric approach. PhD thesis, INRIA, 2004. [171] Xavier Pennec, Pierre Fillard, and Nicholas Ayache. A riemannian framework for tensor computing. IJCV, 2006. [172] Can Pouliquen, Mathurin Massias, and Titouan Vayer. Schur’s positive-definite network: Deep learning in the SPD cone with structure. In ICLR, 2025. [173] John G Ratcliffe. Foundations of Hyperbolic Manifolds. Springer, 2006. [174] Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. Accelerating 3D deep learning with Py- Torch3D. arXiv preprint arXiv:2007.08501, 2020. [175] Herbert Robbins and Sutton Monro. A stochastic approximation method. The annals of mathematical statistics, pages 400–407, 1951. 237 Bibliography [176] Salem Said, Lionel Bombrun, Yannick Berthoumieu, and Jonathan H Manton. Riemannian Gaussian distributions on the space of symmetric positive definite matrices. IEEE TIT, 2017. [177] Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29 (3):93–93, 2008. [178] Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. NTU RGB+ D: A large scale dataset for 3D human activity analysis. In CVPR, 2016. [179] Xianglong Shi, Ziheng Chen, Yunhan Jiang, and Nicu Sebe. Intrinsic Lorentz neural network. In ICLR, 2026. [180] Ryohei Shimizu, Yusuke Mukuta, and Tatsuya Harada. Hyperbolic neural net- works++. In ICLR, 2021. [181] Ryohei Shimizu, YUSUKE Mukuta, and Tatsuya Harada. Hyperbolic neural networks++. In ICLR, 2021. [182] Ondrej Skopek, Octavian-Eugen Ganea, and Gary Bécigneul. Mixed-curvature variational autoencoders. In ICLR, 2020. [183] Yue Song, Nicu Sebe, and Wei Wang. Why approximate matrix square root outperforms accurate svd in global covariance pooling? In ICCV, 2021. [184] Yue Song, Nicu Sebe, and Wei Wang. On the eigenvalues of global covariance pooling for fine-grained visual recognition. IEEE TPAMI, 2022. [185] Yue Song, Nicu Sebe, and Wei Wang. Fast differentiable matrix square root. In ICLR, 2022. [186] Suvrit Sra and Reshad Hosseini. Conic geometric optimization on the manifold of positive definite matrices. SIAM Journal on Optimization, 25(1):713–739, 2015. URL https://doi.org/10.1137/140978168. [187] Shlomo Sternberg. Lectures on differential geometry, volume 316. American Math- ematical Soc., 1999. URL https://w.ams.org/journals/bull/1965-71-02/ S0002-9904-1965-11286-1/S0002-9904-1965-11286-1.pdf. 238 Bibliography [188] Yadong Sun, Xiaofeng Cao, Yu Wang, Wei Ye, Jingcai Guo, and Qing Guo. Ge- ometry awakening: Cross-geometry learning exhibits superiority over individual structures. In NeurIPS, 2024. [189] Tanuj Sur, Samrat Mukherjee, Kaizer Rahaman, Subhasis Chaudhuri, Muham- mad Haris Khan, and Biplab Banerjee. Hyperbolic uncertainty-aware few-shot incremental point cloud segmentation. In CVPR, 2025. [190] Yann Thanwerdas. Riemannian and stratified geometries on covariance and cor- relation matrices. PhD thesis, Université Côte d’Azur, 2022. [191] Yann Thanwerdas. Permutation-invariant log-Euclidean geometries on full-rank correlation matrices. SIAM Journal on Matrix Analysis and Applications, 45(2): 930–953, 2024. doi: 10.1137/22M1538144. [192] Yann Thanwerdas and Xavier Pennec. Is affine-invariance well defined on SPD matrices? a principled continuum of metrics. In Geometric Science of Informa- tion: 4th International Conference, GSI 2019, Toulouse, France, August 27–29, 2019, Proceedings 4, pages 502–510. Springer, 2019. [193] Yann Thanwerdas and Xavier Pennec. Exploration of balanced metrics on sym- metric positive definite matrices. In Geometric Science of Information: 4th In- ternational Conference, GSI 2019, Toulouse, France, August 27–29, 2019, Pro- ceedings 4, pages 484–493. Springer, 2019. [194] Yann Thanwerdas and Xavier Pennec. The geometry of mixed-Euclidean metrics on symmetric positive definite matrices. Differential Geometry and its Applica- tions, 81:101867, 2022. [195] Yann Thanwerdas and Xavier Pennec. Theoretically and computationally con- venient geometries on full-rank correlation matrices. SIAM Journal on Matrix Analysis and Applications, 43(4):1851–1872, 2022. doi: 10.1137/22M1471729. [196] Yann Thanwerdas and Xavier Pennec. O (n)-invariant Riemannian metrics on SPD matrices. Linear Algebra and its Applications, 661:163–201, 2023. [197] Loring W. Tu. An introduction to manifolds. Springer, 2011. [198] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: the missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016. 239 Bibliography [199] Abraham A Ungar. Analytic hyperbolic geometry: Mathematical foundations and applications. World Scientific, 2005. [200] Abraham Albert Ungar. Analytic Hyperbolic Geometry and Albert Einstein’s Spe- cial Theory of Relativity (Second Edition). World Scientific, 2022. [201] Max Van Spengler, Erwin Berkhout, and Pascal Mettes. Poincaré ResNet. In ICCV, 2023. [202] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. [203] Raviteja Vemulapalli, Felipe Arrate, and Rama Chellappa. Human action recog- nition by representing 3D skeletons as points in a Lie group. In CVPR, 2014. [204] Haoqi Wang, Zhizhong Li, and Wayne Zhang. Get the best of both worlds: Improving accuracy and transferability by Grassmann class representation. In ICCV, 2023. [205] Qilong Wang, Jiangtao Xie, Wangmeng Zuo, Lei Zhang, and Peihua Li. Deep cnns meet global covariance pooling: Better representation and generalization. IEEE TPAMI, 2020. [206] Qilong Wang, Zhaolin Zhang, Mingze Gao, Jiangtao Xie, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. Towards a deeper understanding of global covariance pooling in deep learning: An optimization perspective. IEEE TPAMI, 2023. [207] Rui Wang, Xiao-Jun Wu, and Josef Kittler. SymNet: A simple symmetric positive definite manifold deep learning method for image set classification. IEEE TNNLS, 2021. [208] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Tianyang Xu, and Josef Kittler. Dream- Net: A deep Riemannian manifold network for SPD matrix learning. In ACCV, 2022. [209] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Tianyang Xu, and Josef Kittler. Learning a discriminative SPD manifold neural network for image set classification. Neural Networks, 151:94–110, 2022. 240 Bibliography [210] Rui Wang, Chen Hu, Ziheng Chen, Xiao-Jun Wu, and Xiaoning Song. A Grass- mannian manifold self-attention network for signal classification. In IJCAI, 2024. [211] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Cong Hu, and Josef Kittler. SPD manifold deep metric learning for image set classification. IEEE TNNLS, 2024. [212] Rui Wang, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Towards a general attention framework on gyrovector spaces for matrix manifolds. In NeurIPS, 2025. [213] Rui Wang, Jiayao Jin, Ziheng Chen, Cong Wu, Xiao-Jun Wu, and Nicu Sebe. Structural topology refinement network for skeleton-based action recognition. IEEE TIM, 2025. [214] Rui Wang, Shaocheng Jin, Zhenyu Cai, Ziheng Chen, Xiao-Jun Wu, and Josef Kittler. Learning a better SPD network for signal classification: A Riemannian batch normalization method. IEEE TNNLS, 2025. [215] Rui Wang, Shaocheng Jin, Ziheng Chen, Xiaoqing Luo, and Xiao-Jun Wu. Learn- ing to normalize on the SPD manifold under Bures-Wasserstein geometry. In CVPR, 2025. [216] Rui Wang, Zihao Bi, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Riemannian graph convolutional network for skeleton-based two- person interaction recognition. In IJCAI, 2026. [217] Rui Wang, Yuting Jiang, Xiaoqing Luo, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Wasserstein-aligned hyperbolic multi-view clustering. AAAI, 2026. [218] Yunfeng Wang and Gregory S Chirikjian. Error propagation on the euclidean group with applications to manipulator kinematics. IEEE Transactions on Robotics, 22(4):591–602, 2006. [219] Yuxin Wu and Kaiming He. Group normalization. In ECCV, 2018. [220] Or Yair, Mirela Ben-Chen, and Ronen Talmon. Parallel transport on the cone manifold of SPD matrices for domain adaptation. IEEE TIP, 2019. [221] Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, and Rex Ying. Hypformer: Exploring efficient transformer fully in hyperbolic space. In KDD, 2024. 241 Bibliography [222] Menglin Yang, Ram Samarth B B, Aosong Feng, Bo Xiong, Jiahong Liu, Ir- win King, and Rex Ying. Hyperbolic fine-tuning for large language models. In NeurIPS, 2025. [223] Xin Yang, Xingrun Li, Heng Chang, Xihong Yang, Shengyu Tao, Maiko Shigeno, Ningkang Chang, Junfeng Wang, Dawei Yin, Erxue Min, et al. Hgformer: Hy- perbolic graph transformer for collaborative filtering. In ICML, 2025. [224] Ryoma Yataka, Kazuki Hirashima, and Masashi Shiraishi. Grassmann manifold flows for stable shape generation. In NeurIPS, 2023. [225] Florian Yger. A review of kernels on covariance matrices for BCI applications. In 2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), pages 1–6. IEEE, 2013. URL https://doi.org/10.1109/MLSP.2013. 6661972. [226] Hongwei Yong, Jianqiang Huang, Deyu Meng, Xiansheng Hua, and Lei Zhang. Momentum batch normalization for deep learning with small batch size. In Com- puter Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part XII 16, pages 224–240. Springer, 2020. [227] Ying Yuan, Hongtu Zhu, Weili Lin, and James Stephen Marron. Local polynomial regression for symmetric positive definite matrices. Journal of the Royal Statistical Society Series B: Statistical Methodology, 74(4):697–719, 2012. [228] Ernesto Zacur, Matias Bossa, and Salvador Olmos. Left-invariant Riemannian geodesics on spatial transformation groups. SIMAX, 2014. [229] Muhan Zhang and Yixin Chen. Link prediction based on graph neural networks. In NeurIPS, 2018. [230] Tong Zhang, Wenming Zheng, Zhen Cui, Yuan Zong, Chaolong Li, Xiaoyan Zhou, and Jian Yang. Deep manifold-to-manifold transforming network for skeleton- based action recognition. IEEE TMM, 2020. [231] Wei Zhao, Federico Lopez, J Maxwell Riestenberg, Michael Strube, Diaaeldin Taha, and Steve Trettel. Modeling graphs beyond hyperbolic: Graph neural networks in symmetric positive definite matrices. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 122–139. Springer, 2023. 242 Bibliography [232] Xingjian Zhen, Rudrasis Chakraborty, Nicholas Vogt, Barbara B Bendlin, and Vikas Singh. Dilated convolutional neural networks for sequential manifold-valued data. In ICCV, 2019. [233] Runhe Zhou, Shanglin Li, Guanxiang Huang, Xinliang Zhou, Qibin Zhao, Mo- toaki Kawanabe, Yi Ding, and Cuntai Guan. EEG-based multimodal learning via hyperbolic mixture-of-curvature experts. In ICML, 2026. [234] Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu. DNABERT-2: Efficient foundation model and benchmark for multi-species genome. In ICLR, 2024. 243 Bibliography 244 Appendix A Experimental Details and Additional Discussions A.1 Data Sets This section collects descriptions of the data sets used in the experimental chapters. A.1.1 Skeleton-Based Action Recognition and Gesture Data Sets HDM05 [153]. The HDM05 data set 1 consists of 2,273 skeleton-based motion capture sequences executed by different actors. Each frame records the 3D coordinates of 31 joints. Different experiments use the standard task-specific filtered versions adopted by their corresponding backbones, as detailed in the method-specific experimental settings. FPHA [80]. The FPHA data set 2 includes 1,175 skeleton-based first-person hand gesture videos of 45 different categories, with 600 clips for training and 575 for testing. Each frame contains the 3D coordinates of 21 hand joints. G3D [23]. The G3D data set consists of 663 sequences of 20 different gaming actions. Each sequence records the 3D locations of 20 joints, i.e.,, 19 bones. NTU60 [178]. The NTU60 data set 3 contains 56,880 skeleton sequences classified into 60 classes, where each frame includes 3D coordinates of 25 or 50 body joints. NTU120 [138]. The NTU120 data set 4 contains 114,480 sequences in 120 action classes. 1 https://resources.mpi-inf.mpg.de/HDM05/ 2 https://github.com/guiggh/hand_pose_action 3 https://github.com/shahroudy/NTURGB-D 4 https://github.com/shahroudy/NTURGB-D 245 A.1. Data Sets Data Set #Nodes #Edges #Classes #Features Disease1044104321000 Airport31881863144 PubMed19717443383500 Cora2708542971433 Table A.1: Summary statistics for the graph data sets. A.1.2 Radar and EEG Signal Data Sets Radar [31]. The Radar data set 5 contains 3,000 synthetic radar signals equally dis- tributed in 3 classes. Hinss2021 [101]. The Hinss2021 data set 6 is a competition data set contain- ing EEG signals for mental workload estimation. The data set is employed for two tasks, inter-session and inter-subject, which are treated as domain adaptation prob- lems. Geometry-aware methods [220, 123] have demonstrated promising performance in EEG classification. A.1.3 Image Classification Data Sets AFEW [68]. The Acted Facial Expressions in the Wild (AFEW) data set contains seven emotion categories, with 773 videos for training and 383 for validation. CIFAR-10 and CIFAR-100 [126]. The CIFAR-10 and CIFAR-100 data sets each contain 60,000 32× 32 color images from 10 and 100 classes, respectively. We use the standard PyTorch splits: 50,000 training images and 10,000 test images. Tiny-ImageNet [129]. Tiny-ImageNet is a subset of ImageNet with 100,000 im- ages from 200 classes, resized to 64× 64. We use the official validation split for evalu- ation. ImageNet-1k [66]. ImageNet-1k contains 1.28M training images, 50K validation images, and 100K test images distributed across 1k classes. A.1.4 Graph Data Sets Cora [177]. Cora is a citation network where the nodes represent scientific papers in the area of machine learning, the edges are citations between them, and the labels of the nodes are academic subareas. 5 https://w.dropbox.com/s/dfnlx2bnyh3kjwy/data.zip?dl=0 6 https://zenodo.org/record/5055046 246 Appendix A. Experimental Details and Additional Discussions TaskSpeciesData SetsNum. classes Max length Train / Dev / Test RetrotransposonsPlant LTR Copia 2 5007666 / 682 / 568 LINEs100022502 / 2030 / 1782 SINEs50021152 / 1836 / 1784 DNA transposonsPlant CMC-EnSpm 2 20019912 / 1872 / 1808 hAT-Ac100017322 / 1822 / 1428 PseudogenesHuman processed 21000 17956 / 1046 / 1740 unprocessed12938 / 766 / 884 Table A.2: Summary statistics for TEB. Disease [4]. Disease represents a disease propagation tree that simulates the sus- ceptible, infected, and recovered (SIR) disease transmission model, with each node representing either an infection or a non-infection state. Airport [229]. Airport is a transductive data set where nodes represent airports and edges represent airline routes from OpenFlights.org. PubMed [155]. PubMed is a standard benchmark that describes citation networks where nodes represent scientific papers in the area of medicine, the edges are citations between them, and the node labels are academic subareas. Tab. A.1 summarizes the above data sets. A.1.5 Genomic Sequence Data Sets Transposable Elements Benchmark [119]. The Transposable Elements Benchmark (TEB) comprises seven binary classification data sets that investigate transposable el- ements across plant and human genomes. The seven data sets are LTR Copia, LINEs, and SINEs for plant retrotransposons; CMC-EnSpm and hAT-Ac for plant DNA trans- posons; and processed and unprocessed pseudogenes for human pseudogenes. For each data set, positive examples are sequences spanning annotated elements of interest, and negatives are randomly sampled, non-overlapping genomic segments outside these re- gions. We adopt chromosome-level training, validation, and test splits, using chromo- somes 8 and 9 for validation and test in plant genomes, and chromosomes 20–22 and 17–19 for validation and test in human genomes, respectively. Tab. A.2 provides the summary statistics. Genome Understanding Evaluation [234]. The Genome Understanding Eval- uation (GUE) benchmark contains seven biologically significant genome analysis tasks that span 28 data sets. The data sets contain sequences ranging from 70 to 1000 base pairs in length and originate from yeast, mouse, human, and virus genomes. The HBNN experiments select Core promoter detection and Promoter detection, together with the 247 A.2. Backbone Networks TaskSpecies Data Set Num. classes Length Train / Dev / Test Core promoter detectionHuman tata 270 4904 / 613 / 613 notata42452 / 5307 / 5307 all47356 / 5920 / 5920 Promoter detectionHuman tata 2300 4904 / 613 / 613 notata42452 / 5307 / 5307 all47356 / 5920 / 5920 Covid variant classificationVirusCovid9100077669 / 7000 / 7000 Species classification FungiFungi2550008000 / 1000 / 1000 VirusVirus20100004000 / 500 / 500 Table A.3: Summary statistics for the adopted GUE data sets. multi-class tasks Covid variant classification and species classification, totaling nine data sets. Tab. A.3 provides their summary statistics. A.2 Backbone Networks This section collects the backbone network architectures used in the thesis. A.2.1 Backbone Networks on the SPD Manifold SPDNet. SPDNet [106] is a classical SPD neural network. It mimics conventional densely connected feedforward networks and consists of three basic building blocks: BiMap layer: S k = W k S k−1 W k⊤ , with W k semi-orthogonal,(A.1) ReEig layer: S k = U k−1 max(Σ k−1 ,εI n )U k−1⊤ , with S k−1 = U k−1 Σ k−1 U k−1⊤ , (A.2) LogEig layer: S k = log(S k−1 ).(A.3) where max(·) is element-wise maximization. BiMap and ReEig mimic transformation and non-linear activation, while LogEig maps SPD matrices into the tangent space at the identity matrix for classification. SPDNetBN. SPDNetBN [31] further proposed RBN based on AIM: Centering from mean M ∈S n ++ :∀i≤ N, ̄ P i ← M − 1 2 P i M − 1 2 ,(A.4) Biasing towards parameter B ∈S n ++ :∀i≤ N, ˆ P i ← B 1 2 ̄ P i B 1 2 .(A.5) TSMNet and SPDDSMBN. TSMNet [123] can be illustrated as f tc → f sc → f BiMap → f ReEig → f LogEig , where f tc and f sc denote temporal and spatial convolution, 248 Appendix A. Experimental Details and Additional Discussions respectively. SPDDSMBN [123] is an improved version of SPDNetBN. Apart from controlling the mean, it can also control variance. The key operation in SPDDSMBN for controlling the mean and variance is ∀i≤ N, ̄ P i ← Γ I→B [(Γ M→I (P i )) s v ],(A.6) where M is the Riemannian mean, v 2 is the Fréchet variance, B ∈ S n ++ is the biasing parameter, and s∈ R is the scaling factor. Inspired by Yong et al. [226], during training, SPDDSMBN generates running means and running variances for training and testing with distinct momentum parameters. It uses the training running statistics during training and the testing running statistics during testing. SPDDSMBN also applies domain-specific techniques [40], keeping multiple parallel BN layers and distributing observations according to the associated domains. To share cross-domain knowledge, s is uniformly learned across all domains, and B is set to the identity matrix. Kobler et al. [123] adopted SPDDSMBN for domain adaptation in EEG classification. SPDGCN [231]. SPDGCN is used as a Riemannian graph neural network back- bone. GyroSPD [159]. GyroSPD substitutes the BiMap layer in SPDNet with the AIM- based gyrotranslation S k = W k ⊕ AI S k−1 = W k 1 2 S k−1 W k 1 2 , where W k is an SPD matrix parameter. A.2.2 Backbone Networks on the Grassmannian GyroGr. GyroGr [159] mimics conventional densely connected feedforward networks and is composed of three basic building blocks. Given an ONB Grassmannian matrix U k−1 , the gyrotranslation and ProjMap layers are defined as Gyrotranslation: U k = W k ⊕ Gr U k−1 , W k ∈ Gr(p,n),(A.7) ProjMap: P k = U k−1 (U k−1 ) ⊤ .(A.8) In addition, the pooling is performed via the space of projection matrices [108]. Specifically, Grassmannian data are first mapped to the space of projection matrices via the ProjMap layer. A standard mean pooling operation is applied to the result- ing projection matrices. Finally, the pooled matrices are projected back to the ONB 249 A.2. Backbone Networks Grassmannian by SVD. The entire procedure can be expressed as P k = f p U k−1 (U k−1 ) ⊤ , U k = O k 1:p , where P k SVD := O k Σ k (O k ) ⊤ , (A.9) where f p is a regular mean pooling. A.2.3 Backbone Networks on Rotation Matrices LieNet. LieNet [107] is a classical neural network on rotation matrices. Its latent space is the Lie group SO N (3) = SO(3)×·× SO(3), i.e.,, R = (R 1 ,...,R N )∈ SO N (3). The group and manifold structures on SO N (3) are defined component-wise. For instance, R 1 ⊙ R 2 = (R 1 1 R 2 1 ,...,R 1 N R 2 N ). There are three basic layers in LieNet: RotMap layer: R k = W k ⊙ R k−1 , with W k ∈ SO N (3),(A.10) RotPooling layer: R k i = R k−1 m i ,n i , if Θ R k−1 m i ,n i > Θ R k−1 n i ,m i , R k−1 n i ,m i , otherwise, ,(A.11) LogMap layer: R k = log(R k−1 ),(A.12) where Θ(·) is the Euler angle, and (n i ,m i ) are two indices. The RotMap and RotPooling layers mimic the convolution and pooling layers, while the LogMap layer maps rotation matrices into the tangent space for classification. In LieNet, each rotation feature has shape [num, frame, 3, 3], where num and frame denote the spatial and temporal dimensions. The RotPooling layer is applied along either the spatial or temporal dimension, while the RotMap layer is applied along the spatial dimension, with W k of size [num, 3, 3]. A.2.4 Backbone Networks on Hyperbolic Spaces We briefly recap the hyperboloid layers, Poincaré MLR, Poincaré FC layer, and Poincaré β-concatenation used in the experiments. Throughout this subsection, K < 0. Hyperboloid Neural Networks. Let x ∈ H n K be the input vector. The weight parameters are W ∈ R m×(n+1) and v ∈ R n+1 . The Lorentz FC layer [45, Eq. 3] and 250 Appendix A. Experimental Details and Additional Discussions activation layer [15, Eq. 13] are defined in a spacetime manner: Activation: y = q ∥ψ (x s )∥ 2 − 1/K ψ (x s ) ,(A.13) FC: y = " p ∥φ(Wx,v)∥ 2 − 1/K φ(Wx,v) # ,(A.14) with φ(Wx,v) = λσ v ⊤ x + b ′ Wψ(x) + b ∥Wψ(x) + b∥ ,(A.15) where b ∈ R m and b ′ ∈ R are biases, ψ is an activation function, σ is the sigmoid function, and λ > 0 is a learnable scaling parameter. Poincaré MLR. Lebanon and Lafferty [130] first reformulated the Euclidean MLR p(y=k | x)∝ exp (⟨a k ,x⟩− b k ) via the point-to-hyperplane distance: p(y=k | x)∝ exp (sign(⟨a k ,x⟩−b k )∥a k ∥d(x,H a k ,b k )),(A.16) H a,b =x∈ R n |⟨a,x⟩− b = 0, where a∈ R n \0 and b∈ R. (A.17) For x ∈ P n K , Ganea et al. [76, Eqs. 24–25] generalized this formulation via geometric reinterpretation, and Shimizu et al. [181, Sec. 3.1] further derived the resulting closed form: v k (x) = 2∥z k ∥ p |K| asinh λ K x ⟨ p |K|x, [z k ]⟩ cosh(2 p |K|r k )− λ K x − 1 sinh(2 p |K|r k ) , where λ K x = 2(1−|K|∥x∥ 2 ) −1 is the conformal factor, p(y=k | x) ∝ exp (v k (x)), and [z k ] = z k ∥z k ∥ . Here, z k ∈ R n \0 and r k ∈ R are parameters. Under the identification a k = z k and b k =∥z k ∥r k , the Euclidean limit is lim K→0 v k (x) = 4(⟨a k ,x⟩− b k ). Poincaré FC Layer. Shimizu et al. [181] extended the Euclidean FC layer to the Poincaré ball via point-to-hyperplane distances. In Euclidean spaces, the FC can be written element-wise as y k = ⟨a k ,x⟩− b k with x,a k ∈ R n and b k ∈ R. Thus, y k equals ∥a k ∥ times the signed distance from x to the hyperplane H a k ,b k . Under this interpretation, the Poincaré FC layer F : P n K → P m K takes the closed form y = w 1 + p 1 +|K|∥w∥ 2 , w k =|K| −1/2 sinh p |K|v k (x) ,(A.18) where |K| is the magnitude of the curvature. Here, Z = z k m k=1 and r = r k m k=1 parameterize the orientations and biases, and v k (x) is the Poincaré MLR logit. 251 A.2. Backbone Networks Poincaré β-Concatenation. It generalizes the Euclidean concatenation into the hyperbolic Poincaré ball, stabilizing the norm of the Poincaré vector [181, Sec. 3.3]. Given inputs x i ∈ P n i K N i=1 , it is defined as Exp 0 β n β −1 n 1 v ⊤ 1 ,· ,β −1 n N v ⊤ N ⊤ ∈ P n K ,(A.19) where v i = Log 0 (x i ), n = P N i=1 n i , and β α = B ( α /2, 1 /2) is defined using the beta function. A.2.5 Riemannian Residual Network Backbones RResNet. Euclidean residual blocks can be written as x (i) = x (i−1) + n i x (i−1) ,(A.20) where n i is a network. Katsman et al. [118] generalized this to manifolds by replacing addition with the Riemannian exponential map: x (i) = Exp x (i−1) ℓ i (x (i−1) ) ,(A.21) where ℓ i : M → TM outputs a vector field parameterized by the neural network. As Exp x (v) = x +v for the Euclidean space, it can be immediately shown that Eq. (A.21) naturally extends Eq. (A.20) to manifolds. On the SPD manifold, the vector field is generated from the eigenvalues. Specifically, the SPD residual block [118, Eqs. 22–23] is Y = Exp X Q diag (f (spec(X)))Q ⊤ ,(A.22) where X ∈ S n ++ , spec(X) contains all eigenvalues, f : R n → R n is a neural network, and Q is an orthogonal parameter. Here, diag(·) returns a diagonal matrix from the input vector. 252 Appendix A. Experimental Details and Additional Discussions A.3 Experimental Details A.3.1 Lie Group Batch Normalization A.3.1.1 LieBN on the SPD Manifold The SPDNet and TSMNet backbones, including SPDNetBN and SPDDSMBN, are reviewed in Sec. A.2.1. We next describe the domain-specific momentum LieBN and implementation details. SPD modeling and preprocessing. The data set descriptions are collected in Sec. A.1. Throughout this thesis, every sample covariance matrix is made strictly posi- tive definite before subsequent manifold operations by adding a small diagonal pertur- bation, Σ← Σ +εI n with ε > 0. Unless otherwise stated, this preprocessing convention is applied to all covariance-based SPD representations. Following the protocol of Brooks et al. [31], each Radar signal is divided into windows of length 20, whose series yields one 20× 20 SPD covariance matrix. This produces 3,000 covariance matrices equally distributed across 3 classes. For HDM05, each frame consists of 3D coordinates of 31 joints, and each sequence is modeled by a 93× 93 covariance matrix. Following the protocol of Brooks et al. [31], we trim the data set down to 2,086 sequences scattered throughout 117 classes by removing some under-represented classes. For the FPHA, we follow Wang et al. [207] to represent each sequence as a 63× 63 covariance matrix. For Hinss2021, we choose the SOTA method, TSMNet [123], as our baseline model. We follow the Python implementation 7 of Kobler et al. [123] to carry out preprocessing. In detail, the Python packages MOABB [112] and MNE [84] are used to preprocess the data set. The applied steps include resampling the EEG signals to 250/256 Hz, apply- ing temporal filters to extract oscillatory EEG activity in the 4–36 Hz range, extracting short segments (≤ 3s) associated with a class label, and finally obtaining 40× 40 SPD covariance matrices. Domain-specific momentum LieBN for EEG classification. Kobler et al. [123] proposed SPDDSMBN as a domain adaptation approach for EEG classification. SPDDSMBN, based on Eq. (3.6), performed normalization of mean and variance on SPD manifolds under the specific AIM. Additionally, SPDDSMBN utilized separate momentum parameters for updating training and testing running statistics, inspired by Yong et al. [226]. Following Kobler et al. [123, Alg. 1], we also present a momentum 7 https://github.com/rkobler/TSMNet 253 A.3. Experimental Details Algorithm 3: Momentum LieBN (MLieBN) Algorithm Input: A batch of activations P i N i=1 over the Lie group M,⊕,g, and a small positive constant ε running mean ̄ M r = E, running variance ̄v 2 r = 1 for training running mean ̃ M r = E, running variance ̃v 2 r = 1 for testing biasing parameter B ∈M, scaling parameter s∈ R\0, momentum parameters for training and testing η train ,η ∈ [0, 1] Output : Normalized activations ̃ P i N i=1 if training then Compute batch mean M b and variance v 2 b of P i N i=1 ; ̄ M r ← WFM(1− η train ,η train , ̄ M r ,M b ); ̄v 2 r ← (1− η train ) ̄v 2 r + η train v 2 b ; ̃ M r ← WFM(1− η,η, ̃ M r ,M b ); ̃v 2 r ← (1− η) ̃v 2 r + ηv 2 b ; if training then M ← ̄ M r ,v 2 ← ̄v 2 r ; else M ← ̃ M r ,v 2 ← ̃v 2 r ; for i← 1 to N do Centering to the neutral element E: if g is left-invariant then ̄ P i ← L ⊖M (P i ); else ̄ P i ← R ⊖M (P i ); Scaling the dispersion: ˆ P i ← Exp E h s √ v 2 +ε Log E ( ̄ P i ) i Biasing towards parameter B: if g is left-invariant then ̃ P i ← L B ( ˆ P i ); else ̃ P i ← R B ( ˆ P i ); LieBN (MLieBN) in Alg. 3. Here η is fixed and η train is defined as η train = 1− ρ 1 T−1 max(T−t,0) + ρ, where ρ = 1 domains_per_batch ,(A.23) where T and t denote the total number of training epochs and the current epoch, respec- tively. Furthermore, following Kobler et al. [123], we adopt multi-channel mechanisms for domain-specific MLieBN (DSMLieBN), where each domain has its own MLieBN layer. Similar to Kobler et al. [123], we set the biasing parameter equal to the neu- tral element, and the scaling factor is shared across all domains. We denote Alg. 3 as MLieBN(P j | B,s,ε,η,η train ). Then our DSMLieBN follows DSMLieBN(P j ,i) = MLieBN i (P j | E,s,ε,η,η train ), ∀P j ∈P k N k=1 ,(A.24) 254 Appendix A. Experimental Details and Additional Discussions where i is the index of the domain. We follow the official code of SPDDSMBN 8 to imple- ment our DSMLieBN. Thus, the only difference between DSMLieBN and SPDDSMBN is how normalization is performed. Analogous to Thm. 72, computations for DSM- LieBN under pullback metrics can also be performed by mapping, calculating, and then remapping. Implementation details. We use the official code of SPDNetBN 9 [31] and TSM- Net 10 [123] to implement our experiments on the SPDNet and TSMNet backbones. For the SPDNet architecture, we compare our LieBN with SPDNetBN [31], which applies the SPDBN (Eqs. (3.2) and (3.3)) to SPDNet. Similar to SPDNetBN, we apply our LieBN after each transformation layer. In the EEG application, one of the state-of- the-art methods is TSMNet+SPDDSMBN [123], which is a domain adaptation version of Kobler et al. [124]. For a fair comparison, we also implement a domain-specific mo- mentum LieBN, referred to as DSMLieBN. Following Kobler et al. [123], we apply our DSMLieBN before the LogEig layer in TSMNet. We use the standard cross-entropy loss and optimize the parameters with the Riemannian AMSGrad optimizer [18]. The network architectures are represented as d 0 ,d 1 ,...,d L , where the dimension of the parameter in the i-th BiMap layer is d i × d i−1 . The experiments are conducted with a learning rate of 5e −3 , a batch size of 30, and 200 training epochs on the Radar, HDM05, and FPHA data sets. For the Hinss2021 data set, following Kobler et al. [123], we use a learning rate of 1e −3 with a weight decay of 1e −4 , a batch size of 50, and 50 training epochs. Scoring metrics. In line with the previous work [31, 123], we use accuracy as the scoring metric for the Radar, HDM05, and FPHA data sets, and balanced accuracy (i.e., the average recall across classes) for the Hinss2021 data set. Ten-fold experiments on the Radar, HDM05, and FPHA data sets are carried out with randomized initialization and split (the split is officially fixed for the FPHA data set), while on the Hinss2021 data set, models are fit and evaluated with randomized leave-5%-of-sessions-out (inter- session) or leave-5%-of-subjects-out (inter-subject) cross-validation. A.3.1.2 LieBN on Rotation Matrices Data sets and preprocessing. The data set descriptions are collected in Sec. A.1. Following Huang et al. [107], the G3D experiments use the cross-subject setting, with 8 https://github.com/rkobler/TSMNet 9 https://proceedings.neurips.c/paper_files/paper/2019/file/ 6e69ebbfad976d4637b4b39de261bf7-Supplemental.zip 10 https://github.com/rkobler/TSMNet 255 A.3. Experimental Details half of the subjects used for training and the other half for testing, while the NTU60 experiments use the cross-view protocol [178]. We use the code 11 of Vemulapalli et al. [203] to represent each skeleton sequence as a point on the Lie group SO N×T (3), where N and T denote spatial and temporal dimensions. As preprocessed in Huang et al. [107], we set T to 100, 16, and 64 on the G3D, HDM05, and NTU60 data sets, respectively. LieNet. The LieNet backbone is reviewed in Sec. A.2.3. Note that the official code of LieNet 12 is implemented in MATLAB. We use the open-source PyTorch code 13 to implement our experiments. To reproduce LieNet more faithfully, we made the following modifications to this PyTorch code. We reimplemented the LogMap and RotPooling layers to make them consistent with the official MATLAB implementation. In addition, we extended the Riemannian computations of Geoopt [125] to SO(3) to enable the direct Riemannian optimization, which is missing from the current package. We apply our LieBN before the LogMap layer. Note that the dimension of input features in LieNet is B× N × T × 3× 3. We calculate Lie group statistics along the batch and temporal dimensions (B × T). We denote the LieNet models with our LieBN-Left and LieBN- Right as LieNetLieBN-Left and LieNetLieBN-Right, respectively. Implementation details. We find that SGD is the most effective optimizer for LieNet, and thus, we adopt it for our experiments. The learning rate is set to 1e −2 . The batch sizes are 30, 30, and 256 for the G3D, HDM05, and NTU60 data sets, respectively. On the NTU60 data set, the learning rate is reduced by a factor of 10 upon model convergence, specifically at the 5th and 25th epochs for LieNetLieBN and LieNet, respectively. For each model, we apply torch.n.utils.clip_grad_norm_ with max_norm=5 to the transformation matrix in the final FC layer. A.3.1.3 LieBN on Correlation Matrices We follow the same settings as the experiments on the SPD manifold with respect to the backbone architecture, batch size, number of training epochs, optimizer, and learning rate. The network architecture can be denoted as BiMap-[Power-Cov2Cor- LieBN-Cor]-LogEig, where Power denotes the matrix power and Cov2Cor is Cor(·) : S n ++ → Cor + (n). The matrix powers used for each data set are presented in Tab. A.4. A single iteration is sufficient to achieve saturated network performance with respect to calculating D and D ⋆ in OLM and LSM, except for D ⋆ on the HDM05 data set, which requires convergence with a maximum of 20 iterations. 11 https://ravitejav.weebly.com/kbac.html 12 https://github.com/zhiwu-huang/LieNet 13 https://github.com/hjf1997/LieNet 256 Appendix A. Experimental Details and Additional Discussions Data Set Metric ECM LECM OLM LSM HDM050.750.50.5-0.5 FPHA-0.5-0.25-0.25-0.25 Table A.4: Matrix powers in LieBN-Cor under different metrics on each data set. A.3.2 Gyrogroup Batch Normalization A.3.2.1 GyroBN on the Grassmannian For HDM05, under-represented clips are removed, yielding 2,086 instances over 117 classes. For NTU60 and NTU120, we focus on mutual actions and adopt the cross-view and cross-setup protocols, respectively [178, 138]. Following Nguyen and Yang [159], each sequence is represented as a Grassmannian matrix of size 93× 10, 150× 10, and 150× 10 for HDM05, NTU60, and NTU120, respectively. Because GyroBN is inserted after the first pooling layer, the corresponding inputs to GyroBN have sizes 47× 10, 75× 10, and 75× 10, as reported in Tab. 3.16. The GyroGr backbone is reviewed in Sec. A.2.2. Trivialization. Following Nguyen and Yang [159], we apply the trivialization strat- egy reviewed in Sec. 2.7 to the Grassmannian parameters in the gyrotranslation and GyroBN layers. Each Grassmannian parameter U ∈ Gr(p,n) is parameterized by a matrix U∈ R (n−p)×p such that " 0 −U ⊤ U 0 # = h U ⊤ , e I p,n i ,(A.25) where (·) = Log e I p,n (·). The parameter U can be retrieved by U = exp h U ⊤ , e I p,n i I p,n = exp " 0 −U ⊤ U 0 #! I p,n .(A.26) This reparameterization represents the trainable Grassmannian parameters by Eu- clidean coordinates, thereby allowing the direct use of PyTorch optimizers [169] and avoiding direct Riemannian updates of these parameters. 257 A.3. Experimental Details A.3.3 Riemannian Multinomial Logistic Regression A.3.3.1 RMLR on the SPD Manifold This subsection offers additional details on the experiments on SPD MLRs. Backbone networks and LogEig MLR. The SPDNet, TSMNet, SPDNetBN, SPDDSMBN, and SPDGCN backbones are reviewed in Sec. A.2.1, while RResNet is reviewed in Sec. A.2.5. In the SPD baseline models considered here, the Euclidean MLR in the codomain of matrix logarithm (matrix logarithm + FC + softmax) is used for classification. Following the terminology introduced in Sec. 4.2.4, we call this classifier the LogEig MLR. The LogEig MLR is the Euclidean classifier in the tangent space at the identity, which might distort the innate geometry of the SPD manifold. SPD modeling and preprocessing. We use the same SPD modeling and pre- processing as the LieBN experiments in Sec. A.3.1.1. Implementation details. For SPDNet [106] and TSMNet [123], we follow the official PyTorch code of SPDNetBN 14 and TSMNet 15 to implement our experiments. To evaluate the performance of our intrinsic classifiers, we substitute the LogEig MLR in SPDNet and TSMNet with our SPD MLRs. We implement our SPD MLRs induced by five parameterized metrics. On the Radar and HDM05 data sets, the learning rate is 10 −2 , and the batch size is 30. On the Hinss2021 data set, following Kobler et al. [123], the learning rate is 10 −3 with a 10 −4 weight decay, and the batch size is 50. The maximum numbers of training epochs are 200, 200, and 50, respectively. We use the standard cross-entropy loss as the training objective and optimize the parameters with the Riemannian AMSGrad optimizer [18]. RResNet [117]. We focus on the AIM-based RResNet and use the official code 16 and suggested network settings to implement the experiments with RResNet. We con- duct 10-fold and 5-fold experiments on the HDM05 and NTU60 data sets, respectively. Since RResNet is developed based on SPDNet, we use the same learning settings as SPDNet for the action recognition task and borrow the best (θ,α,β) from Tab. 4.4 for our SPD MLRs under the RResNet backbone. NTU60 SPD modeling. For NTU60 [178], each frame contains the 3D coordinates of 25 body joints and is therefore represented by a 25× 3 = 75-dimensional coordinate vector. Following Katsman et al. [117], each sequence is modeled as a 75× 75 temporal covariance matrix, and evaluation follows the cross-view protocol [178]. 14 https://proceedings.neurips.c/paper_files/paper/2019/file/ 6e69ebbfad976d4637b4b39de261bf7-Supplemental.zip 15 https://github.com/rkobler/TSMNet 16 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 258 Appendix A. Experimental Details and Additional Discussions Data Sets(θ,α,β)-AIM(θ,α,β)-EM(α,β)-LEM2θ-BWM θ-LCM Disease(0.25,1,0)(0.25,1,0)(1,1)0.250.5 Cora(0.5,1,0)(0.25,1, 1 /9)(1, 1 /9)0.250.5 Pubmed(0.5,1,0)(0.5,1,0)(1,− 1 /3)0.250.5 Table A.5: (θ,α,β) of SPD MLRs on the SPDGCN backbone. SPDGCN [231]. We use the official code 17 and the suggested network settings in Zhao et al. [231]. Note that SPDGCN with SPD MLR retains the same network settings as vanilla SPDGCN. Tab. A.5 presents the hyperparameters (θ,α,β) on different data sets. Network architectures. We denote the network architecture as [d 0 ,d 1 ,· ,d L ], where the dimension of the parameter in the i-th BiMap layer (Sec. A.2.1) is d i ×d i−1 . For SPDNet, we also validate our SPD MLRs under different network architectures on the Radar and HDM05 data sets. The network architectures on the Radar data set are [20, 16, 8] for the 2-block configuration and [20, 16, 14, 12, 10, 8] for the 5-block configuration, while on the HDM05 data set, the network architectures are [93, 30] for 1-block, [93, 70, 30] for 2-block, and [93, 70, 50, 30] for 3-block. For TSMNet, the 1-block architecture is [40, 20]. Scoring metrics and evaluation protocols. We use the same scoring metrics and evaluation protocols as Sec. A.3.1.1. For the graph-learning experiments, following Zhao et al. [231], we report the 10-fold average and maximum node-classification accuracy. Hyperparameters. We implement the SPD MLRs induced by not only five stan- dard metrics, i.e., LEM, AIM, EM, LCM, and BWM, but also five families of parame- terized metrics. Therefore, in our SPD MLRs, we have a maximum of three hyperpa- rameters, i.e., θ,α,β, where (α,β) are associated with O(n)-invariance and θ controls deformation. For (α,β) in (θ,α,β)-LEM, (θ,α,β)-AIM, and (θ,α,β)-EM, recalling Eq. (2.93), α is a scaling factor, while β measures the relative significance of traces. As scaling is less important [192], we set α = 1. As for the value of β, we select it from a predefined set: 1, 1 /n, 1 /n 2 , 0,− 1 /n + ε,− 1 /n 2 , where n is the dimension of the input SPD matrices in SPD MLRs. The purpose of including ε ∈ R + is to ensure the positive definiteness of the inner product, i.e., α + nβ > 0 and hence (α,β) ∈ ST. These chosen values for β allow for amplifying, neutralizing, or suppressing the trace components, depending on the characteristics of the data sets. For the deformation factor θ, we roughly select its value around its deformation boundary, i.e., [0.25, 1.5] 17 https://github.com/andyweizhao/SPD4GNNs 259 A.3. Experimental Details Metric(θ,α,β)-AIM(θ,α,β)-EMθ-LCM2θ-BWM Candidate Values 0.25, 0.5, 0.75, 1, 1.25, 1.5 0.25, 0.5, 1, 1.5 0.5, 1, 1.5 0.25, 0.5, 0.75 Table A.6: Candidate values for hyperparameters in SPD MLRs. for (θ,α,β)-AIM, [0.5, 1.5] for θ-LCM, [0.25, 1.5] for (θ,α,β)-EM, and [0.25, 0.75] for 2θ-BWM. The detailed values are listed in Tab. A.6. A.3.3.2 RMLR on Rotation Matrices LieNet backbone. The LieNet backbone is reviewed in Sec. A.2.3. In the official MATLAB implementation, the LogMap layer uses the Euler axis–angle representation. Classification is performed using the Euler axis–angle representation, followed by an FC layer and a softmax layer. As the axis–angle representation is equivalent to the matrix logarithm, we call this classifier LogEig MLR as well. This classifier is, therefore, also non-intrinsic. Preprocessing. We use the shared rotation-matrix modeling, preprocessing, LieNet implementation, and SGD optimizer choice described in Sec. A.3.1.2. Following Huang et al. [107], the Lie MLR experiments use the G3D and HDM05 data sets. We trim HDM05 by removing under-represented sequences, resulting in 2,326 sequences across 122 classes. Lie MLR. We use our Lie MLR to replace the axis–angle classifier in LieNet and call the resulting network LieNet+LieMLR. To alleviate the computational burden, we set each P k to have shape [num, 3, 3], where num is the spatial dimension of the input of the Lie MLR layer. In other words, P k is shared in the temporal dimension. We adopt PyTorch3D [174] to calculate the matrix logarithm. Due to the instability of pytorch3d.transforms.so3_log_map, we first use pytorch3d.transforms.matrix_ to_axis_angle to calculate the rotation axis and angle and then convert this represen- tation into the matrix logarithm 18 . Training details. Following Huang et al. [107], we focus on the 3-block and 2- block architectures for the G3D and HDM05 data sets, respectively, which are the suggested architectures for these two data sets. The learning rate is 10 −2 on both data sets, and we further set the weight decay to 10 −5 on the G3D data set. For LieNet and LieNet+LieMLR, we use torch.n.utils.clip_grad_norm_ for gradient clipping with a clipping factor of 5. The clipping is imposed on the dimensionality reduction weight in the final FC linear layer of LieNet or, accordingly, A = A 1 ,...,A C in the 18 https://github.com/facebookresearch/pytorch3d/issues/188 260 Appendix A. Experimental Details and Additional Discussions Lie MLR layer of LieNet+LieMLR. Scoring metrics. For the G3D data set, following LieNet [107], we adopt a 10-fold cross-subject test setting, where half the subjects are used for training and the other half are employed for testing. For the HDM05 data set, following Huang et al. [107], we randomly select half of the sequences for training and the rest for testing. Due to the instability of LieNet, we conduct 20-fold experiments and select the best 10 folds to evaluate the performance. A.3.4 Proper Velocity Neural Networks Common Implementations. We use the trivialization strategy reviewed in Sec. 2.7 in our MLR, FC, and GyroBN layers. Consequently, all trainable manifold-valued parameters in PVNN are represented by Euclidean coordinates and optimized using standard Euclidean optimizers. A.3.4.1 Image Classification The CIFAR-10 and CIFAR-100 data sets and their PyTorch splits are described in Sec. A.1.3. Following Bdeir et al. [15], we use data augmentation that includes random cropping with padding of 4 pixels and random horizontal flipping. Implementation Details. We implement the experiments using the official code 19 of Bdeir et al. [15]. All models share a common backbone, which consists of a ResNet-18 encoder followed by a hyperbolic MLR classifier. Except for PV MLR without Exp 0 , the output embedding of the ResNet-18 backbone is mapped to the target hyperbolic space via the exponential map at the identity e, that is, Exp e (x). Here, e = 0 for the Poincaré and PV spaces, and e = 0 for the hyperboloid model. All models are trained from scratch. Optimization is performed using SGD [175] with an initial learning rate of 0.1, a momentum of 0.9, and a weight decay of 5× 10 −4 . Training is conducted with a batch size of 128 for 200 epochs. The learning rate is decayed by a factor of γ = 0.2 at epochs 60, 120, and 160. The curvature for the PV space is set as K =−0.5. A.3.4.2 Graph Learning The descriptions and summary statistics of the Disease, Airport, PubMed, and Cora data sets are collected in Sec. A.1.4. Implementation Details. We adopt the official code of HGCN 20 [38] to conduct 19 https://github.com/kschwethelm/HyperbolicCV 20 https://github.com/HazyResearch/hgcn 261 A.3. Experimental Details Hyperparameter Disease Airport PubMed Cora Learning rate0.010.010.050.05 Dropout0.40.40.60.6 Curvature-0.3-0.3-1.0-1.0 Table A.7: Hyperparameters for PVNN that vary across graph data sets. Setting Epochs Batch size Weight decay Value20001285× 10 −4 Table A.8: Hyperparameters for PVNN that are shared across graph data sets. experiments. The features of each node are embedded into the hyperbolic space via the exponential map at the identity. The hyperbolic network consists of two FC layers: the first maps the input feature dimension to a 16-dimensional hidden representation, and the second maps from 16 to 16. Each FC layer is followed by an activation function. An MLR layer is then used for classification. All models are trained using the Adam optimizer [122]. We evaluate performance every 10 epochs and employ early stopping with a patience of 200 evaluations, restoring the checkpoint with the best test accuracy. Tabs. A.7 and A.8 summarize the hyperparameters for PVNN. For KNN [147], HNN [76], HNN++ [181], and LNN [15], we follow their original papers to implement the experiments. Tab. A.9 summarizes the hyperbolic layers used in each model. A.3.4.3 Genomic Sequence Learning The description and complete summary of TEB are collected in Sec. A.1.5. We focus on five of its seven data sets. Implementation Details. For the Euclidean CNN and the hyperbolic CNN base- line (HCNN-S), we directly use the results reported in the original paper [119, Tab. 2]. Our PVCNN architecture follows their implementation 21 . Each DNA sequence is repre- sented as a length-L sequence with 4 input channels. We first apply a PV convolution that maps the 4 input channels to 32 channels, followed by PV TBN and a tanh tan- gent activation. A second PV convolution layer is then applied. The final PV feature is concatenated and passed through an FC layer, and finally classified with a PV MLR head. The curvature is initialized at K = −0.5 and learned during training. We train for 100 epochs with a step learning-rate schedule, using milestones at epochs 60 and 21 https://github.com/rrkhan/HGE 262 Appendix A. Experimental Details and Additional Discussions ModelFC layerActivationMLR PVNNPV FC in Thm. 127Exp 0 (σ (Log 0 (x)))PV MLR in Thm. 126 KNNLog 0 (W Exp 0 (x))Exp 0 (σ (Log 0 (x)))Euclidean MLR after Exp 0 HNNLog 0 (W Exp 0 (x))Exp 0 (σ (Log 0 (x)))Poincaré MLR [76] HNN++Poincaré FC [181]Exp 0 (σ (Log 0 (x)))Poincaré MLR [181] LNNLorentz FC [45]Lorentz activation [15]Lorentz MLR [15] Table A.9: Summary of the hyperbolic layers used in the graph node classification models. SettingValue OptimizerAdam Learning rate1e −4 Weight decay2e −2 Batch size100 Dropout0.1 Adam (β 1 ,β 2 )(0.9, 0.999) Table A.10: Hyperparameters for TEB. 85 with a decay factor of 0.1. For the PV FC layer, σ in Eq. (5.29) is set to tanh. All other hyperparameters are summarized in Tab. A.10. A.3.5 Hyperbolic Busemann Neural Networks A.3.5.1 Image Classification Implementation Details. For CIFAR-10/100 and Tiny-ImageNet, we follow Bdeir et al. [15, App. C.1]. For ImageNet-1k, we follow Guo et al. [91, Sec. 4]. Tab. A.11 summarizes the data-set-specific hyperparameters. For the hyperbolic MLR, before mapping into the hyperbolic space, we clip the feature vector by CLIP (x;r) = min 1, r ∥x∥ x(A.27) where r > 0 is a hyperparameter. The clipped Euclidean embedding is projected via the exponential map to the target hyperbolic space: Exp e (CLIP(x;r)). For the Lorentz model, the clipping parameter is r = 1 on CIFAR-10/100 and r = 4 on Tiny-ImageNet and ImageNet-1k. On the Poincaré ball, r = 1 on all four data sets. All methods are implemented in PyTorch and trained with cross-entropy loss. The results of MLR, PMLR, and LMLR on CIFAR-10/100 and Tiny-ImageNet are copied 263 A.3. Experimental Details Hyperparameter CIFAR-10/100 Tiny-ImageNet ImageNet-1k Epochs200100 Batch size128256 Initial learning rate0.10.1 LR schedule60, 120, 160;γ = 0.230, 60, 90;γ = 0.1 Weight decay5e −4 1e −4 OptimizerSGDSGD Precision32-bit32-bit Curvature K−1−1 Table A.11: Summary of hyperparameters used in the image classification task. from Bdeir et al. [15, Tab. 1], while those of PBMLR-P on CIFAR-10/100 are copied from Nguyen et al. [162, Tab. 2]. The remaining results are obtained by our careful implementation. A.3.5.2 Genome Sequence Learning Implementation Details. We mainly follow the official implementations of Khan et al. [119] for data processing, model architecture, and training. We adopt a simple CNN with three convolutional blocks followed by dense, ReLU-activated layers to ex- tract features [119, Fig. 4]. Before classification, the features are clipped and mapped to the target hyperbolic space as in Eq. (A.27), then passed to either a prior hyperbolic MLR or our BMLR head. The clipping factor defaults to r = 1, with the following exceptions on GUE: on the Lorentz model, r = 2.0 for Covid variant classification and r = 5.0 for species classification; on the Poincaré ball, r = 2.0 for species classification. Following Khan et al. [119], we treat the curvature K as a learnable parameter initial- ized as −1. The remaining hyperparameters are listed in Tab. A.12. All models share these hyperparameters, except LMLR on Covid variant classification, for which we set the weight decay to 1e −3 to ensure convergence. All methods are implemented in PyTorch and trained with cross-entropy loss. Re- sults are obtained from our reimplementation. A.3.5.3 Node Classification Implementation Details. We follow the official implementations of HGCN [38] and PBMLR [162] and conduct experiments on both the Poincaré ball and the Lorentz model. We adhere to their experimental settings. The only changes are weight decay 264 Appendix A. Experimental Details and Additional Discussions Batch size100 Epochs100 OptimizerAdam β 1 , β 2 0.9, 0.999 Initial learning rate1e −4 LR schedule60, 85; γ = 0.1 Weight decay0.1 Initial curvature K−1 Table A.12: Hyperparameters for genome sequence learning. SpaceWeight decayDropout P n K 1e− 4, 1e− 5, 1e− 3, 1e− 30.3, 0, 0, 0.2 L n K 1e− 4, 5e− 5, 1e− 3, 1e− 30, 0, 0, 0.3 Table A.13: Hyperparameters for node classification on Disease, Airport, PubMed, and Cora. and dropout. We train with cross-entropy loss and the Adam optimizer [122] for 5000 epochs with a learning rate of 1e −2 , curvature set to −1, embedding dimension 16, and three GCN layers. We tune weight decay and dropout and report the values in Tab. A.13. All methods are implemented in PyTorch. The results of HGCN-PMLR and HGCN-PBMLR-P on the Poincaré ball are taken from Nguyen et al. [162, Tab. 11]. Results for the remaining baselines are obtained from our reimplementation following the original settings. A.3.5.4 Link Prediction Implementation Details. We follow the official implementations of HNN [76], HNN++ [181], and HyboNet [45], and adopt the experimental protocol of Chami et al. [38] for link prediction. The encoder consists of two fully connected layers: the first maps the input features to 16, and the second maps 16 to 16. Each FC layer is instantiated as either our BFC or an existing hyperbolic FC layer. After each FC, we apply the activation Exp e (ReLU(Log e (x))), where e denotes the model origin. Following Chami et al. [38], this activation is disabled on Cora. As in the Möbius layer, we apply a gyro bias after each FC, that is, x⊕ H b. We train with Adam [122] at a learning rate of 1e −2 and tune weight decay and FC dropout. For BFC, we use φ = tanh on Airport and Cora and the identity map on the other two data sets. 265 A.3. Experimental Details A.3.6 Full-Rank Correlation Networks The Radar, HDM05, FPHA, and NTU120 data sets are described in Sec. A.1. For HDM05 and FPHA, each sequence is normalized for body-part length, scale, and view using the preprocessing of Vemulapalli et al. [203]. For NTU120, we follow Chen et al. [46] to preprocess the data. A.3.6.1 Input Data Correlation Input in CorNets. Following Wang et al. [210], Nguyen et al. [161], we first model each sample as a multi-channel SPD tensor. For HDM05 and FPHA, the skeleton preprocessing and multi-channel covariance construction follow the GyroGr pipeline of Nguyen and Yang [159]; CorNet subsequently converts every covariance matrix into a full-rank correlation matrix by Cor :S n ++ ∋ Σ7−→ C = D(Σ) − 1 2 ΣD(Σ) − 1 2 ∈ Cor + (n).(A.28) For Radar, we follow Wang et al. [210] and use temporal convolution followed by co- variance pooling to obtain a multi-channel covariance tensor of shape [c, 20, 20]. After preprocessing, the input correlation tensor shapes are [7, 20, 20], [3, 28, 28], [9, 28, 28], and [6, 28, 28] on Radar, HDM05, FPHA, and NTU120, respectively. A.3.6.2 Implementation Details SPD Baselines. We follow the official PyTorch code of SPDNetBN 22 to implement SPDNet and SPDNetBN. For LieBN 23 , we focus on the instantiation under LCM [137], while for RResNet 24 , we implement the ones induced by AIM [171] and LEM [9]. For SPD MLR 25 , we implement the one based on LCM. Due to the lack of official code, Gyro-based models are carefully reimplemented from their original papers. Following Nguyen et al. [161], GyroSPD++ combines an AIM-based convolution with an LEM- based MLR. Grassmannian Baselines. Since GrNet is officially implemented in MATLAB, we carefully re-implemented it using PyTorch. Additionally, as both GyroGr and GyroGr- Scaling do not release official code, we re-implemented them based on the original paper 22 https://proceedings.neurips.c/paper_files/paper/2019/file/ 6e69ebbfad976d4637b4b39de261bf7-Supplemental.zip 23 https://github.com/GitZH-Chen/LieBN 24 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 25 https://github.com/GitZH-Chen/SPDMLR 266 Appendix A. Experimental Details and Additional Discussions Data SetModelOptimizerlrwdMatrix Power Converged Epoch Radar CorNet-ECMAdam1e −2 N/A1.550 CorNet-LECMAdam1e −2 N/A-0.2550 CorNet-OLMAdam1e −2 N/A-0.2550 CorNet-LSMAdam1e −2 N/A0.7550 CorNet-PHCMAdam1e −2 N/A0.7550 HDM05 CorNet-ECMAdam1e −3 1e −3 0.125100 CorNet-LECMAdam1e −4 1e −3 0.5150 CorNet-OLMSGD5e −2 1e −3 0.25200 CorNet-LSMAdam1e −3 N/A-0.7550 CorNet-PHCMAdam1e −2 N/A-0.2550 FPHA CorNet-ECMAdam5e −3 N/A-0.25150 CorNet-LECMAdam5e −4 1e −4 -0.5150 CorNet-OLMAdam1e −4 N/A-150 CorNet-LSMAdam1e −3 N/A-150 CorNet-PHCMAdam1e −3 1e −4 -0.5150 NTU120 CorNet-ECMSGD1e −2 N/A0.2550 CorNet-LECMSGD1e −2 N/A0.2550 CorNet-OLMSGD5e −3 N/A0.2550 CorNet-LSMSGD1e −3 N/A0.2550 CorNet-PHCMAdam1e −3 N/A0.2550 Table A.14: Hyperparameters in CorNets. [159]. For all Grassmannian comparative methods, we use SGD [175] with a learning rate of 5e −2 . CorNets. On all four data sets, we employ a single convolutional kernel for global convolution, i.e., applying a global receptive field across the channel dimension. The output dimensions of the correlation convolutional layer are 8× 8, 26× 26, 26× 26, and 11× 11 for the Radar, HDM05, FPHA, and NTU120 data sets, respectively. We primarily use the Adam [122] and SGD [175] optimizers. Inspired by the defor- mation effect on the latent SPD geometries by the matrix power over the SPD manifold illustrated in Fig. 4.2, we apply the matrix power before correlation modeling (Cor(·)) as activation. In particular, when the data are centered at zero and power is−1, Cor(Σ −1 ) corresponds to the partial correlation matrix of the covariance matrix Σ [191, Lem. 1.6]. The batch size is set to 30, and training is capped at 200 epochs, although most cases converge in fewer than 150 epochs. Due to the different correlation geometries, the hyperparameters vary for CorNets under different geometries. Tab. A.14 summarizes all the hyperparameters. Extra Computational Details for OLM and LSM Layers. For the MLR, FC, and convolutional layers induced by OLM and LSM, the key computations involve Exp ◦ and Log ⋆ , which depend on the calculations of D and D ⋆ . In our experiments, we empirically observe that iterating until convergence is more effective for D, whereas a single step of Newton’s method generally performs best for D ⋆ . Accordingly, we set D to iterate until convergence, leveraging Thm. 145 for accurate backpropagation. For 267 A.3. Experimental Details D ⋆ , we adopt a single iteration in Newton’s method and use automatic differentiation (autograd) through this single step for backpropagation. A.3.6.3 Additional Details on Visualization We provide additional interpretations of the low-dimensional visualization and clarify how the decision hyperplanes in Fig. 5.5 are obtained. The correlation-in-SPD visual- ization is discussed in Sec. 2.9.2. Visualization of Low-Dimensional SPD and Correlation Matrices. Any 2× 2 covariance matrix in S 2 ++ can be written as Σ = a b b d ! , a > 0, d > 0, ad− b 2 > 0.(A.29) Embedding Σ into R 3 via the map Σ7→ (a,b,d) identifies S 2 ++ with the interior of the quadratic cone (a,b,d)∈ R 3 | a > 0, d > 0, ad− b 2 > 0 ,(A.30) which is an open cone in R 3 . For 2× 2 correlation matrices, any C ∈ Cor + (2) has the form C = 1 r r 1 ! , r ∈ (−1, 1).(A.31) Thus, Cor + (2) is one-dimensional. Embedding C into R 3 as (1,r, 1) yields a line segment inside the cone corresponding to S 2 ++ . For 3× 3 correlation matrices, any C ∈ Cor + (3) is parameterized by its off-diagonal entries (r 12 ,r 13 ,r 23 ): C = 1 r 12 r 13 r 12 1 r 23 r 13 r 23 1 .(A.32) Embedding C into R 3 via C 7→ (r 12 ,r 13 ,r 23 ) produces an open elliptope in R 3 . This is the representation of Cor + (3) used in Fig. 5.5, where each point in the elliptope corresponds to one 3× 3 correlation matrix. Construction of Fig. 5.5. For ECM, LECM, OLM, and LSM, the decision hyper- plane in the correlation MLR is the Riemannian hyperplane in Sec. 4.3.2.1 specialized 268 Appendix A. Experimental Details and Additional Discussions to M = Cor + (n): H A,P = X ∈ Cor + (n)|⟨Log P (X),A⟩ P = 0 , P ∈ Cor + (n), A∈ T P Cor + (n). (A.33) In Fig. 5.5, we focus on Cor + (3) and visualize it as the open elliptope in R 3 via the embedding C ∈ Cor + (3)7−→ (C 21 ,C 31 ,C 32 )∈ R 3 .(A.34) Given a Log-Euclidean metric and parameters (A,P ), each correlation matrix C is first mapped to the tangent space at P by Log P (C), and we evaluate the linear form ⟨Log P (C),A⟩ P . The set of points in the elliptope where this scalar equals zero corre- sponds to the decision hyperplane H A,P and is plotted as the separating surface. For PHCM, the margin hyperplane is defined in the β-concatenated Poincaré em- bedding. Let Φ be the diffeomorphism in Eq. (5.79) that maps C ∈ Cor + (n) to the poly- Poincaré space P n−1 , and let β-concatenation be the Poincaré operation in Sec. 5.4.3.2. We define ̃x(X) = β-concat (Φ(X))∈ P N , N = n(n− 1) 2 ,(A.35) and the PHCM hyperplane H a,p = n X ∈ Cor + (n)| Log p ( ̃x(X)),a p = 0 o , p∈ P N , a∈ T p P N . (A.36) Here P N = x∈ R N |∥x∥ 2 < 1 is the N-dimensional unit Poincaré ball. In Fig. 5.5, we first map each correlation matrix C ∈ Cor + (3) to ̃x(C) ∈ P N , apply the Poincaré logarithm Log p at a reference point p, and then visualize the zero level set of the linear form Log p ( ̃x(C)),a p as the PHCM decision hyperplane. A.3.7 Adaptive Log-Euclidean Metrics A.3.7.1 Data Sets and Settings Although the proposed ALog layers can be plugged into existing SPD networks, we focus on the SPDNet framework [106]. We follow the PyTorch code provided by SPDNetBN 26 to reproduce SPDNet and SPDNetBN and implement our approaches. Following previous work [106, 31], we evaluate our methods on HDM05 [153], FPHA [80], and AFEW [68]. The data set descriptions are given in Sec. A.1. We use the HDM05 and FPHA temporal covariance representations and the filtered HDM05 setting 26 SPDNetBN supplementary code 269 A.3. Experimental Details specified in Sec. A.3.1.1. For AFEW, we use the released pre-trained FAN 27 [149] to extract deep features and establish a 512× 512 temporal covariance matrix for each video. We denote the dimensions of the transformation layers in the SPDNet backbone by d 0 ,d 1 ,...,d L . Following the settings in Brooks et al. [31], all networks are trained using the default RSGD [17] with a fixed learning rate γ and a batch size of 30. To make ALog start from the vanilla matrix logarithm, the parameters in MUL, DIV, and RELU are initialized as 1, 1, and e, respectively. By abuse of notation, SPDNet-ALog-MUL is abbreviated as ALog-MUL, denoting that we substitute the LogEig layer in SPDNet with the proposed ALog optimized by MUL. A.3.7.2 Implementation Details of Additional Applications NTU60 data set. For NTU60, we use the cross-view protocol and 75× 75 covariance representation specified in Sec. A.3.3.1. As reported in Sec. 6.2.4, MUL shows the best performance. Therefore, we view A in Eq. (6.31) as the parameter for all experiments. In the following, we discuss in detail the specific implementation of each method. LieBN. We follow the official code 28 to implement the experiments. The learning rate is 5e −2 . Since our LieBN-ALEM shows early convergence, we set the numbers of training epochs to 150, 50, and 30 for the [93, 30], [93, 70, 30], and [93, 70, 50, 30] architectures. Other settings are the same as Sec. 6.2.4. RResNet. We follow the official code 29 to implement the experiments. For the HDM05 data set, we use RSGD [17] with a 5e −2 learning rate for 200 training epochs. For the NTU60 data set, we use Riemannian AMSGrad [17] with a 1e −2 learning rate for 50 training epochs. We adopt the architectures of [93, 30] and [75, 30] on these two data sets. Gyro MLR. Since the code of gyro MLR is not publicly available, we carefully re-implement the gyro MLR in Nguyen and Yang [159]. We adopt an architecture of [75, 30] under an SGD optimizer. The batch size and number of training epochs are 30 and 200, respectively. 27 https://github.com/Open-Debin/Emotion-FAN 28 https://github.com/GitZH-Chen/LieBN 29 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 270 Appendix A. Experimental Details and Additional Discussions A.3.8 Product Cholesky Metrics A.3.8.1 Implementation Details of SPD Neural Networks Backbone networks. SPDNet and GyroSPD are reviewed in Sec. A.2.1, while RRes- Net is reviewed in Sec. A.2.5. Data sets and preprocessing. The Radar, HDM05, and FPHA data sets are described in Sec. A.1. The global covariance representations for SPDNet and RResNet follow Sec. A.3.1.1; the GyroSPD-specific multichannel construction is given below. SPD MLR. We follow the official PyTorch code 30 to implement the SPD MLR developed in Sec. 4.2.2. Due to the lack of official code, the GyroSPD backbone is carefully reimplemented in PyTorch following the original paper [159]. For simplicity, we set M in (θ, M)-BWCM to the identity matrix. SPDNet. Following Huang and Van Gool [106], we replace the vanilla tangent classifier (LogEig + FC + softmax) in SPDNet with SPD MLRs induced by AIM, LEM, LCM, θ-PCM, and θ-BWCM. We use a Riemannian AMSGrad [17] with a learning rate of 1e −2 , a batch size of 30, and a maximum of 200 epochs. We denote the architectures by [d 0 ,d 1 ,...,d L ], where d i is the output dimension of the i-th BiMap layer. Following prior work [106, 31], we adopt [20, 16, 8] on Radar, [63, 33] on FPHA, and [93, 30], [93, 70, 30], [93, 70, 50, 30] for 1-, 2-, and 3-block variants on HDM05. As shown in Tab. 3.9c, matrix power improves the performance of Cholesky-based metrics on FPHA. We therefore apply a power of −0.25 before SPD MLR layers under LCM, θ-PCM, and θ-BWCM. For θ-PCM on FPHA, we further adopt a weight decay of 1e −4 . The deformation factor θ is reported in Tab. A.15. GyroSPD. Following Nguyen and Yang [159], the backbone consists of one gyro- translation layer followed by an SPD MLR. We compare metrics based on LEM, LCM, AIM, and our proposed geometries under the same settings: Riemannian AMSGrad with a learning rate of 1e −2 and a batch size of 30. The models are trained for up to 100, 100, and 50 epochs on Radar, HDM05, and FPHA, respectively. Similarly, a matrix power of 0.25 is applied on FPHA for LCM and our Cholesky-based metrics. The deformation factor θ is reported in Tab. A.15. SPD input of SPDNet. We use the global covariance representations specified in Sec. A.3.1.1. SPD input of GyroSPD. For Radar, the input is the same as SPDNet. For HDM05 and FPHA, we follow Nguyen and Yang [159] to model each sample into a multi-channel covariance tensor [c,n,n]. Specifically, we first identify the closest left 30 https://github.com/GitZH-Chen/SPDMLR 271 A.3. Experimental Details BackboneMetricRadar HDM05 FPHA SPDNet θ-PCM-1.5-0.50.75 θ-BWCM-0.75-1.51 GyroSPD θ-PCM-0.75-0.75-0.75 θ-BWCM-1.5-1.5-0.5 Table A.15: Hyperparameter θ in θ-PCM and θ-BWCM. It is selected from the candi- date values used for SPDNet and GyroSPD. (right) neighbor of each joint based on its distance to the hip (wrist) joint, and then combine the 3D coordinates of each joint and those of its left (right) neighbor to create a feature vector for the joint. For a given frame t, we compute its Gaussian embedding [143]: Y t = (det Σ t ) − 1 n+1 " Σ t + μ t (μ t ) ⊤ μ t (μ t ) ⊤ 1 # ,(A.37) where μ t and Σ t are the mean vector and covariance matrix computed from the set of feature vectors within the frame. The lower part of the matrix log (Y t ) is flattened to obtain a vector ̃v t . All vectors ̃v t within a time window [t,t + c− 1], where c is determined from a temporal pyramid representation of the sequence (the number of temporal pyramids is set to 2 in our experiments), are used to compute a covariance matrix as e Σ t = 1 c t+c−1 X i=t ( ̃v i − v t ) ( ̃v i −v t ) ⊤ ,(A.38) wherev t = 1 c P t+c−1 i=t ̃v i . The resulting e Σ t are the covariance matrices that we need. On FPHA, we generate the covariances based on three sets of neighbors: left, right, and vertical (bottom) neighbors. After preprocessing, the input covariance matrices are [3, 28, 28] and [9, 28, 28] on HDM05 and FPHA, respectively. RResNet implementation. The backbone architectures of RResNet [118] are similar to those of SPDNet: [20, 16, 8] on Radar, [93, 30] on HDM05, and [63, 33] on FPHA, respectively. The only architectural difference lies in the head: SPDNet directly uses a classification layer, whereas RResNet attaches a residual block before the clas- sification head. Following RResNet, we adopt Log I n + FC + softmax for classification under each metric. We use the official code 31 to re-implement the AIM- and LEM-based RResNet. For the RResNets based on LCM and our θ-PCM, we adopt the following settings: a learning rate of 1e −2 , a batch size of 30, and a maximum of 200 epochs. 31 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 272 Appendix A. Experimental Details and Additional Discussions We use the standard cross-entropy loss as the training objective and optimize the pa- rameters with the Riemannian AMSGrad optimizer [17]. The Cholesky diagonal power is set to −1, 0.5, and −0.5 on the Radar, HDM05, and FPHA data sets, respectively. Additionally, on the FPHA data set, a matrix power of −0.25 is applied before the residual blocks to activate the latent geometry for both LCM and our metric. SPD input. The input SPD matrices are the same as those used in SPDNet. A.4 Additional Discussions A.4.1 Riemannian Multinomial Logistic Regression A.4.1.1 RMLR as a Natural Extension of the Euclidean MLR Proposition 180. When M = R n is the standard Euclidean space, the RMLR defined in Thm. 114 becomes the Euclidean MLR in Eq. (4.1). Proof. On the standard Euclidean space R n , Log y x = x − y,∀x,y ∈ R n . Besides, the differential maps of left translation and parallel transport are the identity maps. Therefore, given x,p k ∈ R n and a k ∈ R n \0 ∼ = T 0 R n \0, we have p(y = k | x∈ R n )∝ exp ⟨Log p k x,a k ⟩ p k ,(A.39) ∝ exp (⟨x− p k ,a k ⟩),(A.40) ∝ exp (⟨x,a k ⟩− b k ),(A.41) where b k =⟨p k ,a k ⟩. A.4.1.2 Gyro SPSD MLR as a Special Case of Our RMLR Gyro SPSD MLR [161] is derived from the product of the Grassmannian and SPD gyro spaces. This section will show that the gyro SPSD MLR is a special case of our RMLR on the product geometry of the SPSD manifold. We first review some necessary results about gyro SPSD MLR and then show the equivalence. Following the notations in Nguyen et al. [161], we denote the Grassmannian with canonical metric under the projector and ONB perspective as f Gr(p,n) and Gr(p,n), respectively. The space of n× n SPSD matrices with a fixed rank p, denoted as S + n,p , forms an SPSD manifold [26]. As shown in Bonnabel et al. [26], Nguyen et al. [161], the SPSD manifold is a product space, i.e., S + n,p ∼ = Gr(p,n)×S p ++ . In other words, every 273 A.4. Additional Discussions P ∈ S + n,p can be decomposed as P = U P S P U ⊤ P with U P ∈ Gr(p,n) and S P ∈ S p ++ . We further denote S p,g ++ as the SPD manifold with metric g, where g could be AIM, LEM, or LCM. As shown in Nguyen et al. [161], the gyro space in S + n,p can be defined by the product of gyro spaces of Gr(p,n) and S p,g ++ . By this product structure, Nguyen et al. [161] proposed the SPSD Pseudo-gyrodistance to a hyperplane. Definition 181 (SPSD hypergyroplanes [161]). Let P,W ∈ Gr(p,n)×S p,g ++ . Then hypergyroplanes in the structure space Gr(p,n)×S p,g ++ are defined as H psd,g W,P = n Q∈ Gr(p,n)×S p,g ++ |⟨⊖ psd,g P ⊕ psd,g Q,W⟩ psd,g = 0 o ,(A.42) where⊕ psd,g and⟨·,·⟩ psd,g are the gyro addition and gyro inner product, respectively, as defined in Nguyen et al. [161]. Theorem 182 (SPSD pseudo-gyrodistance [161]). Let W = (U W ,S W ), P = (U P ,S P ), X = (U X ,S X )∈ Gr(p,n)×S p,g ++ , (A.43) and let H psd,g W,P be a hypergyroplane in the structure space Gr(p,n)×S p,g ++ . Then the pseudo-gyrodistance from X to H psd,g W,P is given by ̄ d X,H psd,g W,P = λ D e ⊖ gr U P e ⊕ gr U X e ⊖ gr U P e ⊕ gr U X ⊤ ,U W U ⊤ W E gr +⟨⊖ g S P ⊕ g S X ,S W ⟩ g q λ U W U ⊤ W gr 2 + (∥S W ∥ g ) 2 ,(A.44) where ∥·∥ gr and ∥·∥ g are the gyro norms on the Grassmann and SPD manifolds [161], and ⟨·,·⟩ gr and ⟨·,·⟩ g are gyro inner products [161]. e ⊕ gr and ⊕ g are gyro additions on Gr(p,n) and S p,g ++ . Denoting g gr as the canonical metric on Gr(p,n) and g as AIM, LEM, or LCM, we can prove that Thm. 182 is a special case of Thm. 113. Theorem 183. Under the product metric g psd,g = λg gr × g, the Riemannian hy- perplane in Eq. (4.19) on the SPSD manifold equals the SPSD hypergyroplane in Thm. 181. Similarly, the Riemannian margin distance in Thm. 113 on the SPSD manifold equals the SPSD Pseudo-gyrodistance in Thm. 182. Proof. Following the notations in Thms. 181 and 182, we further denote P = U P S P U ⊤ P , 274 Appendix A. Experimental Details and Additional Discussions Q = U Q S Q U ⊤ Q , W = U W S W U ⊤ W , and X = U X S X U ⊤ X with U P ,U Q ,U W ,U X ∈ Gr(p,n) and S P ,S Q ,S W ,S X ∈ S p,g ++ . I p is the p × p identity matrix. I p,n = (I p , 0) ⊤ is the gyro identity on Gr(p,n). e Γ gr , g Log gr , and⟨·,·⟩ gr U P denote Riemannian parallel transport along a geodesic, the Riemannian logarithm, and the Riemannian metric at U P on Gr(p,n), respectively. Γ g , Log g , and⟨·,·⟩ g S P denote Riemannian parallel transport along a geodesic, the Riemannian logarithm, and the Riemannian metric at S P on S p,g ++ , respectively. Γ psd,g , Log psd,g , and ⟨·,·⟩ psd,g X denote Riemannian parallel transport along a geodesic, the Riemannian logarithm, and the Riemannian metric at X on S + n,p ∼ = Gr(p,n)×S p ++ , respectively. First, we show that the SPSD hypergyroplane equals our Riemannian hyperplane in Eq. (4.19). We have the following: ⟨⊖ psd,g P ⊕ psd,g Q,W⟩ psd,g (1) = λ⟨⊖ gr U P ⊕ gr U Q ,U W ⟩ gr +⟨⊖ g S P ⊕ g S Q ,S W ⟩ g (2) = λ⟨Log gr U P U Q , ̃ A U W ⟩ gr U P +⟨Log g S P S Q , ̃ A S W ⟩ g S P (3) = ⟨Log psd,g P Q, ̃ A⟩ psd,g P (A.45) where ̃ A U W = e Γ gr I p,n →U P g Log gr I p,n (U W ) ,(A.46) ̃ A S W = Γ g I p →S P Log g I p (S W ) ,(A.47) ̃ A = ( ̃ A U W , ̃ A S W )∈ T P S + n,p ∼ = T U P Gr(p,n)× T S P S p,g ++ ,(A.48) where× denotes the Cartesian product. The above derivation comes from the following. (1) The definition of gyro addition, gyro inverse, and gyro inner product on the SPSD manifold [161, Sec. 3.3]. (2) The proof by Nguyen et al. [161, Prop. 3.2] indicates that similar results also hold on the Grassmannian. Combining that proposition with its Grassmannian counter- parts yields the equation. (3) The Riemannian product geometry. By the product geometry of the SPSD manifold, we can immediately get ̃ A = ( ̃ A U W , ̃ A S W ) = Γ psd,g e I p,n →P Log psd,g e I p,n (W ) ,(A.49) 275 A.4. Additional Discussions where e I p,n = I p,n I ⊤ p,n is the gyro identity on the SPSD manifold. Next, we show the equivalence between SPSD pseudo-gyrodistance and our Rieman- nian margin distance: ̄ d X,H psd,g W,P (1) = ⟨⊖ psd,g P ⊕ psd,g X,W⟩ psd,g ∥W∥ psd,g (2) = ⟨Log psd,g P X, ̃ A⟩ psd,g P ∥W∥ psd,g (3) = ⟨Log psd,g P X, ̃ A⟩ psd,g P Log psd,g e I p,n (W ) psd,g e I p,n (4) = ⟨Log psd,g P X, ̃ A⟩ psd,g P Γ psd,g e I p,n →P Log psd,g e I p,n (W ) psd,g P (5) = ⟨Log psd,g P X, ̃ A⟩ psd,g P ̃ A psd,g P (6) = d(X, ̃ H ̃ A,P ) (A.50) (1) The definition of gyro addition, gyro inverse, gyro inner product, and gyro norm on the SPSD manifold. (2) Eq. (A.45). (3) The definition of SPSD gyro norm [161]. (4) Riemannian parallel transport is norm-preserving [69, Def. 3.1]. (5) Eq. (A.49). (6) Thm. 113. Remark 184. We make the following remarks regarding the gyro MLR and our MLR on the SPSD manifold. (1) Eq. (A.49) indicates that when generating ̃ A in our RMLR by parallel trans- porting a tangent vector A ∈ T e I p,n S + n,p , ̃ A is the initial velocity of W in Eq. (A.44). 276 Appendix A. Experimental Details and Additional Discussions (2) Putting the pseudo-gyrodistance and Riemannian margin distance into Eq. (4.18) yields the gyro MLR and our Riemannian MLR. Therefore, Thm. 183 indicates the equivalence of the gyro MLR with our RMLR on the SPSD man- ifold. (3) As a metric g is required to induce gyro-structures, the metric g in gyro SPSD MLR is confined to AIM, LEM, and LCM. However, our SPSD MLR can be defined on the product space of the Grassmannian and SPD manifold under other metrics, such as BWM and PEM, as our framework does not require gyro-structures. A.4.1.3 Theories on the Deformed Metrics Limiting cases of the deformed metrics. Thanwerdas and Pennec [192] gener- alized (α,β)-AIM to a three-parameter family of metrics by power deformation, i.e., (θ,α,β)-AIM. The family of (θ,α,β)-AIM comprises (α,β)-AIM for θ = 1 and ap- proaches (α,β)-LEM with θ → 0 [192]. As reviewed in Sec. 3.2.5.1, LCM and (α,β)-LEM are extended to power-deformed metrics, denoted as (θ,α,β)-LEM and θ-LCM. Thm. 67 shows that (θ,α,β)-LEM is equal to (α,β)-LEM, and θ-LCM interpolates between ̃g-LEM (θ → 0) and LCM (θ = 1), with ̃g-LEM defined as ⟨V 1 ,V 2 ⟩ P = 1 2 ⟨ e V 1 , e V 2 ⟩− 1 4 ⟨D( e V 1 ), D( e V 2 )⟩,∀V i ∈ T P S n ++ ,(A.51) where e V i = log ∗,P (V i ) with log ∗,P as the differential map of matrix logarithm, and D( e V i ) is a diagonal matrix consisting of the diagonal elements of e V i . Thanwerdas and Pennec [194] identified the Alpha-Procrustes metric [151] with power-deformed BWM, denoted as 2θ-BWM. Similarly, 2θ-BWM becomes BWM with θ = 0.5 [194]. We further show the limiting case of 2θ-BWM under θ → 0. Proposition 185. 2θ-BWM tends to ( 1 4 , 0)-LEM as θ → 0. Before starting the proof, we first recall a well-known property of deformed metrics [194]. Lemma 186. Let 1 θ 2 φ ∗ θ g be the deformed metric on SPD manifolds pulled back from g by the matrix power φ θ and scaled by 1 θ 2 . Then, as θ tends to 0, for all P ∈S n ++ 277 A.4. Additional Discussions and all V ∈ T P S n ++ , we have ( 1 θ 2 φ ∗ θ g) P (V,V )→ g I (log ∗,P (V ), log ∗,P (V )).(A.52) Now, we present our proof for the limiting cases of deformed metrics. Proof of Thm. 185. First, we have g BWM I (V,V ) = 1 4 ⟨V,V⟩.(A.53) By Thm. 186, we have the following: g 2θ-BWM P (V,V ) θ→0 −→ g BWM I log ∗,P (V ), log ∗,P (V ) = 1 4 ⟨log ∗,P (V ), log ∗,P (V )⟩ = g ( 1 4 ,0)-LEM P (V,V ). (A.54) Proof of the properties of the deformed metrics (Tab. 4.2). In this subsec- tion, we prove the properties presented in Tab. 4.2. We first present a useful lemma and then present our detailed proof. This lemma will be useful in the proof of our SPD MLRs as well. Lemma 187. Given a Riemannian manifold (M,g) and a positive real scalar a > 0, the scaled metric ag on M has the same Riemannian logarithmic maps, exponential maps, and parallel transport as g. Proof. Since the Christoffel symbols of ag are identical to those of g, the geodesics and parallel transport under both ag and g remain unchanged. The equivalence of geodesics implies that the Riemannian exponential maps are the same for ag and g. As the inverse of the Riemannian exponential maps, the Riemannian logarithmic maps under ag and g are also identical. According to Thm. 187, geodesic completeness is independent of the scaling factor a > 0. By the definition of O(n)-, left-, right-, and bi-invariance, these invariant properties are also independent of the scaling factor a > 0. Without loss of generality, we will omit the scaling factor in the following proof. 278 Appendix A. Experimental Details and Additional Discussions Proof. First, we prove O(n)-invariance of (θ,α,β)-LEM, (θ,α,β)-EM, (θ,α,β)-AIM, and 2θ-BWM. Since the differential of φ θ is O(n)-equivariant, and (α,β)-LEM, (α,β)- EM, (α,β)-AIM, and BWM are O(n)-invariant [196], O(n)-invariance is thus acquired. Next, we focus on geodesic completeness. It can be easily proven that Riemannian isometries preserve geodesic completeness. On the other hand, (α,β)-LEM, (α,β)- AIM, and LCM are geodesically complete [196, 137]. As a direct corollary, geodesic completeness can be obtained since φ θ is a Riemannian isometry. Finally, we deal with Lie group invariance. Similarly, it can be readily proved that Lie group invariance is preserved under isometries. LCM, LEM, and (α,β)-AIM are Lie group bi-invariant [137], bi-invariant [9], and left-invariant [195]. As an isometric pullback metric from the standard LEM [196], (α,β)-LEM is, therefore, Lie group bi- invariant. As pullback metrics, (θ,α,β)-LEM, (θ,α,β)-AIM, and θ-LCM are therefore bi-invariant, left-invariant, and bi-invariant, respectively. A.4.1.4 Computational Details on the SPD MLR under Power-Deformed BWM Matrix square roots in the SPD MLR under power-deformed BWM. In the case of MLRs induced by 2θ-BWM, computing square roots like (BA) 1 2 and (AB) 1 2 with B,A ∈ S n ++ poses a challenge. Eigendecomposition cannot be directly applied since BA and AB are no longer symmetric, let alone positive definite. Instead, we use the following formulas to compute these square roots [151]: (BA) 1 2 = B 1 2 (B 1 2 AB 1 2 ) 1 2 B − 1 2 and (AB) 1 2 = [(BA) 1 2 ] ⊤ ,(A.55) where the involved square roots can be computed using eigendecomposition or singular value decomposition (SVD). Numerical stability of the SPD MLR under power-deformed BWM. Let us first explain why we abandon parallel transport for the SPD MLR derived from 2θ- BWM. Then, we propose our numerically stable methods for computing the SPD MLR based on 2θ-BWM. Instability of parallel transport under power-deformed BWM. As discussed in Thm. 114, there are two ways to generate ̃ A in SPD MLR: parallel transport and Lie group translation. However, parallel transport under 2θ-BWM can cause numerical problems. Without loss of generality, we focus on the standard BWM because 2θ-BWM is isometric to the BWM. Although general parallel transport under BWM is obtained by solving an ODE, 279 A.4. Additional Discussions transport starting from the identity matrix has a closed-form expression [196]: Γ I→P (V ) = U " r σ i + σ j 2 U ⊤ V U ij # U ⊤ ,(A.56) where P = U ΣU ⊤ is the eigendecomposition of P ∈S n ++ . The forward computation of Eq. (A.56) is numerically stable. However, its backpropagation requires differentiating through an eigendecomposition, involving the calculation of 1 /(σ i −σ j ) [110, Prop. 2]. When σ i is close to σ j , this backpropagation can be problematic. Numerically stable methods for the SPD MLR under power-deformed BWM. To bypass the instability of parallel transport under BWM, we use Lie group left translation to generate ̃ A in MLRs induced by 2θ-BWM. However, there is another problem that could cause instability. The computation of the Riemannian metric of 2θ-BWM requires solving the Lyapunov operator, i.e., L P [V ]P + PL P [V ] = V . For symmetric matrices, the Lyapunov operator can be obtained by eigendecomposition: L P [V ] = U V ′ ij σ i + σ j i,j U ⊤ ,(A.57) where V ∈ S n , UV ′ U ⊤ = V , and P = U ΣU ⊤ is the eigendecomposition of P ∈ S n ++ . Similar to Eq. (A.56), the backpropagation of Eq. (A.57) requires 1 /(σ i −σ j ), undermining numerical stability. To remedy this problem, we propose the following formula to stably compute the backpropagation of Eq. (A.57). Proposition 188. For all P ∈ S n ++ and all V ∈ S n , we denote the Lyapunov equation as XP + PX = V,(A.58) where X = L P [V ]. Given the gradient ∂L ∂X of loss L with respect to X, the back- propagation of the Lyapunov operator can be computed by ∂L ∂V =L P [ ∂L ∂X ],(A.59) ∂L ∂P =−XL P [ ∂L ∂X ]−L P [ ∂L ∂X ]X,(A.60) where L P [·] can be computed by Eq. (A.57). 280 Appendix A. Experimental Details and Additional Discussions Proof. Differentiating both sides of Eq. (A.58), we obtain dXP + X dP + dPX + P dX = dV,(A.61) =⇒ dXP + P dX = dV − X dP − dPX,(A.62) =⇒ dX =L P [dV − X dP − dPX].(A.63) Besides, easy computations show that L P [V ] : W = V :L P [W ],∀W,V ∈S n ,(A.64) where · :· denotes the standard Frobenius inner product. Then we have the following: ∂L ∂X : dX = ∂L ∂X :L P [dV − X dP − dPX],(A.65) =⇒ ∂L ∂X : dX =L P [ ∂L ∂X ] : dV + −XL P [ ∂L ∂X ]−L P [ ∂L ∂X ]X : dP.(A.66) Remark 189. Eq. (A.57) needs to be computed in the Lyapunov operator’s forward and backward processes. Therefore, in the forward process, we can save the inter- mediate matrices U and K with K i,j = h 1 σ i +σ j i i,j , and then use them to compute the backward process efficiently. 281 A.4. Additional Discussions MethodPoint-to-hyperplane distance Applied manifolds Dist Euclidean MLR |⟨a,x⟩ + b| ∥a∥ R n Real Poincaré MLR [76] 1 √ −K sinh −1 2 √ −K|⟨−p⊕ M x,a⟩| 1 + K∥−p⊕ M x∥ 2 ∥a∥ ! [76, Thm. 5] P n K Real Pseudo-Busemann MLR [162] d(x,p) B v (−p⊕ M x) ∥−p⊕ M x∥ [162, Cor. 4.3] P n K Pseudo Lorentz MLR [15] 1 √ −K sinh −1 √ −K ⟨v,x⟩ L ∥v∥ L [15, Eq. (44)] L n K Real BMLR |−αB v (x) + b| α P n K , L n K Real Table A.17: Comparison of point-to-hyperplane distances. Real means the point-to- hyperplane distance is the real distance, obtained by inf y∈H d(x,y) with H as a hyper- plane and d as the geodesic distance. Instead, Pseudo means the point-to-hyperplane distance is a surrogate, which only equals the real distance in Euclidean geometry. A.4.2 Hyperbolic Busemann Neural Networks A.4.2.1 Comparison with Existing Hyperbolic MLRs MethodHyperplaneFormulation Applied manifolds Compact params Euclidean MLREuclidean x∈ R n |⟨a,x⟩ + b = 0 a∈ R n , b∈ R R n ✓ Poincaré MLR [76] Geodesic [76, Def. 3.1] n x∈ P n K | Log p (x),a p = 0 o p∈ P n K , a∈ T p P n K P n K ✗ Pseudo-Busemann MLR [162] Busemann & gyro [162, Def. 4.1] x∈ P n K | B v (−p⊕ M x) = 0 v ∈ S n−1 , p∈ P n K P n K ✗ Lorentz MLR [15] Ambient Minkowski [15, Eq. (7)] x∈ L n K |⟨w,x⟩ L = 0 p∈ L n K , w ∈ T p L n K L n K ✗ BMLRHorosphere x∈H n K |−αB v (x) + b = 0 α > 0, v ∈ S n−1 , b∈ R P n K , L n K ✓ Table A.16: Comparison of hyperplanes. Compact params indicate whether the pa- rameterization requires an additional manifold-valued point. All the considered hyperbolic MLRs follow a point-to-hyperplane formulation. There- fore, the key difference lies in hyperplanes and point-to-hyperplane distances across methods. In addition to Tab. 5.11, Tabs. A.16 and A.17 further make this comparison. 282 Appendix A. Experimental Details and Additional Discussions We draw the following three conclusions. (1) Hyperplanes. Our BMLR uses Busemann-based horospheres that simultaneously satisfy three desiderata: (i) compact parameterization without a per-class manifold- valued point, whereas other hyperbolic ones 32 are over-parameterized; (i) a natural generalization of Euclidean hyperplanes via horospheres, while the Lorentz MLR relies on the ambient Minkowski space, failing to fully respect the intrinsic geom- etry; and (i) applicability across hyperbolic models, whereas the Lorentz MLR is tailored to the Lorentz model. (2) Point-to-Hyperplane Distances. Although pseudo-Busemann MLR also ex- ploits Busemann functions, it relies on a pseudo point-to-hyperplane distance that only coincides with the real distance in Euclidean geometry. In contrast, our BMLR calculates the real point-to-horosphere distance, ensuring geometric fidelity across hyperbolic models. (3) Batch Efficiency. Recalling Tab. 5.11, the Poincaré MLR [76] computes logits using ⟨−p k ⊕ M x,a k ⟩ and ∥−p k ⊕ M x∥ 2 . For a batch X ∈ R bs×n and C classes, evaluating −p k ⊕ M X for every k yields an intermediate tensor of shape [bs,C,n], and materializing this tensor can cause GPU out of memory (OOM) when n or C is large. The same limitation holds for the pseudo-Busemann MLR. Consequently, their official implementations compute per class in a for-loop, which is batch in- efficient. By contrast, BMLR uses logits −α k B v k (x) + b k , whose Busemann term reduces to class-wise inner products ⟨v k ,x⟩ (or ⟨v k ,x s ⟩). With X ∈ R bs×n and V = [v 1 ,...,v C ] ∈ R n×C , such inner products can be efficiently implemented as a single matrix multiplication XV without any [bs,C,n] intermediate, yielding high throughput and low memory usage. A.4.2.2 Busemann Fully Connected Layers and Point-to-Horosphere Dis- tances An apparently natural attempt to define a hyperbolic FC layer is to replace the LHS of Eq. (5.59) by the signed point-to-horosphere distance. The Euclidean hyperplane passing through the origin and orthogonal to e k ∈ R m is H e k ,0 =y ∈ R m |⟨e k ,y⟩ = 0. The corresponding hyperbolic horosphere, following Eq. (5.55), is H e k ,1,0 = y ∈ H m K | −B e k (y) = 0. By Eq. (5.57), the signed point-to-horosphere distance to H e k ,1,0 equals 32 We note that Shimizu et al. [181], Bdeir et al. [15] mitigate this issue through re-parameterization: a = PT e→p (z) and p = Exp e b z ∥z∥ , where e∈H n K denotes the origin, z ∈ R n , and b∈ R. Neverthe- less, the underlying definitions are over-parameterized. 283 A.4. Additional Discussions −B e k (y). Accordingly, we can define an alternative FC mappingF :H n K ∋ x7→ y ∈H m K via B e k (y) = α k B v k (x)− b k , k = 1,...,m,(A.67) where u k (x) =−α k B v k (x) + b k , and α k > 0, v k ∈ S n−1 , b k ∈ R are learnable. Although this only differs from Eq. (5.59) on the LHS, the following discussion shows that such a definition is infeasible in general and fails to deliver a valid hyperbolic FC layer. Poincaré Model. Using Eq. (5.40) with v = e k , Eq. (A.67) becomes 1 √ −K log e k − √ −Ky 2 1 + K∥y∥ 2 ! =−u k (x).(A.68) Define t k = exp − √ −Ku k (x) > 0. Exponentiating Eq. (A.68) gives t k 1 + K∥y∥ 2 = 1− 2 √ −Ky k − K∥y∥ 2 , k = 1,...,m.(A.69) Writing R =∥y∥ 2 , Eq. (A.69) yields an affine expression for each coordinate y k = c k + d k R, c k = 1− t k 2 √ −K , d k = √ −K 2 (1 + t k ).(A.70) Imposing R = P m k=1 y 2 k = P m k=1 (c k + d k R) 2 gives a quadratic in R: A 2 R 2 + (A 1 − 1)R + A 0 = 0,(A.71) where, denoting T = P m k=1 t k and q = P m k=1 t 2 k , A 2 = −K 4 (m + 2T + q), A 1 = m− q 2 , A 0 = m− 2T + q 4(−K) ≥ 0.(A.72) 284 Appendix A. Experimental Details and Additional Discussions Its discriminant is ∆ P = (A 1 − 1) 2 − 4A 2 A 0 = m− q 2 − 1 2 − 4 −K 4 (m + 2T + q) m− 2T + q 4(−K) = m− 2− q 2 2 − (m + 2T + q) (m− 2T + q) 4 = 1 4 (m− 2− q) 2 − (m + q) 2 − (2T ) 2 = 1 4 ((m− 2− q)− (m + q)) ((m− 2− q) + (m + q)) + 4T 2 = 1 4 (−2− 2q)(2m− 2) + 4T 2 = T 2 − (m− 1) (1 + q). (A.73) A real solution R exists only if ∆ P ≥ 0. In addition, feasibility requires 0≤ R <−1/K so that y ∈ P m K . Such conditions can fail for generic u k (x), in which case no y ∈ P m K satisfies Eq. (A.67). Lorentz Model. Using Eq. (5.41) with v = e k , Eq. (A.67) becomes 1 √ −K log √ −K (y t − (y s ) k ) =−u k (x).(A.74) Define t k = exp − √ −Ku k (x) > 0 and t = (t 1 ,...,t m ) ⊤ . Denoting T = P m k=1 t k and q = P m k=1 t 2 k , we obtain from Eq. (A.74), (y s ) k = y t − t k √ −K , k = 1,...,m,=⇒ y s = y t 1− 1 √ −K t.(A.75) Enforcing the hyperboloid constraint ∥y s ∥ 2 − y 2 t = 1/K yields a quadratic in y t : (m− 1)Ky 2 t + 2 √ −KTy t − (1 + q) = 0.(A.76) The discriminant is ∆ L = 2 √ −KT 2 − 4(m− 1)K (− (1 + q)) = 4(−K)T 2 + 4(m− 1)K (1 + q) = 4(−K) T 2 − (m− 1) (1 + q) . (A.77) Hence, a real y t exists only if T 2 − (m− 1) (1 + q) ≥ 0. In addition, feasibility re- 285 A.4. Additional Discussions quires y t > 0. Such conditions can fail for generic u k (x), where no y ∈ L m K satisfies Eq. (A.67). Summary. Equating Busemann coordinates as in Eq. (A.67) requires nontrivial inequalities on the responses u k (x). These constraints are not guaranteed during learning, so the system can become infeasible and the output y undefined. This moti- vates our choice in Eq. (5.59) to use signed point-to-hyperplane distances on the LHS, which admit closed-form solutions that are feasible for all inputs and parameters in both the Poincaré and Lorentz models. A.4.3 Full-Rank Correlation Networks A.4.3.1 Connections among FC Layers: Correlation, SPD, Poincaré, and Euclidean We clarify the correspondence between our FC formulation in Eq. (5.69) and previous FC layers. SPD Manifold. Nguyen et al. [161, Props. 3.4–3.6] introduced three SPD FC layers based on the gyrovector spaces under LEM, LCM, and AIM, respectively. These gyro SPD FC layers share the same definition as Eq. (5.69), except that their signed distance and v k are defined by gyrovector spaces. Poincaré Ball. We show that the Poincaré FC layer F (·) : P n K → P m K reviewed in Sec. A.2.4 is also defined as our correlation FC layer in Thm. 141. We define the zero vector 0∈ P n K as the Poincaré origin, as it is the identity element of the Poincaré gyrovector space [76]. Obviously, e k m k=1 is the orthogonal basis over T 0 P m K , where e k = (δ ik ) m i=1 . Corresponding to Eq. (5.69), we have sign(⟨Log 0 (y),e k ⟩) d(y,H e k ,0 ) = v k (x).(A.78) Compared with Shimizu et al. [181, Eq. (56)], we only need to show the LHS. The sign can be calculated as sign(⟨Log 0 (y),e k ⟩ 0 ) (1) = sign 4 tanh −1 √ −K∥y∥ y √ −K∥y∥ ,e k (2) = sign (⟨y,e k ⟩) = sign (y k ). (A.79) The above follows from the following. 286 Appendix A. Experimental Details and Additional Discussions (1) λ K 0 = 2 1+K∥0∥ 2 = 2 and Log 0 (y) = tanh −1 √ −K∥y∥ y √ −K∥y∥ . (2) tanh −1 (a) > 0, ⇐⇒ a > 0. Therefore, the LHS of Eq. (A.78) is simplified as sign(⟨Log 0 (y),e k ⟩) d(y,H e k ,0 ) = sign (y k ) d(y,H e k ,0 ) (1) = sign (y k ) 1 √ −K sinh −1 2 √ −K|y k | 1 + K∥y∥ 2 (2) = 1 √ −K sinh −1 2 √ −Ky k 1 + K∥y∥ 2 . (A.80) The above comes from the following. (1) Ganea et al. [76, Thm. 5]. (2) 1 + K∥y∥ 2 > 0 by definition and sign(a) sinh −1 (|a|) = sinh −1 (a),∀a∈ R. The last equation in Eq. (A.80) is the LHS of Shimizu et al. [181, Eq. (56)], indicating the equality. Euclidean Space. We show that the Euclidean FC layer F (·) : R n ∋ x → y = Ax + b∈ R m can also be defined as our correlation FC layer in Thm. 141. In Euclidean space R n , the zero vector 0 ∈ R n is the origin, and e k m k=1 is the orthogonal basis over T 0 R m ∼ = R m . Then, the RHS of Eq. (5.69) becomes v k (x) =⟨x− p k ,a k ⟩ (1) = ⟨x,z k ⟩− γ k ∥z k ∥. (A.81) where (1) comes from Exp 0 (γ k [z k ]) = γ k [z k ] and PT 0→p k (z k ) = z k . The above takes the form of ⟨x,a k ⟩ + b k . On the other hand, the LHS of Eq. (5.69) becomes sign(⟨Log 0 (y),e k ⟩) d(y,H e k ,0 ) (1) = sign(y k ) d(y,H e k ,0 ) (2) = y k . (A.82) The above comes from the following. (1) Log 0 (y) = y and ⟨y,e k ⟩ = y k . (2) d(y,H e k ,0 ) = |⟨y,e k ⟩| ∥e k ∥ =|y k |. 287 A.4. Additional Discussions A.4.3.2 Log-Euclidean Layers under Product Geometry We first review some basic facts of the product geometry, and then discuss the Log- Euclidean correlation MLR and FC layer under the product geometry. Product of Correlation Manifolds. Given a manifold (M,g), the n-fold product is (M n ,g) = Q n i=1 (M,g). Each point and tangent vector over M n are M n ∋ P = (P 1 ∈M,· ,P n ∈M),(A.83) T P M n ∋ V = (V 1 ∈ T P 1 M,· ,V n ∈ T P n M).(A.84) The product metric is ⟨V,W⟩ P = n X i=1 ⟨V i ,W i ⟩ P i , ∀V,W ∈ T P M n .(A.85) Correlation MLR. Following Thm. 138, the logit for the k-th class and the input X =X l ∈ Cor + (n) c l=1 ∈ (Cor + (n)) c is v k (X) = c X l=1 v kl (X l ;Z kl ,γ kl ) = c X l=1 (⟨φ n (X l ), (φ n ) ∗,I n (Z kl )⟩− γ kl ∥(φ n ) ∗,I n (Z kl )∥), (A.86) where d n = n(n− 1)/2, φ n : Cor + (n) → R d n is the diffeomorphism of the selected Log-Euclidean geometry, I = (I n ,· ,I n ), Z k = (Z k1 ,· ,Z kc ) ∈ T I (Cor + (n)) c , Z kl ∈ T I n Cor + (n), and γ kl ∈ R. Each v kl is the response of the l-th component given by Thm. 138, with the unit direction [Z kl ] = Z kl /∥Z kl ∥ I n . Correlation FC Layer. Following Thm. 201 and Eq. (A.86), the FC layer F (·) : (Cor + (n)) c → Cor + (m) for the input X is Y = φ −1 m d m X i=1 c X l=1 v il (X l ;Z il ,γ il )e i ! ,(A.87) where d m = m(m− 1)/2, φ m : Cor + (m) → R d m is the output diffeomorphism, Z il ∈ T I n Cor + (n) ∼ = Hol(n), and γ il ∈ R. Eq. (A.87) implies thatF (·) : (Cor + (n)) c → Cor + (m) differs fromF (·) : Cor + (n)→ Cor + (m) only in each scalar response v i , where the former is a summation over the c input components. For example, considering the FC layerF (·) : (Cor + (n)) c → Cor + (m) 288 Appendix A. Experimental Details and Additional Discussions under ECM, its matrix-coordinate response v ij for the input C =C l ∈ Cor + (n) c l=1 is v ij (C) = c X l=1 v EC ijl (C l ;Z ijl ,γ ijl ),(A.88) where Z ijl ∈ Hol(n) and γ ijl ∈ R, for i,j = 1,· ,m with i > j, and 1≤ l ≤ c. A.4.4 Adaptive Log-Euclidean Metrics A.4.4.1 Well-definedness of General Matrix Logarithm Eq. (6.12) does not explicitly specify the correspondence between eigenvalues and diag- onal logarithms. Here, we present a detailed explanation. In implementations such as PyTorch or MATLAB, eigendecomposition routines return eigenpairs in a prescribed order. We rewrite the eigendecomposition as S = P σ i E i , where E i = u i u ⊤ i and u i is the corresponding eigenvector in U. Let S be an n× n SPD matrix, and let P n be the permutation group on 1,...,n. Changing the order of 1,...,n can be viewed as a permutation, so we use π ∈ P n to represent the corresponding changed order. Assume that the eigenvalues σ i are sorted in ascending order, i.e., σ 1 ≤ · ≤ σ n . We use “the i-th eigenvalue” to refer to the eigenvalue in the i-th pair of the ordered eigenpair sequence (σ 1 ,u 1 ),..., (σ n ,u n ). Since each eigenvector u i is unique, it is safe to say the eigenvalues are ordered, and the i-th eigenvalue/eigenvector pair is unique. Let log α (Σ) denote applying the scalar logarithm log a i to the i-th eigenvalue σ i . Then log α (S) is rewritten as log α (S) = P log a i (σ i )E i , where S = P σ i E i . In this way, log α is clearly well-defined. By definition, we can observe that the output of log α does not depend on the order in eigendecomposition. Suppose there are two eigendecompositions with different orders, i.e., S = U ΣU ⊤ = ̃ U ̃ Σ ̃ U ⊤ , where ̃ U and ̃ Σ are rearrangements of U and Σ. There exists a π ∈ P n such that, for each j, there is a unique i satisfying ̃u j = u π(i) and ̃σ j = σ (i) . We then have P log a i (σ i )E i for S = U ΣU ⊤ and P log a π(i) (σ π(i) )E π(i) for S = ̃ U ̃ Σ ̃ U ⊤ , which indicates that the two eigendecompositions are equivalent. 289 A.4. Additional Discussions A.4.5 Product Cholesky Metrics A.4.5.1 Riemannian Operators under (θ, M)-DBWM We first define a map φ θ :L n ++ →L n ++ as φ θ (L) =⌊L⌋ + L θ ,∀L∈L n ++ .(A.89) Its differential at L∈L n ++ is given by (φ θ ) ∗,L (X) =⌊X⌋ + θL θ−1 X, ∀X ∈ T L L n ++ .(A.90) Let g M-DBW and g (θ,M)-DBW be M-DBWM and (θ, M)-DBWM, respectively. Since constant scaling of a Riemannian metric preserves the Christoffel symbols, Riemannian operators such as the Riemannian logarithm, exponential map, and parallel transport under g (θ,M)-DBW are the same as those under the pullback metric φ ∗ θ g M-DBW . Follow- ing Thm. 33, these Riemannian operators under (θ, M)-DBWM can be obtained by φ ∗ θ g M-DBW . Besides, as constant scaling does not affect the weighted Fréchet mean (WFM), the WFM under (θ, M)-DBWM is the same as that under φ ∗ θ g M-DBW . The lat- ter can be calculated by the properties of isometries presented in Thm. 33. Therefore, by the properties of isometry in Thm. 33 and Eqs. (A.89) and (A.90), we can obtain all the Riemannian operators. 290 Appendix B Proofs B.1 Mathematical Background B.1.1 Proof of Thm. 57 Proof. The gyration-preservation property is shown by Ungar [200, Eq. (6.323)]. The proofs for the gyro identity and gyroinverse follow directly from the isomorphism and the uniqueness of inverse and identity [200, Thm. 2.10]. B.2 Lie Group Batch Normalization B.2.1 Proof of Thm. 61 Proof. Case (1). The MLE of M is M MLE = argmax log(k(v))− N X i=1 d(P i ,M ) 2 2v 2 = argmin N X i=1 d(P i ,M ) 2 . (B.1) Case (2). We denote Y = L B (X), and p X and p Y as the density of X and Y , 291 B.2. Lie Group Batch Normalization respectively. The density of Y is p Y (Q) (1) = p X (L ⊖B (Q)) = k(σ) exp − d(L ⊖B (Q),M ) 2 2σ 2 (2) = k(σ) exp − d(Q, L B (M )) 2 2σ 2 . (B.2) The above comes from: (1) Pennec [170, Thm. 7]; (2) The isometry of the left translation. B.2.2 Proof of Thm. 62 Proof. The isometry of L B directly implies the homogeneity of the sample mean. Now let us focus on Eq. (3.14). We have the following: X N i=1 w i d 2 (φ s (P i ),E) = X N i=1 w i ∥s Log E P i ∥ 2 E = s 2 X N i=1 w i ∥ Log E P i ∥ 2 E = s 2 X N i=1 w i d 2 (P i ,E), (B.3) where ∥·∥ E is the norm on T E M. B.2.3 Proof of Thm. 65 Proof. As the right-invariant metric has properties analogous to those of the left- invariant metric, this proof follows the same logic as the two proofs above. Gaussian homogeneity. We denote Y = R B (X), and p X and p Y as the density of X and Y , respectively. The density of Y is p Y (Q) (1) = p X (R ⊖B (Q)) = k(σ) exp − d(R ⊖B (Q),M ) 2 2σ 2 (2) = k(σ) exp − d(Q, R B (M )) 2 2σ 2 . (B.4) 292 Appendix B. Proofs The above comes from: (1) Pennec [170, Thm. 13]; (2) The isometry of the right translation. Sample mean homogeneity. This is a direct corollary of the isometry of right translation. B.2.4 Proof of Thm. 66 Proof. As R n is an abelian group and the Euclidean inner product is bi-invariant, we focus on left translation in the following. The core of this proof lies in the fact that on R n , (1) the Fréchet mean and variance are reduced to the familiar Euclidean statistics. (2) the calculation of the running mean becomes the weighted arithmetic mean. (3) Eqs. (3.8) to (3.10) become Eq. (3.1). We prove these three points one by one. As stated by Lou et al. [142, Prop. G.1 and Cor. G.2], from the view of the product manifold, the element-wise Fréchet mean and variance on R n are equivalent to the vector-valued Euclidean mean and variance. Besides, by an argument similar to that of Lou et al. [142, Prop. G.1], the weighted Fréchet mean on R n is simplified to the weighted arithmetic average. Therefore, on R n , the calculation of running statistics in our Alg. 1 becomes the familiar moving average. Thirdly, on R n , we know that L x (y) = x + y, Exp x v = x + v, Log x y = y− x, and the neutral element is 0. Since statistics, as well as the Euclidean BN, are calculated element-wise, it suffices to consider a single coordinate, i.e., R. For a batch of activations x i N i=1 ⊂ R with batch mean μ b and batch variance v 2 b , Eqs. (3.8) to (3.10) are rewritten as follows: L β Exp 0 " γ p v 2 b + ε Log 0 (L −μ b (x i )) #! = γ x i − μ b p v 2 b + ε + β.(B.5) The above equation is the exact core computation of the standard Euclidean BN. B.2.5 Proof of Thm. 67 Proof. We first prove the case of (θ,α,β)-LEM, and then proceed to the case of θ-LCM. (θ,α,β)-LEM. For clarity, we denote the metric tensor of (θ,α,β)-LEM as g (θ,α,β)-LE = 1 θ 2 P ∗ θ g (α,β)-LE ,(B.6) where g (α,β)-LE is the metric tensor of (α,β)-LEM. Let P ∈ S n ++ and V,W ∈ T P S n ++ . 293 B.2. Lie Group Batch Normalization Then we have g (θ,α,β)-LE P (V,W ) = 1 θ 2 g (α,β)-LE P θ (P ) ((P θ ) ∗,P (V ), (P θ ) ∗,P (W )) = 1 θ 2 ⟨(log◦ P θ ) ∗,P (V ), (log◦ P θ ) ∗,P (W )⟩ (α,β) =⟨log ∗,P (V ), log ∗,P (W )⟩ (α,β) = g (α,β)-LE P (V,W ). (B.7) θ-LCM. Let us first review a well-known fact of deformed metrics [194]. Let ̃g = 1 θ 2 P ∗ θ g be the power-deformed metric on SPD. Then when θ tends to 0, for all P ∈S n ++ and all V ∈ T P S n ++ , we have ̃g P (V,V )→ g I (log ∗,P (V ), log ∗,P (V )).(B.8) By Eq. (B.8), we can readily obtain the results. B.2.6 Proof of Thm. 68 Proof. (α,β)-AIM is left-invariant [195]. As the pullback of (α,β)-AIM, (θ,α,β)-AIM is left-invariant as well. Besides, Thm. 147 shows that LCM is the pullback metric from the Euclidean space of LT n . Therefore, θ-LCM is bi-invariant. B.2.7 Proof of Thm. 69 Proof. In the following, we denote the Riemannian operators under the left-invariant metric, i.e., AIM, by d L , Log L , and Exp L . Note that the Cholesky decomposition pulls back the group operation of matrix product from the Cholesky manifoldL n ++ [195]. For simplicity, we abbreviate ⊕ LieAI as ⊕. The differential maps of the Cholesky decomposition and its inverse are reviewed in Eq. (2.92). Following the notation in this theorem, we further denote X ∈ T L L n ++ . Specifically, for the differential map at I, we have Chol ∗,I (V ) = (V ) 1 2 ,∀V ∈ T I S n ++ .(B.9) (Chol −1 ) ∗,I (X) = (X) sym+ ,∀X ∈ T I L n ++ .(B.10) Denoting e L and e R as the group translation on the Cholesky manifoldL n ++ , we have the 294 Appendix B. Proofs following w.r.t. the differential maps of left and right translation: ((L P ) ∗,⊖P ) −1 (1) = Chol −1 ∗,I ( e L L ) ∗,L −1 ◦ Chol ∗,⊖P −1 (B.11) = (Chol ∗,⊖P ) −1 ( e L L ) ∗,L −1 −1 ◦ Chol ∗,I ,(B.12) (R ⊖P ) ∗,P (2) = Chol −1 ∗,I ◦ ( e R L −1 ) ∗,L ◦ Chol ∗,P .(B.13) The above derivation comes from the following: (1) L P = Chol −1 ◦ e L L ◦ Chol; (2) R ⊖P = Chol −1 ◦ e R L −1 ◦ Chol. Riemannian metric. For the differential of right translation, we have the following: (R ⊖P ) ∗,P (V ) = Chol −1 ∗,I ◦ ( e R L −1 ) ∗,L ◦ Chol ∗,P (V ) = L(L −1 V L −⊤ ) 1 2 L −1 sym+ . (B.14) By Eq. (B.14), one can obtain the expression for the Riemannian metric tensor. Riemannian geodesic and exponential map. According to Zacur et al. [228], we have the following for the operators between left- and right-invariant metrics: Exp P (V ) =⊖ Exp L ⊖P − ((L P ) ∗,⊖P ) −1 ◦ (R ⊖P ) ∗,P (V ) ,(B.15) d(P,Q) = d L (⊖P,⊖Q).(B.16) Putting the AIM-based geodesic distance into the RHS of Eq. (B.16), one can obtain the geodesic distance under CRIM. Now, we simplify Eq. (B.15). Putting Eqs. (B.12) and (B.13) into Eq. (B.15), we have the following: Exp P (V ) =⊖ Exp L ⊖P − (Chol ∗,⊖P ) −1 ◦ ( e L L ) ∗,L −1 −1 ◦ ( e R L −1 ) ∗,L ◦ Chol ∗,P (V ) =⊖ Exp L ⊖P − (Chol ∗,⊖P ) −1 L −1 Chol ∗,P (V )L −1 (1) = ⊖ n Exp L ⊖P − (Chol ∗,⊖P ) −1 h L −1 V L −⊤ 1 2 L −1 io (2) = ⊖ Exp L ⊖P − L −1 V L −⊤ 1 2 L −1 L −⊤ sym+ . (B.17) The above comes from the following: 295 B.2. Lie Group Batch Normalization (1) The first identity follows from Eq. (2.92); (2) The second identity follows from Eq. (2.92). Riemannian logarithm. From the second equality in Eq. (B.17), we have the following: Log P (Q) =− Chol −1 ∗,P L Chol ∗,⊖P Log L ⊖P (⊖Q) L (1) = − Chol −1 ∗,P L e V L ⊤ 1 2 L =− L ⊤ L e V L ⊤ ⊤ 1 2 sym+ (B.18) The above comes from the following: (1) Chol ∗,⊖P (V ) = L −1 LV L ⊤ 1 2 ,∀V ∈ T ⊖P S n ++ . B.2.8 Proof of Thm. 70 Proof. Completeness. Eq. (3.21) indicates that Exp I is defined over the whole T I S n ++ . Besides, the SPD manifold is connected [171]. By Lee [131, Cor. 6.20], CRIM is com- plete. Geodesic. For simplicity, we abbreviate ⊕ LieAI as ⊕. The geodesic connecting P and Q can be obtained as follows: γ(t;P,Q) = Exp P (t Log P (Q)) (1) = ⊖ Exp L ⊖P − (Chol ∗,⊖P ) −1 L −1 Chol ∗,P (t Log P (Q))L −1 (2) = ⊖ Exp L ⊖P t Log L ⊖P (⊖Q) =⊖ n γ AI (t; e P, e Q) o . (B.19) The above comes from the following: (1) The second equality in Eq. (B.17); (2) The first equality in Eq. (B.18). 296 Appendix B. Proofs B.2.9 Proof of Thm. 72 Proof. Without loss of generality, we focus on the case of the left-invariant metric. The results for the right-invariant metric can be proven similarly. We denote Eqs. (3.8) to (3.10) on M i ,i = 1, 2 as the mapping ξ i (·|M,v 2 ,B,s). Let B =P i N i=1 and f (B) =f (P i ) N i=1 . The core of this proof lies in three points: (1) The Fréchet mean and variance ofB inM 1 correspond to the counterparts of f (B) in M 2 . (2) ξ 1 (P i |M,v 2 ,B,s) in M 1 is equal to f −1 (ξ 2 (f (P i )|f (M ),v 2 ,f (B),s)). (3) The updates of running statistics in M 1 correspond to the counterparts in M 2 . We denote M as the Fréchet mean of B, and v 2 as the Fréchet variance of B. Then, by the isometry of f, the Fréchet mean and variance of f (B) are f (M ) and v 2 , respectively. On M i ,i = 1, 2, we denote L i ,⊕ i , Exp i , Log i as the Lie group and Riemannian operators, E i as the neutral element, and Eq. (3.9) as φ i s (·). With the isometry and Lie group isomorphism of f, we have the following equations: L 1 ⊖ 1 M = f −1 ◦ L 2 ⊖ 2 f (M ) ◦f,(B.20) φ 1 s = Exp 1 E 1 s Log 1 E 1 (·) (B.21) = f −1 Exp 2 E 2 s Log 2 E 2 (f (·)) (B.22) = f −1 ◦ φ 2 s ◦ f,(B.23) L 1 B = f −1 ◦ L 2 f (B) ◦f.(B.24) Then we have ξ 1 (P i |M,v 2 ,B,s) = f −1 (ξ 2 (f (P i )|f (M ),v 2 ,f (B),s))(B.25) Lastly, we show the correspondence between running statistics. Since the Fréchet variance is the same for both B and f (B), we focus on the running mean. Let M r and f (M r ) denote the initial values of the running means in M 1 and M 2 , respectively, and let WFM i represent the weighted Fréchet mean in M i . Then the updated running mean in M 1 is WFM 1 (1− η,η,M r ,M) = f −1 (WFM 2 (1− η,η,f (M r ),f (M ))) (B.26) 297 B.3. Gyrogroup Batch Normalization We can further simplify the above equation as WFM 1 = f −1 ◦ WFM 2 ◦f(B.27) Denoting LieBN i as the LieBN algorithm on M i , Eq. (B.25) and Eq. (B.27) imply that LieBN 1 (P i |B,s,ε,η) = f −1 LieBN 2 (f (P i )|f (B),s,ε,η) .(B.28) B.3 Gyrogroup Batch Normalization B.3.1 Proof of Thm. 74 Proof. By Nguyen and Yang [159, Lem. 2.3], easy computations show that Eq. (3.33) holds for Gr(p,n) if and only if it holds for f Gr(p,n). Without loss of generality, we prove the case for the projector perspective. Given any P,Q∈ f Gr(p,n), Nguyen [157, Def. 3.18] gives the expression for gyration: gyr[⊖P,P ]Q = F (⊖P,P )Q (F (⊖P,P )) −1 ,(B.29) with F (⊖P,P ) defined as F (⊖P,P ) = exp − h ⊖P ⊕ P, e I p,n i exp h ⊖P, e I p,n i exp h P, e I p,n i ,(B.30) where(·) = Log e I p,n (·). This equation can be further simplified as F (⊖P,P ) (1) = exp(0) exp h ⊖P, e I p,n i exp h P, e I p,n i (2) = exp h ⊖P, e I p,n i exp h P, e I p,n i (3) = I n . (B.31) The above derivation follows from (1)⊖P ⊕ P = e I p,n = 0 n×n ∈ R n×n . (2) exp(0) = I n . (3)⊖P =−P and exp h −P, e I p,n i = exp h P, e I p,n i −1 . 298 Appendix B. Proofs Therefore, gyr[⊖P,P ] is the identity map. B.3.2 Proof of Thm. 75 Proof. This theorem follows from Ungar [200, Thms. 2.10–2.11], which presents some useful properties for gyrogroups. We argue that all the properties except gyr[a,a] = id are independent of the left reduction law (G4), and are therefore satisfied on pseudo- reductive gyrogroups. All the properties can be proven in the same way as in Ungar [200, Thms. 2.10–2.11]. We summarize the logic in the following: • left gyroassociativity ⇒ Case (1) • left gyroassociativity + Case (1) ⇒ Case (2) • definition ⇒ Case (3) • left gyroassociativity + Case (1) + Case (3) ⇒ Case (4) • definition ⇒ Case (5) • left gyroassociativity + (G2) + Case (1) + Case (3) + Case (4) + Case (5) ⇒ Case (6) • Case (1)+Case (6) ⇒ Case (7) • left gyroassociativity +Case (3) ⇒ Case (8) • left gyroassociativity + left cancellation in Case (8) ⇒ Case (9) • gyro identity in Case (9) ⇒ Case (10) • Case (10) ⇒ Case (11) • left cancellation in Case (8) + gyro identity in Case (9) ⇒ Case (12) • left cancellation in Case (8) + gyro identity in Case (9) ⇒ Case (13) 299 B.3. Gyrogroup Batch Normalization B.3.3 Proof of Thm. 77 Proof. y (1) = x⊕ (⊖x⊕ y) (2) = Exp x (PT e→x (Log e (⊖x⊕ y))) (3) ⇒ Log x (y) = PT e→x (Log e (⊖x⊕ y)). (B.32) The above comes from the following. (1) Left cancellation law. (2) Definition of gyroaddition. (3) Applying Log x (·) to both sides. By the last equation, we have d(x,y) =∥Log x (y)∥ x =∥PT e→x (Log e (⊖x⊕ y))∥ x (1) = ∥Log e (⊖x⊕ y)∥ e = d gyr (x,y), (B.33) where (1) comes from • Parallel transport preserving the norm [69, Sec. 3.1]. • PT x→e ◦ PT e→x (v) = v,∀v ∈ T e M. B.3.4 Proof of Thm. 78 Proof. We denote the Riemannian logarithm, gyration, and gyronorm on M,g by Log, gyr, and ∥·∥ gyr , respectively, while g Log, fgyr, and ∥·∥ fgyr are their counterparts on f M,eg. We recall the following from Nguyen and Yang [159, Lems. 2.1–2.3]: x⊕ y = φ −1 φ(x) e ⊕φ(y) ,(B.34) t⊙ x = φ −1 t e ⊙φ(x) ,(B.35) gyr[x,y]z = φ −1 (fgyr[φ(x),φ(y)]φ(z)),(B.36) where x,y,z ∈M. 300 Appendix B. Proofs Gyrodistance. g d gyr (φ(x),φ(y)) = e ⊖φ(x) e ⊕φ(y) fgyr (1) = ∥φ(⊖x⊕ y)∥ fgyr = r D ] Log e (φ(⊖x⊕ y)), ] Log e (φ(⊖x⊕ y)) E e (2) = q ⟨Log e (⊖x⊕ y), Log e (⊖x⊕ y)⟩ e =∥⊖x⊕ y∥ gyr = d gyr (x,y). (B.37) The derivation above comes from the following. (1) By Eqs. (B.34) and (B.35): φ(⊖x⊕ y) = φ(⊖x) e ⊕φ(y).(B.38) (2) By the isometry: Log x (y) = (φ ∗,x ) −1 g Log φ(x) (φ(y)) ,∀x,y ∈M,(B.39) ⟨v,w⟩ x =⟨φ ∗,x (v),φ ∗,x (w)⟩ φ(x) ,∀x∈M and ∀v,w ∈ T x M,(B.40) where φ ∗,x is the differential map. Here, the RHSs contain the operators over f M, whereas the LHSs involve those over M. Gyroisometry. Given any x,y,z,a∈M, we have the following by the isometry of φ. For the gyroinverse: d fgyr ( e ⊖φ(x), e ⊖φ(y)) = d fgyr (φ(⊖x),φ(⊖y)) = d gyr (⊖x,⊖y) = d gyr (x,y) = d fgyr (φ(x),φ(y)). (B.41) 301 B.3. Gyrogroup Batch Normalization For the gyration: d fgyr (fgyr[φ(z),φ(a)]φ(x), fgyr[φ(z),φ(a)]φ(y)) = d fgyr (φ(gyr[z,a]x),φ(gyr[z,a]y)) = d gyr (gyr[z,a]x, gyr[z,a]y) = d gyr (x,y) = d fgyr (φ(x),φ(y)). (B.42) For the left gyrotranslation: d fgyr (φ(z) e ⊕φ(x),φ(z) e ⊕φ(y)) = d fgyr (φ(z⊕ x),φ(z⊕ y)) = d gyr (z⊕ x,z⊕ y) = d gyr (x,y) = d fgyr (φ(x),φ(y)). (B.43) B.3.5 Proof of Thm. 79 We first prove a useful lemma. Lemma 190 (Left gyrotranslation law). Every pseudo-reductive gyrogroup G,⊕ verifies the left gyrotranslation law: ⊖ (x⊕ y)⊕ (x⊕ z) = gyr[x,y] (⊖y⊕ z), ∀x,y,z ∈ G.(B.44) Proof. This lemma generalizes Nguyen and Yang [159, Lems. I.1 and L.1], which prove the left gyrotranslation law on the specific gyrogroups of the SPD and Grassmannian manifolds. Their proof only relies on the left cancellation and the basic axioms (G1– G3). Note that the original proof of left gyrotranslation on the Grassmannian [159, Lem. I.1] is questionable, as it relies on the left cancellation of gyrogroups, and the Grassmannian is not a gyrogroup but a non-reductive gyrogroup. Fortunately, as we show in Thm. 75, the general pseudo-reductive gyrogroups, including the Grassmannian, enjoy left cancellation. Therefore, the proof in Nguyen and Yang [159, Lem. I.1] can be readily generalized to pseudo-reductive gyrogroups. Proof of Thm. 79. ⇒: For any z,a ∈ G, the gyroautomorphism can be expressed by 302 Appendix B. Proofs the gyrator identity in Thm. 75: gyr[x,y]z = X ⊕ ̄z,(B.45) gyr[x,y]a = X ⊕ ̄a,(B.46) where X =⊖(x⊕y), ̄z = x⊕ (y⊕z), and ̄a = x⊕ (y⊕a). Then Nguyen and Yang [159, Eq. (30)], derived for the specific Grassmannian, can be directly extended to pseudo- reductive gyrogroups, as it only relies on left gyrotranslation, invariance of the norm under gyroautomorphisms, and the axioms (G1–G3). ⇐: ∥gyr[x,y](z)∥ gyr =∥gyr[x,y](⊖e⊕ z)∥ gyr (Case (7) in Thm. 75 indicates ⊖e = e) =∥⊖ gyr[x,y](e)⊕ gyr[x,y](z)∥ gyr (automorphism) = d (gyr[x,y](e), gyr[x,y](z)) = d (e,z) =∥⊖e⊕ z∥ gyr =∥z∥ gyr . (B.47) B.3.6 Proof of Thm. 80 Proof. Given any x,y,z ∈ G, we prove the two claims as follows. Gyroisometry of the left gyrotranslation. This property generalizes Nguyen and Yang [159, Thms. 2.12 and 2.16], which deal with the gyrotranslation in the SPD and Grassmannian, respectively. We have the following: d(L x (y),L x (z)) = d(x⊕ y,x⊕ z) =∥⊖(x⊕ y)⊕ (x⊕ z)∥ gyr =∥gyr[x,y] (⊖y⊕ z)∥ gyr (left gyrotranslation law) =∥⊖y⊕ z∥ gyr (gyroisometry of the automorphism) = d(y,z). (B.48) 303 B.3. Gyrogroup Batch Normalization Gyroisometry of the gyroinverse. d(⊖x,⊖y) =∥x⊖ y∥ gyr =∥⊖y⊕ x∥ gyr (gyrocommutativity and gyroisometry of the automorphism) = d(y,x) = d(x,y) (symmetry of the geodesic distance). (B.49) B.3.7 Proof of Thm. 81 We give a concise argument based on Thms. 77, 79 and 80. Proof. First, all these gyrospaces are characterized by Eqs. (2.77) and (2.78). By Thm. 77, their gyrodistances agree with the geodesic distances. It thus remains to establish the gyroisometries. As shown by Thms. 79 and 80, it suffices to show that gyrations in each space preserve the gyronorm. This argument on the SPD and ONB Grassmannian has already been proven [159, Lems. L.2 and I.2]. Since the ONB Grassmannian is isometric to the P Grassmannian via Eq. (2.138), Thm. 78 implies that the same arguments apply to the P. We therefore only need to treat st n K with K ≤ 0. As the Euclidean case R n is trivial, we only need to show the Poincaré ball. In the following, a,b,x,y are arbitrary points in P n K . Norm invariance under gyrations. As the Poincaré ball forms a real inner- product gyrovector space [200, Def. 6.2 and Thm. 6.85], any gyration preserves the Euclidean norm: ∥gyr[a,b](x)∥ =∥x∥, ∀x∈ st n K .(B.50) For the gyronorm, we further have ∥gyr[a,b]x∥ gyr =∥ Log 0 (gyr[a,b]x)∥ 0 = 2 √ |K| tanh −1 p |K|∥ gyr[a,b]x∥ = 2 √ |K| tanh −1 p |K|∥x∥ =∥x∥ gyr . (B.51) 304 Appendix B. Proofs B.3.8 Proof of Thm. 83 Proof. According to Thm. 80, any left gyrotranslation is a gyroisometry. For any y ∈ M, we have the following: d (β⊕ x i ,y) (1) = d (⊖β⊕ (β⊕ x i ),⊖β⊕ y) (2) = d ((⊖β⊕ β)⊕ gyr[⊖β,β](x i ),⊖β⊕ y) (3) = d (x i ,⊖β⊕ y). (B.52) The above comes from the following. (1) Any left gyrotranslation is a gyroisometry. (2) Left gyroassociative law. (3) ⊖β⊕ β = e and pseudo-reduction. Denoting the gyromean of x i and β⊕ x i as μ and eμ, we have the following: β⊕ μ (1) = β⊕ (⊖β⊕ eμ) (2) = gyr[β,⊖β](eμ) (3) = eμ. (B.53) The above comes from the following. (1) Eq. (B.52) indicates that μ =⊖β⊕ eμ. (2) Left gyroassociative law. (3) Pseudo-reduction. 305 B.3. Gyrogroup Batch Normalization Algorithm 4: ONB Grassmannian logarithm [19, Alg. 5.3] Input: U,Y ∈ Gr(p,n) are Stiefel representatives under the ONB perspective. QSR ⊤ SVD := Y ⊤ U with S in ascending order, and Q and R flipped column-wise accordingly ˆ S = p I p − S 2 ∆ = (I n − U ⊤ )Y Q arcsin( ˆ S) ˆ S R ⊤ Output: Log U (Y ) = ∆ Now, we proceed to deal with the second property. We have the following: d(t⊙ x i ,e) (1) = ∥⊖e⊕ (t⊙ x i )∥ gyr (2) = ∥t⊙ x i ∥ gyr =∥t Log e (x i )∥ e =|t|∥Log e (x i )∥ e =|t|∥x i ∥ gyr (3) = |t|∥⊖e⊕ x i ∥ gyr =|t| d(e,x i ) (4) = |t| d(x i ,e) (B.54) The above follows from the following. (1) Symmetry of gyrodistance (as geodesic distance). (2) ⊖e = e. (3) x i =⊖e⊕ x i . (4) Symmetry of gyrodistance (as geodesic distance). The last equation in Eq. (B.54) indicates the homogeneity of dispersion from e. B.3.9 Proof of Thm. 86 We first review a fast and stable algorithm for the ONB Grassmannian logarithm [19, Alg. 5.3], and the calculation of Grassmannian logarithm under the projector perspec- tive by the ONB Grassmannian logarithm [161, Prop. 3.12]. Alg. 4 reviews a fast and stable algorithm for the Grassmannian Riemannian log- arithm under the ONB perspective Gr(p,n). The vanilla Riemannian logarithm in Tabs. 2.10 and 2.11 requires an n×p SVD and a p×p matrix inverse, while Alg. 4 only requires a p× p SVD. Therefore, Alg. 4 is more efficient than the vanilla logarithm. 306 Appendix B. Proofs Besides, Alg. 4 can also return a unique tangent vector when Y is in the cut locus of U. For more details, please refer to Bendokat et al. [19, Sec. 5.2]. As the projector perspective is isometric to the ONB perspective, the Grassmannian logarithm under the projector perspective can be calculated by the ONB Grassmannian logarithm [161, Prop. 3.12]. Proposition 191 ([161]). Given any P,Q ∈ f Gr(p,n) with U = π −1 (P ) and V = π −1 (Q), the Riemannian logarithm g Log P (Q) on f Gr(p,n) is given as g Log P (Q) = π ∗,U (Log U V ),(B.55) where Log is the Riemannian logarithm under the ONB perspective, π ∗,U : T U Gr(p,n)→ T P f Gr(p,n) is the differential map of π at U, which is defined as π ∗,U (∆) = ∆U ⊤ + U ∆ ⊤ ,∀∆∈ T U Gr(p,n).(B.56) Now, we begin to present the proof. Proof of Thm. 86. We first show the expression for Log I p,n and g Log e I p,n . First note the following: (I n − I p,n I ⊤ p,n ) = 0 0 0 I n−p ! ,(B.57) U ⊤ I p,n = U ⊤ 1 ,U ⊤ 2 I p 0 ! = U ⊤ 1 , (B.58) By the above two equations, the ONB Grassmannian logarithm at I p,n is Log I p,n (U ) = 0 0 0 I n−p ! U 1 U 2 ! Q arcsin( ˆ S) ˆ S R ⊤ (Alg. 4) = 0 U 2 Q arcsin( ˆ S) ˆ S R ⊤ ! = 0 e U 2 ! , (B.59) where QSR ⊤ SVD := U ⊤ 1 with S in ascending order, and Q and R flipped column-wise 307 B.3. Gyrogroup Batch Normalization accordingly, and ˆ S = p I p − S 2 . For g Log e I p,n , we have g Log e I p,n (U ⊤ ) (1) = π ∗,I p,n Log I p,n (U ) (2) = π ∗,I p,n 0 e U 2 !! (3) = 0 e U ⊤ 2 e U 2 0 ! . (B.60) The above derivation comes from the following. (1) Thm. 191 (2) Eq. (B.59) (3) For any ∆ = (∆ ⊤ 1 , ∆ ⊤ 2 ) ⊤ ∈ T I p,n Gr(p,n), where ∆ 1 is p× p, we have the following: π ∗,I p,n ∆ 1 ∆ 2 !! = ∆ 1 ∆ 2 ! (I p ,0) + I p 0 ! ∆ ⊤ 1 , ∆ ⊤ 2 = ∆ 1 0 ∆ 2 0 ! + ∆ ⊤ 1 ∆ ⊤ 2 0 0 ! = ∆ 1 + ∆ ⊤ 1 ∆ ⊤ 2 ∆ 2 0 ! . (B.61) Combining all the above results together, we have the following: [ U ⊤ , e I p,n ] = h g Log e I p,n (U ⊤ ), e I p,n i = " 0 e U ⊤ 2 e U 2 0 ! , e I p,n # = 0 e U ⊤ 2 e U 2 0 ! I p 0 0 0 ! − I p 0 0 0 ! 0 e U ⊤ 2 e U 2 0 ! = 0 − e U ⊤ 2 e U 2 0 ! . (B.62) 308 Appendix B. Proofs B.3.10 Proof of Thm. 88 Proof. Given v ∈ T 0 st n K , we have the following: λ K 0 = 2,(B.63) Log 0 (y) = tan −1 K p |K|∥y∥ y p |K|∥y∥ ,(B.64) PT 0→x (v) = λ K 0 λ K x gyr[x,−0](v) = 2 λ K x v,(B.65) Exp 0 (v) = tan K p |K|∥v∥ v p |K|∥v∥ .(B.66) Therefore, we have Exp x (PT 0→x (Log 0 (y))) = Exp x 2 p |K|λ K x tan −1 K p |K|∥y∥ y ∥y∥ ! = x⊕ K y, (B.67) which implies Exp 0 (t Log 0 (x)) = t⊙ K x. (B.68) B.3.11 Proof of Thm. 89 Proof. We only need to show the case of K > 0. Let x,y,z,w be any vectors in st n K , and s,t∈ R be real scalars. We first notice the gyroaddition is a linear combination: x⊕ K y = (1− 2K⟨x,y⟩− K∥y∥ 2 )x + (1 + K∥x∥ 2 )y 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 = Ax + By D ,(B.69) where A = 1− 2K⟨x,y⟩− K∥y∥ 2 , B = 1 + K∥x∥ 2 , D = 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 . (B.70) Recalling Eq. (2.146), the gyration gyr[x,y](z) is also a linear expression: gyr[x,y](z) = z + f 1 · x + f 2 · y.(B.71) 309 B.3. Gyrogroup Batch Normalization Axioms G1–G3. Since e = 0 and ⊖ K x = −x, G1 and G2 can be immediately verified. The left gyroassociative law follows from the definition of gyration [13, Eq. 36]: gyr[x,y](z) =− (x⊕ K y)⊕ K (x⊕ K (y⊕ K z)).(B.72) To confirm that any gyration is an automorphism of (st n K ,⊕ K ), we verify the following identity using symbolic computation: gyr[x,y](w⊕ K z) = gyr[x,y](w)⊕ K gyr[x,y](z).(B.73) Expanding all necessary inner products (e.g.,⟨x,w⊕ K z⟩,⟨y,w⊕ K z⟩) and norms in terms of ⟨x,w⟩, ⟨x,z⟩, etc., we express both sides of Eq. (B.73) as linear combinations: gyr[x,y](w⊕ K z) = f 1 x + f 2 y + f 3 w + f 4 z,(B.74) gyr[x,y](w)⊕ K gyr[x,y](z) = f ′ 1 x + f ′ 2 y + f ′ 3 w + f ′ 4 z.(B.75) We use SymPy to compare the coefficients of x, y, w, and z on both sides. The above is exposed in stereographic_gyr_automorphism.py. (G4) Left reduction law. Similarly, we expand gyr[x⊕ K z,y](w) and gyr[x,y](w) and compare the coefficients, as implemented in stereographic_left_reduction.py. Gyrocommutative law. This has been verified by Bachmann et al. [13, Lem. 11]. (V1) Identity scalar multiplication. This can be directly verified by definition: t⊙ K x = tan t tan −1 ( √ K∥x∥) √ K x ∥x∥ ,∀x̸= 0.(B.76) (V2) Scalar distributive law. We want to verify (s + t)⊙ K x = (s⊙ K x)⊕ K (t⊙ K x).(B.77) We only need to show the case of x ̸= 0. Let θ = tan −1 ( √ K∥x∥). Then scalar multiplication is simplified as t⊙ K x = tan(tθ) √ K∥x∥ x.(B.78) Using symbolic computation (see stereographic_gyr_v2.py), we obtain the following 310 Appendix B. Proofs w.r.t. Eq. (B.77): LHS = tan ((s + t)θ) √ K∥x∥ ,(B.79) RHS = tan(sθ) + tan(tθ) √ K∥x∥ (1− tan(sθ) tan(tθ)) .(B.80) The identity follows from the tangent addition formula: tan((s + t)θ) = tan(sθ) + tan(tθ) 1− tan(sθ) tan(tθ) .(B.81) (V3) Scalar associative law. This can be directly verified by definition. (V4) Gyroautomorphism. We now verify gyr[x,y](t⊙ K z) = t⊙ K gyr[x,y](z).(B.82) Let us denote α (z) t = tan t tan −1 √ K∥z∥ √ K∥z∥ .(B.83) By the linearity of gyration [13, Lem. 11], the left-hand side of Eq. (B.82) becomes LHS = α (z) t gyr[x,y](z),(B.84) while the right-hand side reads RHS = α gyr[x,y](z) t gyr[x,y](z).(B.85) Since α (·) t depends only on the norm of its argument, it suffices to show ∥z∥ =∥gyr[x,y](z)∥,(B.86) which holds as proven by Bachmann et al. [13, Lem. 11, iv)]. (V5) Identity gyroautomorphism. In each of the three special cases when (i) x = 0, or (i) y = 0, or (i) x and y are parallel in V, x∥ y, we have gyr[0,x](z) (1) = z,(B.87) gyr[x,0](z) (2) = z,(B.88) 311 B.3. Gyrogroup Batch Normalization gyr[x,y](z) (3) = z, x∥ y,(B.89) where (1–3) come from A st x + B st y = 0 in Eq. (2.146). Therefore, we have gyr[s⊙ K x,t⊙ K x] = gyr[α (x) s x,α (x) t x] = id.(B.90) Note that (1–2) are also implied by the first gyrogroup theorem [200, Thm. 2.10]. B.3.12 Proof of Thm. 90 This proof largely follows the one for Thm. 81. Proof. Norm invariance under gyrations. By Bachmann et al. [13, Lem. 11], any gyration preserves the Euclidean norm: ∥ gyr[a,b](x)∥ =∥x∥, ∀x∈ st n K .(B.91) For the gyronorm, we further have ∥gyr[a,b]x∥ gyr =∥ Log 0 (gyr[a,b]x)∥ 0 = 2 √ |K| tan −1 K p |K|∥ gyr[a,b]x∥ = 2 √ |K| tan −1 K p |K|∥x∥ =∥x∥ gyr , (B.92) with tan K = tanh for K < 0 and tan K = tan for K > 0. Isometry of left gyrotranslation and gyroinverse. As shown in Thm. 89, st n K forms a gyrocommutative gyrogroup. By Thm. 80, the left gyrotranslation and gyroinverse are gyroisometries. B.3.13 Proof of Thm. 91 Proof. Non-singular cases. Denoting ∥x∥ s =∥x s ∥,∀x∈M n K , we have (t Log 0 x) s = t cos −1 K ( p |K|x t ) p |K|∥x s ∥ x s ,(B.93) ∥t Log 0 x∥ s =|t| cos −1 K ( p |K|x t ) p |K| ,(B.94) 312 Appendix B. Proofs p |K|∥t Log 0 x∥ s =|t| cos −1 K ( p |K|x t ),(B.95) For t̸= 0, the normalized spatial direction satisfies (t Log 0 x) s ∥t Log 0 x∥ s = sgn(t) x s ∥x s ∥ .(B.96) Since cos K is even and sin K is odd, substituting the above expressions into Exp 0 in Tab. 2.14 recovers the signed argument t cos −1 K ( p |K|x t ) in Eq. (3.55). The cases t = 0 and x =0 are handled by the first branch of Eq. (3.55). For t =−1, we have −1⊙ M K x = 1 p |K| cos K − cos −1 K ( p |K|x t ) sin K − cos −1 K ( √ |K|x t ) ∥x s ∥ x s (1) = x t − sin K cos −1 K ( √ |K|x t ) √ |K|∥x s ∥ x s , (B.97) where (1) comes from cos K (−θ) = cos K (θ) and sin K (−θ) = − sin K (θ). The rest is to show sin K cos −1 K ( √ |K|x t ) √ |K|∥x s ∥ = 1: sin K cos −1 K ( p |K|x t ) p |K|∥x s ∥ (1) = p sign(K)(1−|K|x 2 t ) p |K|∥x s ∥ (2) = 1.(B.98) The derivation above is based on the following: (1) cos 2 K (θ) + sign(K) sin 2 K (θ) = 1 implies sin K (θ) = q sign(K)(1− cos 2 K (θ)),∀θ > 0.(B.99) Considering cos −1 K (θ)≥ 0 for any θ ∈ dom(cos −1 K (·)), we have (1). (2) ⟨x,x⟩ K = sign(K)x 2 t +∥x s ∥ 2 = 1 K ⇒ K∥x s ∥ 2 = 1−|K|x 2 t ⇒|K|∥x s ∥ 2 = sign(K)(1−|K|x 2 t ). (B.100) 313 B.3. Gyrogroup Batch Normalization Equivalence in the singular cases. By Eq. (3.50), √ K∥u∥ = √ K∥x s ∥ 1 + √ Kx t .(B.101) Writing θ = cos −1 ( √ Kx t ) or √ Kx t = cosθ, the sphere constraint gives √ K∥x s ∥ = sinθ. Hence √ K∥u∥ = sinθ 1 + cosθ = tan θ 2 , ⇒ tan −1 √ K∥u∥ = θ 2 .(B.102) Therefore, (i) is equivalent to (i) since t tan −1 ( √ K∥u∥) = tθ 2 . Next, we use the closed form of gyromultiplication Eq. (3.58). For K > 0, the gyromultiplication reads t⊙ M K x = 1 √ K cos (tθ) sin(tθ) ∥x s ∥ x s .(B.103) If (i) holds, then sin(tθ) = 0 and cos(tθ) = −1, whence t⊙ M K x = [−1/ √ K, 0] ⊤ = − 0. B.3.14 Proof of Thm. 92 Proof. We denote x⊕ M K y = [z t ,z ⊤ s ] ⊤ . As the results are trivial under x = 0 or y =0, we assume x̸= 0 and y ̸=0 in the following. Non-singular cases. We first consider the non-singular case: i) K < 0; i) K > 0, x,y ̸= ±0, and u ̸= v K∥v∥ 2 . Since gyroadditions on M n K and st n K are both defined by Eq. (2.77), we have the following under isometries [159, Lem. 2.2]: x⊕ M K y = π st n K →M n K π M n K →st n K (x) ⊕ K π M n K →st n K (y) .(B.104) Following Eq. (B.70), we rewrite the gyroaddition on the stereographic model as u⊕ K v = (1− 2K⟨u,v⟩− K∥v∥ 2 )u + (1 + K∥u∥ 2 )v 1− 2K⟨u,v⟩ + K 2 ∥u∥ 2 ∥v∥ 2 = Au + Bv ∆ ,(B.105) where A = 1− 2K⟨u,v⟩− K∥v∥ 2 , B = 1 + K∥u∥ 2 ,∆ = 1− 2K⟨u,v⟩ + K 2 ∥u∥ 2 ∥v∥ 2 . (B.106) 314 Appendix B. Proofs Using ⟨u,v⟩ = s xy /(ab), ∥u∥ 2 = n x /a 2 , ∥v∥ 2 = n y /b 2 , one obtains A = ab 2 − 2Kbs xy − Kan y ab 2 , B = a 2 + Kn x a 2 ,∆ = D a 2 b 2 .(B.107) Hence st n K ∋ w = u⊕ K v = Au + Bv ∆ = 1 ∆ A a x s + B b y s .(B.108) Apply π st n K →M n K to w: z t = 1 p |K| 1− K∥w∥ 2 1 + K∥w∥ 2 , z s = 2w 1 + K∥w∥ 2 .(B.109) A direct expansion (radius_gyroaddition.py) yields ∥w∥ 2 = A 2 ∥u∥ 2 + B 2 ∥v∥ 2 + 2AB⟨u,v⟩ ∆ 2 = N D .(B.110) Substituting the above into Eq. (B.109) gives z t = 1 p |K| 1− KN/D 1 + KN/D = 1 p |K| D− KN D + KN .(B.111) For z s , using Eq. (B.108), Eq. (B.107) and 1 + K∥w∥ 2 = (D + KN )/D, z s = 2 1 + KN/D · 1 ∆ A a x s + B b y s = 2 D + KN (Aab 2 )x s + (Ba 2 b)y s = 2 (A s x s + A y y s ) D + KN . (B.112) Gyroaddition in singular cases (K > 0). We show that the definition Eq. (3.59) indeed returns − 0 in this case. Step 1: Log 0 (y). Write R = 1/ √ K and choose the polar angle θ ∈ (0,π) so that x = " R cosθ R sinθˆs # , y = " −R cosθ R sinθˆs # ,ˆs = x s ∥x s ∥ .(B.113) 315 B.3. Gyrogroup Batch Normalization Using Log 0 in Tab. 2.14, Log 0 (y) = 0 cos −1 ( √ Ky t ) √ K∥y s ∥ y s = " 0 (π− θ)Rˆs # .(B.114) Step 2: Parallel transport. With 1 + √ Kx t = 1 + cosθ ̸= 0 (since x̸=−0) and PT 0→x in Tab. 2.14, we have PT 0→x (Log 0 (y)) = " 0 (π− θ)Rˆs # − K⟨x s , (π− θ)Rˆs⟩ 1 + √ Kx t " x t + 1 √ K x s # .(B.115) Using ⟨x s , ˆs⟩ = ∥x s ∥ = R sinθ, KR 2 = 1, and x t = R cosθ, the scalar factor in Eq. (B.115) equals K(π− θ)R⟨x s , ˆs⟩ 1 + √ Kx t = (π− θ) sinθ 1 + cosθ .(B.116) A short simplification then yields the compact form w = PT 0→x (Log 0 (y)) = R(π− θ) " − sinθ cosθˆs # .(B.117) Note that ∥w∥ x = R(π− θ), so with α = √ K∥w∥ x we have α = π− θ. Step 3: Exponential at x. By Tab. 2.14, Exp x (w) = cos(α)x + sin(α) α w, α = √ K∥w∥ x ,(B.118) and here cos(α) = cos(π − θ) = − cosθ, sin(α) = sin(π − θ) = sinθ. Substituting Eq. (B.113) and Eq. (B.117) into Eq. (B.118), Exp x (w) = R " − cosθ " cosθ sinθˆs # + sinθ " − sinθ cosθˆs ## = R " −1 0 # =−0.(B.119) We conclude that x⊕ M K y =− 0 in the singular configuration. Equivalence in singular cases (K > 0). Recalling the isometry between M n K and st n K , we write u = x s a and v = y s b . 316 Appendix B. Proofs (i) ⇒ (i). Write α = K∥v∥ 2 > 0. From u = v α and Eq. (3.51) we get y t = 1 √ K 1− α 1 + α , x t = 1 √ K 1− 1 α 1 + 1 α =− 1 √ K 1− α 1 + α =−y t , x s = 2u 1 + K∥u∥ 2 = 2v/α 1 + K∥v∥ 2 /α 2 = 2v 1 + α = y s , (B.120) Hence x s = y s and x t =−y t . (i) ⇒ (i). Under x s = y s and x t =−y t , set n =∥x s ∥ 2 =∥y s ∥ 2 and note s xy = n, n x = n y = n. Then, D = a 2 b 2 − 2Kabs xy + K 2 n x n y = (ab− Kn) 2 .(B.121) On the sphere S n K , we have the constraint x 2 t + ∥x s ∥ 2 = 1 /K. Since y t = −x t and ∥y s ∥ 2 =∥x s ∥ 2 = n, ab = (1 + √ Kx t )(1 + √ Ky t ) = (1 + √ Kx t )(1− √ Kx t ) = 1− Kx 2 t = 1− K 1 K − n = Kn. (B.122) Therefore D = (ab− Kn) 2 = 0. (i) ⇒ (i). This can be obtained using the Cauchy–Schwarz inequality, as in Bach- mann et al. [13, App. C.2.1]. Consequences in singular cases (K > 0). Under (i) we also have a + b = 2, so N = a 2 n + 2abn + b 2 n = (a + b) 2 n = 4n > 0.(B.123) Using D = 0 and N > 0, Eq. (B.111) gives z t = 1 √ K −KN KN =− 1 √ K .(B.124) Besides, since for x s = y s one checks A s + A y = (a + b)(ab− Kn) = 0, Eq. (B.112) yields z s = 0. On the other hand, u⊕ K v =∞. 317 B.3. Gyrogroup Batch Normalization B.3.15 Proof of Thm. 93 Proof. This is already implied by the proofs of Thms. 91 and 92. B.3.16 Proof of Thm. 95 Proof. The gyrovector space over the stereographic model is also defined by Eqs. (2.77) and (2.78): x⊕ K y = Exp x (PT 0→x (Log 0 (y))), r⊙ K x = Exp 0 (r Log 0 (x)), (B.125) with x,y ∈ st n K and r ∈ R. Since π M n K →st n K ( 0) = 0 and (st n K ,⊕ K ,⊙ K ) is a gyrovector space, (M n K ,⊕ M K ,⊙ M K ) is also a gyrovector space [159, Thm. 2.4]. B.3.17 Proof of Thm. 97 Proof. Isometries. The Beltrami–Klein model is isometric to the hyperboloid model by the following diffeomorphisms [131, Thm. 3.7]: π K n K →H n K : K n K ∋ x7−→ 1 √ −K p 1 + K∥x∥ 2 , x p 1 + K∥x∥ 2 ! ∈ H n K ,(B.126) π H n K →K n K : H n K ∋ " x t x s # 7−→ x s √ −Kx t ∈ K n K .(B.127) Combining the isometries Eqs. (B.126), (B.127), (3.50) and (3.51), one can readily obtain the isometries between Beltrami–Klein and Poincaré ball models. Now, we turn to the differential maps. Given a curve over c(t)∈M with c(0) = x and c ′ (0) = v, the 318 Appendix B. Proofs differential maps can be calculated by dπ K n K →P n K (c(t)) dt t=0 = d dt c(t) 1 + p 1 + K∥c(t)∥ 2 t=0 = 1 + p 1 + K∥x∥ 2 v− (1 + K∥x∥ 2 ) − 1 2 K⟨x,v⟩x 1 + p 1 + K∥x∥ 2 2 = 1 1 + p 1 + K∥x∥ 2 v− K⟨x,v⟩ 1 + p 1 + K∥x∥ 2 2 p 1 + K∥x∥ 2 x, dπ P n K →K n K (c(t)) dt t=0 = d dt 2c(t) 1− K∥c(t)∥ 2 t=0 = 2v (1− K∥x∥ 2 ) + 4K⟨x,v⟩x (1− K∥x∥ 2 ) 2 = 2 (1− K∥x∥ 2 ) v + 4K⟨x,v⟩ (1− K∥x∥ 2 ) 2 x. (B.128) Homomorphism. The homomorphism w.r.t. the scalar product can be readily obtained by the Riemannian isometry. Therefore, we first address it before proceeding to the addition. For simplicity, we denote φ = π P n K →K n K . Scalar product. As shown by Ungar [200], the geodesics under the Beltrami–Klein and Poincaré ball models are γ K φ(x)→φ(y) (t) = φ(x)⊕ E t⊙ E (−φ(x)⊕ E φ(y)), γ P x→y (t) = x⊕ M t⊙ M (−x⊕ M y). (B.129) Here, we use the fact that the gyro inverses in the Möbius and Einstein gyrovector spaces are exactly the familiar vector inverse. The above geodesics satisfy φ(x⊕ M t⊙ M (−x⊕ M y)) = φ(x)⊕ E t⊙ E (−φ(x)⊕ E φ(y)). (B.130) The above comes from φ(γ P x→y (t)) = γ K φ(x)→φ(y) (t). In particular, the geodesic starting from the identity element yields φ(t⊙ M x) = φ(γ P 0→x (t)) = γ K φ(0)→φ(x) (t) (1) = t⊙ E φ(x), (B.131) where (1) comes from φ(0) = 0. The above holds for all x ∈ P n K and ∀t ∈ R, as every hyperbolic geometry is geodesically complete [131, p. 139]. 319 B.3. Gyrogroup Batch Normalization Addition. We first expand the LHS of Eq. (3.69). Inspired by Mao et al. [147, Eqs. 73–76], we express the Möbius addition as x⊕ M y = Bx+Cy A by denoting A = 1− 2K⟨x,y⟩ + K 2 ∥x∥ 2 ∥y∥ 2 , B = 1− 2K⟨x,y⟩− K∥y∥ 2 , C = 1 + K∥x∥ 2 . (B.132) Then, the Möbius addition is φ(x⊕ M y) = 2 Bx+Cy A 1− K Bx+Cy A 2 = 2ABx + 2ACy A 2 − K∥Bx + Cy∥ 2 = 2ABx + 2ACy A 2 − B 2 K∥x∥ 2 − C 2 K∥y∥ 2 − 2BCK⟨x,y⟩ . (B.133) Inspired by Mao et al. [147, Eq. 77], we denote ∥x∥ = a,∥y∥ = b, and ⟨x,y⟩ = ab cos(θ). In this way, we can resort to the symbolic computation package SymPy [150] for the heavy algebra computation, which brings φ(x⊕ M y) = 2 (−2Kab cos (θ)− Kb 2 + 1) K 2 a 2 b 2 − Ka 2 − 4Kab cos (θ)− Kb 2 + 1 x + 2Ka 2 + 2 K 2 a 2 b 2 − Ka 2 − 4Kab cos (θ)− Kb 2 + 1 y. (B.134) Now we turn to the RHS of Eq. (3.69). For any u,v ∈ K n K , the Einstein addition can be rewritten as u⊕ E v = 1 1− K⟨u,v⟩ u + 1 γ u v− K γ u 1 + γ u ⟨u,v⟩u = 1− K γ u 1+γ u ⟨u,v⟩ 1− K⟨u,v⟩ u + 1 γ u (1− K⟨u,v⟩) v. (B.135) 320 Appendix B. Proofs The gamma factor and inner product under isometry can be rewritten as γ φ(x) = 1 p 1 + K∥φ(x)∥ 2 = 1 r 1 + K 2 1−K∥x∥ 2 2 ∥x∥ 2 = 1− K∥x∥ 2 q (1− K∥x∥ 2 ) 2 + 4K∥x∥ 2 (1) = 1− K∥x∥ 2 1 + K∥x∥ 2 , ⟨φ(x),φ(y)⟩ = 4 (1− K∥x∥ 2 )(1− K∥y∥ 2 ) ⟨x,y⟩. (B.136) Here, (1) comes from∥x∥ 2 <− 1 K ⇒ K∥x∥ 2 + 1 > 0. Putting the above into Eq. (B.135) and following the same notation as Eq. (B.134), we can obtain the following by SymPy: φ(x)⊕ E φ(y) = 2· (2Kab cos (θ) + Kb 2 − 1) 4Kab cos (θ)− (Ka 2 − 1) (Kb 2 − 1) x + −2Ka 2 − 2 4Kab cos (θ)− (Ka 2 − 1) (Kb 2 − 1) y, (B.137) which is clearly equal to Eq. (B.134). B.3.18 Proof of Thm. 98 Proof. For simplicity, we denote φ = π P n K →K n K . The homomorphism and bijection of φ imply the homomorphism of its inverse φ −1 . Also note that φ(0) = φ −1 (0) = 0. This yields x⊕ E y = φ φ −1 (x)⊕ M φ −1 (y) (1) = φ Exp P φ −1 (x) PT P 0→φ −1 (x) Log P 0 (φ −1 (y)) (2) = Exp K x PT K 0→x Log K 0 (y) . (B.138) The above comes from the following. (1) For any Poincaré vectors u,w ∈ P n K , the following holds: PT P 0→u (v) = Log P u (u⊕ M Exp P 0 (v))⇒ u⊕ M w = Exp P u PT P 0→u (Log P 0 (w)) , (B.139) where v = Log P 0 (w) and the LHS comes from Ganea et al. [76, Thm. 4]; 321 B.3. Gyrogroup Batch Normalization (2) The second identity follows from the isometry: Exp P x (v) = φ −1 Exp K φ(x) (φ ∗,x (v)) , Log P x (y) = φ −1 ∗,x Log K φ(x) (φ(y)) , PT P x→y (v) = φ −1 ∗,y PT K φ(x)→φ(y) (φ ∗,x (v)) . (B.140) Similar to the gyro addition, the isometry of φ implies the following with respect to the gyro scalar product: t⊙ E x = φ t⊙ M φ −1 (x) (1) = φ Exp P 0 t Log P 0 (φ −1 (x)) (2) = Exp K 0 t Log K 0 (x) . (B.141) The above comes from the following. (1) Ganea et al. [76, Lem. 3]; (2) The isometry of φ. B.3.19 Proof of Thm. 99 Proof. Following Thm. 98, we denote φ = π P n K →K n K and ψ = π K n K →P n K . The results can be obtained by the properties of isometry and isomorphism of φ. Riemannian exponential and logarithmic maps at the zero vector. First, we recall that the expressions of Exp P 0 (v) and Log P 0 (v) under the Poincaré ball model are exactly those in Eqs. (3.72) and (3.73) [76, Eq. 13]. 322 Appendix B. Proofs By Riemannian isometry, we have the following: Exp K 0 (v) = φ Exp P 0 (ψ ∗,0 (v)) = φ tanh √ −K 2 ∥v∥ v √ −K∥v∥ = 2 1− K∥v∥ 2 tanh √ −K 2 ∥v∥ √ −K∥v∥ 2 tanh √ −K 2 ∥v∥ √ −K∥v∥ v = 2 tanh √ −K 2 ∥v∥ 1 + tanh √ −K 2 ∥v∥ 2 v √ −K∥v∥ (1) = tanh( √ −K∥v∥) v √ −K∥v∥ , (B.142) where (1) comes from tanh(2x) = 2 tanhx 1+tanh 2 x . As the inverse of Exp K 0 (v), Log K 0 (x), therefore, shares the expression with its coun- terpart under the Poincaré ball model. Here, we use the properties of isometry to further validate this result: Log K 0 (x) = φ ∗,0 Log P 0 (ψ(x)) = 2 tanh −1 √ −K∥x∥ 1 + q 1 + K∥x∥ 2 x √ −K∥x∥ . (B.143) We only need to show 2 tanh −1 √ −K∥x∥ 1 + p 1 + K∥x∥ 2 ! = tanh −1 ( √ −K∥x∥).(B.144) We denote a = √ −K∥x∥ < 1. Both sides are LHS: 2 tanh −1 a 1 + √ 1− a 2 = ln 1 + a 1+ √ 1−a 2 1− a 1+ √ 1−a 2 ! = ln 1 + √ 1− a 2 + a 1 + √ 1− a 2 − a , RHS: tanh −1 (a) = ln √ 1 + a √ 1− a = ln √ 1− a 2 1− a . (B.145) 323 B.3. Gyrogroup Batch Normalization Note that the below equation holds: 1 + √ 1− a 2 + a 1 + √ 1− a 2 − a = √ 1− a 2 1− a .(B.146) Geodesic distances. d K (x,y) (1) = d P (ψ(x),ψ(y)) = 2 √ −K tanh −1 √ −K∥−ψ(x)⊕ M ψ(y)∥ (2) = 2 √ −K tanh −1 √ −K∥ψ (−x⊕ E y)∥ = 2 √ −K tanh −1 √ −K ∥−x⊕ E y∥ 1 + q 1 + K∥−x⊕ E y∥ 2 , (B.147) where (1) comes from the isometry, while (2) comes from the isomorphism. Exponential maps. Exp K x (v) (1) = ψ −1 Exp P ψ(x) (ψ ∗,x (v)) (2) = ψ −1 ψ(x)⊕ M Exp 0 λ K ψ(x) 2 ψ ∗,x (v) !! (3) = x⊕ E ψ −1 Exp 0 λ K ψ(x) 2 ψ ∗,x (v) !! (4) = x⊕ E Exp 0 λ K ψ(x) ψ ∗,x (v) . (B.148) The above comes from the following. (1) The isometry of ψ; (2) Exp P x (v) = x⊕ M Exp 0 λ K x 2 v ; (3) The isomorphism of ψ; (4) ψ −1 ◦ Exp 0 (v) = ψ −1 ◦ Exp P 0 (v) = ψ −1 ◦ ψ◦ Exp K 0 (2v) = Exp 0 (2v). (B.149) 324 Appendix B. Proofs It remains to calculate λ K ψ(x) ψ ∗,x (v): λ K ψ(x) = 2 1 + K ∥x∥ 2 1+ √ 1+K∥x∥ 2 2 = 2 1 + p 1 + K∥x∥ 2 2 1 + p 1 + K∥x∥ 2 2 + K∥x∥ 2 = 2 1 + p 1 + K∥x∥ 2 2 2 + 2 p 1 + K∥x∥ 2 + 2K∥x∥ 2 = 1 + p 1 + K∥x∥ 2 2 1 + p 1 + K∥x∥ 2 + K∥x∥ 2 = 1 + p 1 + K∥x∥ 2 p 1 + K∥x∥ 2 , (B.150) λ K ψ(x) ψ ∗,x (v) = λ K ψ(x) 1 1 + p 1 + K∥x∥ 2 v− K⟨x,v⟩ 1 + p 1 + K∥x∥ 2 2 p 1 + K∥x∥ 2 x = 1 p 1 + K∥x∥ 2 v− K⟨x,v⟩ 1 + p 1 + K∥x∥ 2 (1 + K∥x∥ 2 ) x. (B.151) Logarithmic maps. Log K x (y) (1) = φ ∗,ψ(x) Log P ψ(x) ψ(y) (2) = φ ∗,ψ(x) 2 λ K ψ(x) Log 0 (−ψ(x)⊕ M ψ(y)) ! (3) = 2 λ K ψ(x) φ ∗,ψ(x) (Log 0 (−ψ(x)⊕ M ψ(y))) (4) = 1 λ K ψ(x) φ ∗,ψ(x) (Log 0 (−x⊕ E y)). (B.152) The above comes from the following. (1) The isometry of ψ; (2) Log P x (y) = 2 λ K x Log 0 (−x⊕ M y); (3) The linearity of differential maps; 325 B.4. SPD Multinomial Logistic Regression (4) Log 0 (−ψ(x)⊕ M ψ(y)) = Log P 0 (−ψ(x)⊕ M ψ(y)) = 1 2 Log K 0 ψ −1 (−ψ(x)⊕ M ψ(y)) = 1 2 Log K 0 (−x⊕ E y). (B.153) Remark 192. Mao et al. [147, Thm. 9] extended the Möbius matrix-vector mul- tiplication [76, Lem. 6] to the Einstein space of the Beltrami–Klein model under K = −1, namely Exp 0 (M Log 0 (x))). Although their presented formulations are different from the Möbius one, this theorem indicates that matrix-vector multi- plications over these two spaces are identical under any negative curvature. The equality can also be readily observed from Eqs. (B.145) and (B.146). B.4 SPD Multinomial Logistic Regression B.4.1 Proof of Thm. 104 This claim can be proven either by definition [197, Def. 9.1] or by the constant rank level set theorem [197, Thm. 11.2]. We focus on the latter. Proof. Consider any P ∈S n ++ and A∈ T P S n ++ \0. Define the function f :S n ++ → R, S 7→⟨Log P S,A⟩ P .(B.154) For the SPD hyperplane ̃ H A,P , we have ̃ H A,P = f −1 (0). By assumption, Log P :S n ++ → T P S n ++ is a global diffeomorphism, and f is therefore well-defined. We can rewrite f as a composition, i.e., f = h◦ Log P , where h(·) =⟨·,A⟩ P is a linear map. Since Log P is a diffeomorphism and h(·) is a nonzero linear map, the rank of f is globally constant. Therefore, there exists a neighborhood, e.g., the whole SPD manifold, of f −1 (0) where the rank of f is constant. According to the constant rank level set theorem [197, Thm. 11.2], we can obtain the claim. 326 Appendix B. Proofs B.4.2 Proof of Thm. 105 Proof. By Thm. 149, we have ⟨Log P Q,A⟩ P = D φ ∗,P φ −1 ∗,φ(P ) (φ(Q)− φ(P )),φ ∗,P A E (B.155) =⟨φ(Q)− φ(P ),φ ∗,P A⟩.(B.156) Therefore, due to the isometry of φ, the SPD hyperplane ̃ H A k ,P k corresponds to the Euclidean hyperplane H φ ∗,P k (A k ),φ(P k ) .(B.157) Furthermore, the distance to the margin hyperplane is equivalent to inf φ(Q) ∥φ(S)− φ(Q)∥ F (B.158) s.t.⟨φ(Q)− φ(P k ),φ ∗,P k A k ⟩ = 0.(B.159) The problem above is the familiar Euclidean distance from a point to a hyperplane. By simple computation, one can obtain the result. B.4.3 Proof of Thm. 106 Proof. For simplicity, we abbreviate ⊙ φ and g φ as ⊙ and g. By abuse of notation, we further denote Q⊙ P −1 ⊙ as QP −1 , where P −1 ⊙ is the inverse of P under ⊙. According to Thm. 149, (S n ++ ,⊙) is an abelian group and g is a bi-invariant Riemannian metric. By Lin [137, Lem. 6], any parallel transport can be expressed as a differential of left translation, PT P→Q = (L QP −1 ) ∗,P , ∀P,Q∈S n ++ .(B.160) B.4.4 Proof of Thm. 107 Proof. By Eq. (6.9), parallel transport under a pullback Euclidean metric is path- independent and satisfies PT Q→P (V ) = φ −1 ∗,φ(P ) (φ ∗,Q (V )).(B.161) 327 B.4. SPD Multinomial Logistic Regression For any ̃ A 1,k ∈ T Q 1 S n ++ , define ̃ A 2,k = φ −1 ∗,φ(Q 2 ) φ ∗,Q 1 ̃ A 1,k ∈ T Q 2 S n ++ .(B.162) Since φ is a diffeomorphism, its differential at every point is a linear isomorphism. Hence, PT Q 2 →P k ̃ A 2,k = φ −1 ∗,φ(P k ) φ ∗,Q 2 ̃ A 2,k = φ −1 ∗,φ(P k ) φ ∗,Q 1 ̃ A 1,k = PT Q 1 →P k ̃ A 1,k . (B.163) Moreover, if ̃ B 2,k ∈ T Q 2 S n ++ satisfies the same equality, applying φ ∗,P k to both sides gives φ ∗,Q 2 ̃ B 2,k = φ ∗,Q 1 ̃ A 1,k .(B.164) The invertibility of φ ∗,Q 2 then yields ̃ B 2,k = ̃ A 2,k , proving uniqueness. B.4.5 Proof of Thm. 108 Proof. A k = PT I→P k ( ̃ A k )(B.165) = φ −1 ∗,φ(P k ) φ ∗,I ( ̃ A k ) .(B.166) One can obtain the result by putting Eq. (B.166) into Eq. (4.10). B.4.6 Proof of Thm. 109 Proof. Denoting the matrix power as P θ :S n ++ →S n ++ , we have P θ (I) = I,(B.167) (P θ ) ∗,I (A) = θA, ∀A∈ T I S n ++ .(B.168) Next, we prove the two cases separately. (α,β)-LEM. We define the following map: ψ LEM = f ◦ log,(B.169) where f :S n →S n is the linear isometry between the standard Frobenius inner product 328 Appendix B. Proofs and the O(n)-invariant inner product ⟨·,·⟩ (α,β) . Then ψ LEM pulls back the standard Euclidean metric on S n to (α,β)-LEM on S n ++ . Putting Eqs. (B.168) and (B.169) into Eq. (4.12), we have exp D ψ LEM (S)− ψ LEM (P k ),ψ LEM ∗,I ( ̃ A k ) E = exp hD f (log(S)− log(P k )),f ( ̃ A k ) Ei = exp D log(S)− log(P k ), ̃ A k E (α,β) , (B.170) where the last equality follows from the linearity and isometry property of f, with f = f ∗ . θ-LCM. We denote ψ LCM = Dlog◦ Chol◦ P θ .(B.171) Then ψ LCM pulls back the Euclidean metric 1 θ 2 g E on the Euclidean space LT n of lower triangular matrices to θ-LCM on S n ++ . The differential of Cholesky decomposition is presented by Lin [137, Prop. 4], while the differential of Dlog is obtained as the natural- logarithm specialization of Thm. 154. Then, simple computations show that ψ LCM ∗,I (A) = θ ⌊A⌋ + 1 2 D(A) , ∀A∈ T I S n ++ .(B.172) Putting Eqs. (B.171) and (B.172) into Eq. (4.12), we can obtain the result. B.4.7 Proof of Thm. 111 To prove Thm. 111, we first present two lemmas about the general cases under pullback Euclidean metrics. One can observe that Eqs. (4.12) and (4.13) are very similar to those of a Euclidean MLR. However, since φ is normally nonlinear and P k is an SPD parameter, Eq. (4.12) cannot hastily be identified with a Euclidean MLR. Under some special circumstances, SPD MLR can be reduced to the familiar Euclidean MLR. To show this result, we first specialize the RSGD strategy reviewed in Sec. 2.7 to pullback Euclidean metrics. In ambient-gradient notation, the required update is W t+1 = Exp W t (−γ t Π W t (∇ W t f )),(B.173) where Π W t denotes the projection mapping the Euclidean gradient ∇ W t f to the Rie- 329 B.4. SPD Multinomial Logistic Regression mannian gradient, and γ t denotes the learning rate. We have already obtained the formula for the Riemannian exponential map in Eq. (6.7). We proceed to formulate Π. Lemma 193. For a smooth function f :S n ++ → R on S n ++ endowed with any kind of pullback Euclidean metric, the projection map Π P :S n → T P S n ++ at P ∈S n ++ is Π P (∇ P f ) = φ −1 ∗,φ(P ) φ −∗ ∗,φ(P ) (∇ P f ) ,(B.174) where φ −∗ ∗,φ(P ) := φ −1 ∗,φ(P ) ∗ : T P S n ++ → T φ(P ) S n is the Frobenius adjoint of φ −1 ∗,φ(P ) : T φ(P ) S n → T P S n ++ . Specifically, for all U ∈ T P S n ++ and Z ∈ T φ(P ) S n , it satisfies D U,φ −1 ∗,φ(P ) (Z) E = D φ −∗ ∗,φ(P ) (U ),Z E .(B.175) Proof. Given any smooth function f :S n ++ → R, denote its Riemannian gradient at P as grad P f ∈ T P S n ++ . ⟨grad P f,V⟩ P = f ∗,P (V ), ∀V ∈ T P S n ++ .(B.176) Let ∇ P f denote the Euclidean gradient in the canonical chart. For any Z ∈ T φ(P ) S n , set V = φ −1 ∗,φ(P ) (Z). By Eq. (6.5) and the definition of the Frobenius adjoint, we have ⟨φ ∗,P (grad P f ),Z⟩ = D grad P f,φ −1 ∗,φ(P ) (Z) E P = f ∗,P φ −1 ∗,φ(P ) (Z) = D ∇ P f,φ −1 ∗,φ(P ) (Z) E = D φ −∗ ∗,φ(P ) (∇ P f ),Z E . (B.177) Since Z is arbitrary, the nondegeneracy of the Frobenius inner product gives φ ∗,P (grad P f ) = φ −∗ ∗,φ(P ) (∇ P f ).(B.178) Applying φ −1 ∗,φ(P ) to both sides yields grad P f = φ −1 ∗,φ(P ) φ −∗ ∗,φ(P ) (∇ P f ) .(B.179) By definition, Π P (∇ P f ) = grad P f, which yields Eq. (B.174). We can describe the special case mentioned above with this lemma. 330 Appendix B. Proofs Lemma 194. Suppose the differential map φ ∗,I is the identity map, and P k in Eq. (4.12) is optimized by pullback-Euclidean-metric-based RSGD. Then Eq. (4.12) can be reduced to a Euclidean MLR in the codomain of φ updated by Euclidean SGD. Proof. Define a Euclidean MLR in the codomain of φ as p(y = k | S)∝ exp φ(S)− ̄ P k , ̄ A k ,(B.180) where ̄ P k , ̄ A k ∈S n . We call this classifier φ-EMLR. Define the SPD MLR under the pullback Euclidean metric induced by φ as p(y = k | S)∝ exp D φ(S)− φ(P k ), ̃ A k E ,(B.181) where P k ∈S n ++ and ̃ A k ∈S n . Suppose the SPD MLR and φ-EMLR satisfy ̄ P k = φ(P k ). Other settings of the network are all the same, indicating that the Euclidean gradients satisfy ∂L ∂ ̄ P k = ∂L ∂φ(P k ) .(B.182) The update of ̄ P k in the φ-EMLR is ̄ P ′ k = ̄ P k − γ ∂L ∂ ̄ P k .(B.183) The update of P k in the SPD MLR is P ′ k = Exp P k −γΠ P k ∂L ∂P k (B.184) = φ −1 φ(P k )− γφ −∗ ∗,P k ∂L ∂P k .(B.185) Therefore, φ(P ′ k ) satisfies φ(P ′ k ) = φ(P k )− γφ −∗ ∗,P k ∂L ∂P k (B.186) = φ(P k )− γφ −∗ ∗,P k φ ∗ ∗,P k ∂L ∂φ(P k ) (B.187) = φ(P k )− γ ∂L ∂φ(P k ) (B.188) 331 B.5. Riemannian Multinomial Logistic Regression = ̄ P ′ k .(B.189) Eq. (B.187) comes from the Euclidean chain rule of the differential. Let Y = φ(X). Then we have ∂L ∂Y : dY = ∂L ∂Y : φ ∗,X dX = φ ∗ ∗,X ∂L ∂Y : dX,(B.190) where : denotes the Frobenius inner product. The equivalence of ̄ A k and ̃ A k is obvious. By mathematical induction, the claim can be proven. Proof. The claim follows directly from Thm. 194. B.5 Riemannian Multinomial Logistic Regression B.5.1 Proof of Thm. 113 Proof. Let us first solve Y ∗ in Eq. (4.22), which is the solution to the following con- strained optimization problem: max Y ⟨Log P Y, Log P S⟩ P ∥ Log P Y∥ P ∥ Log P S∥ P s.t.⟨Log P Y, ̃ A⟩ P = 0.(B.191) Note that Eq. (B.191) is well-defined due to the existence of the Riemannian logarithm. Although Eq. (B.191) is normally non-convex, Eq. (B.191) and Eq. (4.22) can be reduced to a Euclidean problem: max ̃ Y ⟨ ̃ Y , ̃ S⟩ P ∥ ̃ Y∥ P ∥ ̃ S∥ P s.t.⟨ ̃ Y , ̃ A⟩ P = 0,(B.192) d(S, ̃ H ̃ A,P ) = sin(∠SPY ∗ )∥ ̃ S∥ P ,(B.193) where ̃ Y = Log P Y and ̃ S = Log P S. Let us first discuss Eq. (B.192). Denote the solution of Eq. (B.192) as ̃ Y ∗ . Note that ̃ Y ∗ is not necessarily unique. Note that Exp P is only well-defined locally. More precisely, Exp P is well-defined in an open ball B ε (0) centered at 0∈ T P M. Therefore, ̃ Y ∗ might not be in B ε (0). In this case, we can scale ̃ Y ∗ into B ε (0), and the scaled ̃ Y ∗ is still the maximizer of Eq. (B.192). Therefore, without loss of generality, we assume ̃ Y ∗ ∈ B ε (0). Putting ̃ Y ∗ into Eq. (B.193), Eq. (B.193) is reduced to the distance to the hyperplane 332 Appendix B. Proofs ⟨ ̃ Y , ̃ A⟩ P = 0 in the Euclidean space (T P M,⟨·,·⟩ P ), which has a closed-form solution: d(S, ̃ H ̃ A,P ) = |⟨ ̃ S, ̃ A⟩ P | ∥ ̃ A∥ P (B.194) = |⟨Log P S, ̃ A⟩ P | ∥ ̃ A∥ P .(B.195) B.5.2 Proof of Thm. 114 Proof. Putting the margin distance (Eq. (4.25)) into Eq. (4.18), we have the following: p(y = k | S)∝ exp sign(⟨ ̃ A k , Log P k (S)⟩ P k )∥ ̃ A k ∥ P k d(S, ̃ H ̃ A k ,P k ) = exp sign(⟨ ̃ A k , Log P k (S)⟩ P k )∥ ̃ A k ∥ P k |⟨Log P k (S), ̃ A k ⟩ P k | ∥ ̃ A k ∥ P k ! = exp ⟨Log P k S, ̃ A k ⟩ P k . (B.196) B.5.3 Proof of Thm. 116 Proof. The Riemannian metric (α,β)-EM at I is g (α,β)-EM I (V,V ) =⟨V,V⟩ (α,β) .(B.197) By Thm. 186, we have the following: g (θ,α,β)-EM P (V,V ) θ→0 −→ g (α,β)-EM I log ∗,P (V ), log ∗,P (V ) =⟨log ∗,P (V ), log ∗,P (V )⟩ (α,β) = g (α,β)-LEM P (V,V ). (B.198) B.5.4 Proof of Thm. 117 As the five families of metrics presented in Thm. 117 are pullback metrics, we first present a general result regarding Riemannian MLRs under pullback metrics. 333 B.5. Riemannian Multinomial Logistic Regression Lemma 195 (Riemannian MLRs under pullback metrics). Suppose (N,g) is a Riemannian manifold and φ : M → N is a diffeomorphism between manifolds. The Riemannian MLR by parallel transport, obtained by combining Eqs. (4.26) and (4.27), on M under ̃g = φ ∗ g can be obtained using g: p(y = k | S ∈M)∝ exp h ⟨ ̃ Log P k S, ̃ Γ Q→P k A k ⟩ P k i (B.199) = exp h ⟨Log φ(P k ) φ(S), ̃ A k ⟩ φ(P k ) i ,(B.200) where ̃ A k = Γ φ(Q)→φ(P k ) φ ∗,Q (A k ) with A k ∈ T Q M; ̃ Log and ̃ Γ are the Riemannian logarithm and parallel transport under ̃g, respectively; and Log and Γ are their counterparts under g. Furthermore, if N has a Lie group operation ⊙, M could be endowed with a Lie group structure ̃ ⊙ by φ. The Riemannian MLR by left translation, obtained by combining Eqs. (4.26) and (4.28), on M under ̃g and ̃ ⊙ can be calculated using g and ⊙: p(y = k | S ∈M)∝ exp " ̃ Log P k S, ̃ L ̃ R k ∗,Q A k P k # (B.201) = exp h ⟨Log φ(P k ) φ(S), ̃ A k ⟩ φ(P k ) i ,(B.202) where ̃ A k = (L R k ) ∗,φ(Q) (φ ∗,Q (A k )), ̃ R k = P k ̃ ⊙Q −1 ̃ ⊙ , R k = φ(P k ) ⊙ φ(Q) −1 ⊙ , and ̃ L P k ̃ ⊙Q −1 ̃ ⊙ is the left translation under ̃ ⊙. Proof. Before starting, we should point out that since φ is a diffeomorphism, ̃ ⊙ and ̃g are indeed well defined, and (M, ̃g) forms a Riemannian manifold and (M, ̃ ⊙) forms a Lie group. We first focus on the Riemannian MLR by parallel transport: p(y = k | S ∈M) ∝ exp ̃g P k ( ̃ Log P k S, ̃ Γ Q→P k A k ) = exp h g φ(P k ) φ ∗,P k ◦ φ −1 ∗,φ(P k ) Log φ(P k ) φ(S),φ ∗,P k ◦ φ −1 ∗,φ(P k ) Γ φ(Q)→φ(P k ) φ ∗,Q (A k ) i = exp g φ(P k ) (Log φ(P k ) φ(S), Γ φ(Q)→φ(P k ) φ ∗,Q (A k )) . (B.203) In the case of the Riemannian MLR by left translation, we first note that ̃ L ̃ R k = φ −1 ◦ L φ(P k )⊙φ(Q) −1 ⊙ ◦ φ.(B.204) 334 Appendix B. Proofs Therefore, the associated differential is ̃ L ̃ R k ∗,Q = φ −1 ∗,φ(P k ) ◦ (L R k ) ∗,φ(Q) ◦ φ ∗,Q .(B.205) Putting Eq. (B.205) into Eq. (B.201), we can obtain the result. Now, we apply Thm. 195 to derive the expressions for our SPD MLRs presented in Thm. 117. In our cases of SPD MLRs, we set Q = I. For simplicity, we will omit the subscript k for P k and A k . We will first derive the expressions for SPD MLRs under (θ,α,β)-LEM, θ-LCM, (θ,α,β)-EM, and (θ,α,β)-AIM from Eq. (B.200). Then we will derive the expression for MLR under 2θ-BWM from Eq. (B.202). According to Thm. 187, the scaled metric ag shares the same Riemannian operators as g. We will use this fact throughout the following proof. Proof. For simplicity, we abbreviate φ θ as φ during the proof. Note that for 2θ-BWM, φ should be understood as φ 2θ . We first show φ(I) and the differential map φ ∗,I , which will be frequently required in the following proof: φ(I) = I,(B.206) φ ∗,I (A) = θA,∀A∈ T I S n ++ .(B.207) Let φ : (S n ++ , ̃g)→ (S n ++ ,g). Then the SPD MLR under ̃g by parallel transport with Q = I is p(y = k | S ∈S n ++ )∝ exp g φ(P ) (Log φ(P ) φ(S), Γ I→φ(P ) θA) .(B.208) Next, we begin to prove the five SPD MLRs one by one. (α,β)-LEM. As shown in Thm. 147, the standard LEM is the pullback metric from the Euclidean space S n . Similarly, (α,β)-LEM is also a pullback metric: (S n ++ ,g (α,β)-LEM ) log −→ (S n ,g (α,β) ).(B.209) By Eq. (B.200), we have p(y = k | S ∈S n ++ )∝ exp ⟨log(S)− log(P ), log ∗,I (A)⟩ (α,β) (B.210) = exp ⟨log(S)− log(P ),A⟩ (α,β) .(B.211) θ-LCM. Simple computations show that θ-LCM is the scaled pullback metric of the 335 B.5. Riemannian Multinomial Logistic Regression standard Euclidean metric in the Euclidean space of lower triangular matrices LT n : (S n ++ ,θ 2 g θ-LCM ) φ −→ (S n ++ ,g LCM ) Chol −→ (L n ++ ,g CM ) Dlog −→ (LT n ,g E ),(B.212) where g E is the standard Frobenius inner product, and g CM is the Cholesky metric on the Cholesky space L n ++ [137]. Denoting ζ = Dlog◦ Chol◦φ, we have ζ ∗,I (A) = θ ⌊A⌋ + 1 2 D(A) ,∀A∈ T I S n ++ .(B.213) Similar to the case of (θ,α,β)-LEM, we have p(y = k | S ∈S n ++ )∝ exp 1 θ 2 ⟨ζ(S)− ζ(P ),ζ ∗,I A⟩ (B.214) = exp 1 θ * ⌊ ̃ K⌋−⌊ ̃ L⌋ + h Dlog(D( ̃ K))− Dlog(D( ̃ L)) i , ⌊A⌋ + 1 2 D(A) + , (B.215) where ̃ K = Chol(S θ ), ̃ L = Chol(P θ ), D( ̃ K) is a diagonal matrix with diagonal elements from ̃ K, and ⌊ ̃ K⌋ is a strictly lower triangular matrix from ̃ K. (θ,α,β)-EM. Let η = 1 |θ| φ. A simple computation shows that (θ,α,β)-EM is the pullback metric of (α,β)-EM: (S n ++ ,g (θ,α,β)-EM ) η −→ (S n ++ ,g (α,β)-EM ).(B.216) Besides, we have the following for η: η ∗,I (A) = sgn(θ)A,∀A∈ T I S n ++ .(B.217) According to Eq. (B.200), we have p(y = k | S ∈S n ++ )∝ exp ⟨η(S)− η(P ), sgn(θ)A⟩ (α,β) (B.218) = exp 1 θ ⟨S θ − P θ ,A⟩ (α,β) .(B.219) 336 Appendix B. Proofs (θ,α,β)-AIM. Putting g (α,β)-AIM into Eq. (B.208), we have p(y = k | S ∈S n ++ )∝ exp 1 θ 2 g (α,β)-AIM φ(P ) P θ 2 log P − θ 2 S θ P − θ 2 P θ 2 ,P θ 2 θAP θ 2 (B.220) = exp 1 θ D log P − θ 2 S θ P − θ 2 ,A E (α,β) .(B.221) 2θ-BWM. We first simplify Eq. (B.202) for SPD manifolds and then proceed to focus on the case of g = g BWM . Denote φ : (S n ++ , ̃g, ̃ ⊙) → (S n ++ ,g,⊙), where the Lie group operation ⊙ [195] is defined as S 1 ⊙ S 2 = L 1 S 2 L ⊤ 1 ,∀S 1 ,S 2 ∈S n ++ , with L 1 = Chol(S 1 ).(B.222) Note that I is the identity element of (S n ++ ,⊙), and for any S ∈ S n ++ , the differential map of the left translation L S under ⊙ is (L S ) ∗,Q (V ) = LV L ⊤ ,∀Q∈S n ++ ,∀V ∈ T Q S n ++ , with L = Chol(S).(B.223) For the induced Lie group (S n ++ , ̃ ⊙), the left translation ̃ L P ̃ ⊙I −1 ̃ ⊙ under ̃ ⊙ is ̃ L P ̃ ⊙I −1 ̃ ⊙ = φ −1 ◦ L φ(P )⊙φ(I) −1 ⊙ ◦ φ,(B.224) = φ −1 ◦ L P 2θ ◦ φ, φ(P )⊙ φ(I) −1 ⊙ = P 2θ .(B.225) The associated differential at I is ̃ L P ̃ ⊙I −1 ̃ ⊙ ∗,I (A) = φ −1 ∗,φ(P ) ◦ (L P 2θ ) ∗,φ(I) ◦ φ ∗,I (A)(B.226) = 2θφ −1 ∗,φ(P ) ( ̄ LA ̄ L ⊤ ),(B.227) where ̄ L = Chol(P 2θ ). Then the SPD MLR under ̃g and ̃ ⊙ by left translation is p(y = k | S ∈S n ++ )∝ exp 2θg φ(P ) Log φ(P ) φ(S), ̄ LA ̄ L ⊤ .(B.228) Setting g = g BWM (we omit the scaling factor), we obtain the SPD MLR under 2θ-BWM: p(y = k | S ∈S n ++ )∝ exp 2θ· 1 4θ 2 g BWM φ(P ) Log BWM φ(P ) φ(S), ̄ LA ̄ L ⊤ (B.229) 337 B.6. Proper Velocity Neural Networks = exp 1 4θ ⟨(P 2θ S 2θ ) 1 2 + (S 2θ P 2θ ) 1 2 − 2P 2θ ,L P 2θ [ ̄ LA ̄ L ⊤ ]⟩ . (B.230) B.5.5 Proof of Thm. 119 Proof. During this proof, we use the ambient representation of tangent vectors. Given rotation matrices P,Q∈ SO(n) and a tangent vector H ∈ T Q SO(n), let c(t) be a curve on SO(n) satisfying c(0) = Q and c ′ (0) = H. The differential of the left translation L PQ −1 at Q is (L PQ −1 ) ∗,Q (H) = dPQ −1 c(t) dt t=0 = PQ −1 H = PQ ⊤ H,(B.231) which is precisely the vector transport T Q→P (H) in Boumal and Absil [28, Tab. 1]. B.5.6 Proof of Thm. 120 Proof. Setting Q = I in Eq. (4.28) and using Thm. 119 give ̃ A k = P k A k . Combining this identity with the logarithmic map and metric in Tab. 2.12, we obtain D Log P k R, ̃ A k E P k = P k log P ⊤ k R ,P k A k = log P ⊤ k R ,A k .(B.232) Substituting this expression into Eq. (4.26) yields the result. B.6 Proper Velocity Neural Networks B.6.1 Derivation of the Proper Velocity Metric The PV line element at x ∈ PV n K can be written in terms of the curvature parameter K < 0 as Q x (u) =∥u∥ 2 + Kβ 2 x ⟨x,u⟩ 2 , ∀u∈ T x PV n K ≃ R n ,(B.233) where β x = 1 √ 1−K∥x∥ 2 . This is equivalent to the expression in Ungar [200, Eq. (7.76)] after substituting s 2 = −1/K. Given u,v ∈ T x PV n K , the bilinear form g x (u,v) is 338 Appendix B. Proofs obtained by the polarization identity: g x (u,v) = 1 4 (Q x (u + v)− Q x (u− v)).(B.234) We first expand the two terms in the polarization identity: Q x (u + v) =∥u + v∥ 2 + Kβ 2 x ⟨x,u + v⟩ 2 =∥u∥ 2 + 2⟨u,v⟩ +∥v∥ 2 + Kβ 2 x ⟨x,u⟩ 2 + 2⟨x,u⟩⟨x,v⟩ +⟨x,v⟩ 2 , Q x (u− v) =∥u− v∥ 2 + Kβ 2 x ⟨x,u− v⟩ 2 =∥u∥ 2 − 2⟨u,v⟩ +∥v∥ 2 + Kβ 2 x ⟨x,u⟩ 2 − 2⟨x,u⟩⟨x,v⟩ +⟨x,v⟩ 2 . (B.235) Taking the difference yields Q x (u + v)− Q x (u− v) = 4⟨u,v⟩ + 4Kβ 2 x ⟨x,u⟩⟨x,v⟩. (B.236) Substituting this expression into the polarization identity, we obtain g x (u,v) = 1 4 (Q x (u + v)− Q x (u− v)) =⟨u,v⟩ + Kβ 2 x ⟨x,u⟩⟨x,v⟩, (B.237) which coincides with the expression of the PV metric in Eq. (5.1). B.6.2 Proof of Thm. 121 Proof. Differential of π PV n K →P n K . Consider the curve c : (−ε,ε)→ PV n K which satisfies c(0) = x and c ′ (0) = v. By definition of the differential, d x π PV n K →P n K (v) = d dt t=0 π PV n K →P n K (c(t)).(B.238) Using π PV n K →P n K (x) = β x 1+β x x with β x = 1 √ 1−K∥x∥ 2 , we write π PV n K →P n K (c(t)) = h(t)c(t), h(t) := β c(t) 1 + β c(t) .(B.239) Let r(t) =∥c(t)∥ 2 , so that β c(t) = (1− Kr(t)) −1/2 . Then r ′ (0) = 2⟨x,v⟩, β ′ c(0) = 1 2 (1− Kr(0)) −3/2 Kr ′ (0) = Kβ 3 x ⟨x,v⟩.(B.240) 339 B.6. Proper Velocity Neural Networks Differentiating h(t) at t = 0 gives h ′ (0) = β ′ c(0) (1 + β x ) 2 = K β 3 x (1 + β x ) 2 ⟨x,v⟩.(B.241) Finally, differentiating h(t)c(t) at t = 0 yields d x π PV n K →P n K (v) = h ′ (0)x + h(0)v = K β 3 x (1 + β x ) 2 ⟨x,v⟩x + β x 1 + β x v. (B.242) In particular, at x = 0 one has β 0 = 1 and ⟨x,v⟩ = 0. Thus, we have d 0 π PV n K →P n K (v) = β 0 1 + β 0 v = 1 2 v.(B.243) Differential of π P n K →PV n K . Consider the curve c : (−ε,ε) → P n K which satisfies c(0) = y and c ′ (0) = w. By definition of the differential, d y π P n K →PV n K (w) = d dt t=0 π P n K →PV n K (c(t)).(B.244) Using the explicit expression π P n K →PV n K (y) = 2γ 2 y y with γ y = 1 √ 1+K∥y∥ 2 , we obtain π P n K →PV n K (c(t)) = 2γ 2 c(t) c(t).(B.245) Let r(t) =∥c(t)∥ 2 so that γ 2 c(t) = (1 + Kr(t)) −1 . Then r ′ (0) = 2⟨y,w⟩, d dt t=0 γ 2 c(t) =− Kr ′ (0) (1 + Kr(0)) 2 =−2Kγ 4 y ⟨y,w⟩.(B.246) Differentiating 2γ 2 c(t) c(t) at t = 0 yields d y π P n K →PV n K (w) = 2 d dt t=0 γ 2 c(t) y + 2γ 2 y w =−4Kγ 4 y ⟨y,w⟩y + 2γ 2 y w. (B.247) In particular, at y = 0 we have γ 0 = 1 and ⟨y,w⟩ = 0. Thus, we have d 0 π P n K →PV n K (w) = 2w.(B.248) 340 Appendix B. Proofs B.6.3 Proof of Thm. 122 Proof. It suffices to show that for any x∈ PV n K and v,w ∈ T x PV n K , g P y d x π PV n K →P n K (v),d x π PV n K →P n K (w) = g PV x (v,w),(B.249) where y = π PV n K →P n K (x). We first recall the following equations from Eq. (5.1), Sec. 2.9.5, and Thm. 121: g PV x (v,w) =⟨v,w⟩ + Kβ 2 x ⟨x,v⟩⟨x,w⟩, ∀x∈ PV n K ,∀v,w ∈ T x PV n K , g P y (u,z) = λ K y 2 ⟨u,z⟩, ∀y ∈ P n K ,∀u,z ∈ T y P n K , π PV n K →P n K (x) = β x 1 + β x x, ∀x∈ PV n K , d x π PV n K →P n K (v) = β x 1 + β x v + K β 3 x (1 + β x ) 2 ⟨x,v⟩x, ∀x∈ PV n K ,∀v ∈ T x PV n K . (B.250) Let a = β x 1 + β x , b = K β 3 x (1 + β x ) 2 .(B.251) Then d x π PV n K →P n K (v) = av + b⟨x,v⟩x and d x π PV n K →P n K (w) = aw + b⟨x,w⟩x. Thus, g P y d x π PV n K →P n K (v),d x π PV n K →P n K (w) = λ K y 2 ⟨av + b⟨x,v⟩x,aw + b⟨x,w⟩x⟩ = λ K y 2 a 2 ⟨v,w⟩ + ab⟨x,w⟩⟨v,x⟩ + ab⟨x,v⟩⟨x,w⟩ + b 2 ⟨x,v⟩⟨x,w⟩⟨x,x⟩ = λ K y 2 a 2 ⟨v,w⟩ + 2ab + b 2 ∥x∥ 2 ⟨x,v⟩⟨x,w⟩ . (B.252) Using y = π PV n K →P n K (x) and the relation between λ K y , β x , and ∥x∥ from Sec. 2.9.5, we simplify the coefficients. First, K∥y∥ 2 = K β x 1 + β x x, β x 1 + β x x = K β 2 x (1 + β x ) 2 ∥x∥ 2 = β 2 x − 1 (1 + β x ) 2 , (B.253) 341 B.6. Proper Velocity Neural Networks where we use β 2 x = 1 1−K∥x∥ 2 . Hence 1 + K∥y∥ 2 = 1 + β 2 x − 1 (1 + β x ) 2 = 2β x 1 + β x ,(B.254) which implies λ K y = 2 1+K∥y∥ 2 = 1+β x β x . Therefore λ K y 2 a 2 = 1 + β x β x 2 β x 1 + β x 2 = 1.(B.255) Next, we compute 2ab = 2 β x 1 + β x K β 3 x (1 + β x ) 2 = 2Kβ 4 x (1 + β x ) 3 , b 2 ∥x∥ 2 = K 2 β 6 x (1 + β x ) 4 ∥x∥ 2 = K β 4 x (β 2 x − 1) (1 + β x ) 4 , (B.256) which brings us to 2ab + b 2 ∥x∥ 2 = K β 4 x (1 + β x ) 4 2(1 + β x ) + β 2 x − 1 = K β 4 x (β x + 1) 2 (1 + β x ) 4 = K β 4 x (1 + β x ) 2 . (B.257) Multiplying by λ K y 2 = 1+β x β x 2 yields λ K y 2 2ab + b 2 ∥x∥ 2 = 1 + β x β x 2 K β 4 x (1 + β x ) 2 = Kβ 2 x .(B.258) Substituting these identities into the expression for g P gives g P y d x π PV n K →P n K (v),d x π PV n K →P n K (w) =⟨v,w⟩ + Kβ 2 x ⟨x,v⟩⟨x,w⟩ = g PV x (v,w). (B.259) B.6.4 Proof of Thm. 123 Following the notation in the main theorem, we further denote: ̄x = π(x)∈ P n K , ̄y = π(y)∈ P n K , ̄v = d x π(v)∈ T ̄x P n K ,(B.260) 342 Appendix B. Proofs Recalling Eq. (5.8) and Thm. 121, we have the following: π(x) = β x 1 + β x x, ∀x∈ PV n K ,(B.261) π −1 ( ̄y) = 2γ 2 ̄y ̄y, ∀ ̄y ∈ P n K ,(B.262) d x π(v) = K β 3 x (1 + β x ) 2 ⟨x,v⟩x + β x 1 + β x v, ∀x∈ PV n K ,∀v ∈ T x PV n K ,(B.263) d ̄y π −1 (w) =−4Kγ 4 ̄y ⟨ ̄y,w⟩ ̄y + 2γ 2 ̄y w, ∀ ̄y ∈ P n K ,∀w ∈ T ̄y P n K .(B.264) Next, we derive the expressions for each PV operator. B.6.4.1 PV Exponential Map We recall from Tab. 2.14 that the Riemannian exponential on the Poincaré ball is Exp P ̄x ( ̄v) = ̄x⊕ M 1 √ −K tanh √ −Kλ K ̄x ∥ ̄v∥ 2 ̄v ∥ ̄v∥ .(B.265) By the Riemannian isometry and the gyrovector isomorphism of π, for any x ∈ PV n K and v ∈ T x PV n K we have Exp x (v) (1) = π −1 Exp P ̄x ( ̄v) (2) = x⊕ U π −1 1 √ −K tanh √ −Kλ K ̄x ∥ ̄v∥ 2 ̄v ∥ ̄v∥ , (B.266) The above equalities follow from the following facts. (1) Isometry. (2) Gyrovector isomorphism. Let u = 1 √ −K tanh √ −Kλ K ̄x ∥ ̄v∥ 2 ̄v ∥ ̄v∥ .(B.267) We have ∥u∥ = 1 √ −K tanh √ −Kλ K ̄x ∥ ̄v∥ 2 .(B.268) 343 B.6. Proper Velocity Neural Networks Let t = √ −Kλ K ̄x ∥ ̄v∥ 2 so that √ −K∥u∥ = tanh(t). Then π −1 (u) = 2 1 + K∥u∥ 2 u = 2 1 + K∥u∥ 2 1 √ −K tanh(t) ̄v ∥ ̄v∥ = 2 1− tanh 2 (t) 1 √ −K tanh(t) ̄v ∥ ̄v∥ = 2 √ −K tanh(t) 1− tanh 2 (t) ̄v ∥ ̄v∥ (1) = 2 √ −K tanh(t) cosh 2 (t) ̄v ∥ ̄v∥ = 2 √ −K sinh(t) cosh(t) cosh 2 (t) ̄v ∥ ̄v∥ = 2 √ −K sinh(t) cosh(t) ̄v ∥ ̄v∥ (2) = 1 √ −K (2 sinh(t) cosh(t)) ̄v ∥ ̄v∥ = 1 √ −K sinh (2t) ̄v ∥ ̄v∥ = 1 √ −K sinh √ −Kλ K ̄x ∥ ̄v∥ ̄v ∥ ̄v∥ (3) = 1 √ −K sinh √ −K(1 + β x ) β x ∥ ̄v∥ ̄v ∥ ̄v∥ . (B.269) The above equalities use: (1) 1− tanh 2 (t) = 1/ cosh 2 (t); (2) sinh(2t) = 2 sinh(t) cosh(t); (3) λ K ̄x = 1+β x β x . B.6.4.2 PV Logarithmic Map We recall from Tab. 2.14 that the Riemannian logarithm on the Poincaré ball is Log P ̄x ( ̄y) = 2 √ −Kλ K ̄x tanh −1 √ −K∥ ̄z∥ ∥ ̄z∥ ̄z, ̄z = (− ̄x)⊕ M ̄y,(B.270) 344 Appendix B. Proofs where λ K ̄x = 2 1+K∥ ̄x∥ 2 . We define z = (−x)⊕ U y, ̄z = π(z).(B.271) By the Riemannian isometry of π, we have Log x (y) = d ̄x π −1 Log P ̄x ( ̄y) = d ̄x π −1 2 √ −Kλ K ̄x tanh −1 √ −K∥ ̄z∥ ∥ ̄z∥ ̄z ! = α(x,y)d ̄x π −1 ( ̄z), (B.272) where α(x,y) = 2 √ −Kλ K ̄x tanh −1 √ −K∥ ̄z∥ ∥ ̄z∥ .(B.273) The differential of π −1 at ̄x = π(x) is d ̄x π −1 (h) =−4Kγ 4 ̄x ⟨ ̄x,h⟩ ̄x + 2γ 2 ̄x h, ∀h∈ T ̄x P n K ,(B.274) where γ ̄x = 1 √ 1+K∥ ̄x∥ 2 . Using ̄x = β x 1+β x x and the relation 1− K∥x∥ 2 = 1 β 2 x , we have K∥ ̄x∥ 2 = K β 2 x (1 + β x ) 2 ∥x∥ 2 = β 2 x (1 + β x ) 2 β 2 x − 1 β 2 x = β 2 x − 1 (1 + β x ) 2 , 1 + K∥ ̄x∥ 2 = 1 + β 2 x − 1 (1 + β x ) 2 = 2β x 1 + β x . (B.275) Hence γ 2 ̄x = 1 1 + K∥ ̄x∥ 2 = 1 + β x 2β x , γ 4 ̄x = 1 + β x 2β x 2 = (1 + β x ) 2 4β 2 x .(B.276) Substituting these into d ̄x π −1 (h) yields d ̄x π −1 (h) =−4K (1 + β x ) 2 4β 2 x ⟨ ̄x,h⟩ ̄x + 2 1 + β x 2β x h =−K (1 + β x ) 2 β 2 x ⟨ ̄x,h⟩ ̄x + 1 + β x β x h = 1 + β x β x h− K⟨x,h⟩x, (B.277) where the last equality uses that ̄x = β x 1+β x x. Applying Eq. (B.277) with h = ̄z and 345 B.6. Proper Velocity Neural Networks using that ̄z = π(z) is collinear with z = (−x)⊕ U y, we obtain Log x (y) = α(x,y)(d x π) −1 ( ̄z) = α(x,y) 1 + β x β x ̄z− K⟨x, ̄z⟩x . (B.278) Since ̄z = π(z) and π is given by Eq. (5.8), z and ̄z are collinear and ̄z = ρz, ρ = β z 1 + β z ,(B.279) which also implies ⟨x, ̄z⟩ = ρ⟨x,z⟩. Substituting these into Eq. (B.278) yields Log x (y) = α(x,y) 1 + β x β x ρz− Kρ⟨x,z⟩x = α(x,y) 1 + β x β x ρ | z σ(x,y) z + (−Kα(x,y)ρ) |z τ (x,y) ⟨x,z⟩x. (B.280) Using the definition of α(x,y) in Eq. (B.273) together with λ K ̄x = 1+β x β x and ρ = β z 1+β z , a straightforward simplification yields σ(x,y) = 2 √ −K tanh −1 √ −K∥ ̄z∥ ∥z∥ , τ (x,y) = 2β x 1 + β x √ −K tanh −1 √ −K∥ ̄z∥ ∥z∥ . (B.281) Thus, Log x (y) = σ(x,y)z + τ (x,y)⟨x,z⟩x.(B.282) B.6.4.3 PV Parallel Transport We recall from Tab. 2.14 that the parallel transport on the Poincaré ball is PT P ̄x→ ̄y (w) = λ K ̄x λ K ̄y gyr M [ ̄y,− ̄x](w), with w ∈ T ̄x P n K .(B.283) 346 Appendix B. Proofs We have PT x→y (v) (1) = d ̄y π −1 PT P ̄x→ ̄y (d x π(v)) = d ̄y π −1 λ K ̄x λ K ̄y gyr M [ ̄y,− ̄x] (d x π(v)) (2) = λ K ̄x λ K ̄y d ̄y π −1 (gyr M [ ̄y,− ̄x] (d x π(v))) (3) = λ K ̄x λ K ̄y 1 + β y β y gyr M [ ̄y,− ̄x] (d x π(v))− K⟨y, gyr M [ ̄y,− ̄x] (d x π(v))⟩y (4) = (1 + β x )β y (1 + β y )β x 1 + β y β y gyr M [ ̄y,− ̄x] (d x π(v))− K⟨y, gyr M [ ̄y,− ̄x] (d x π(v))⟩y = 1 + β x β x gyr M [ ̄y,− ̄x] (d x π(v))− K (1 + β x )β y (1 + β y )β x ⟨y, gyr M [ ̄y,− ̄x] (d x π(v))⟩y. (B.284) The above equalities use: (1) the isometry property of π; (2) linearity of d ̄y π −1 ; (3) Eq. (B.277). (4) Using the relation between λ K ̄x and β x in the proof of Thm. 122, λ K ̄x = 1 + β x β x λ K ̄y = 1 + β y β y .(B.285) B.6.4.4 PV Geodesic Distance We recall from Tab. 2.14 that the geodesic distance on the Poincaré ball P n K is d P (y 1 ,y 2 ) = 2 √ −K tanh −1 √ −K∥(−y 1 )⊕ M y 2 ∥ , y 1 ,y 2 ∈ P n K .(B.286) By isometry and isomorphism, the PV geodesic distance is d(x,y) = d P (π(x),π(y)) = 2 √ −K tanh −1 √ −K∥(−π(x))⊕ M π(y)∥ = 2 √ −K tanh −1 √ −K∥π(−x⊕ U y)∥ . (B.287) 347 B.6. Proper Velocity Neural Networks B.6.4.5 Special Cases at the Identity Exponential Map at the Identity. Exp 0 (v) (1) = 1 √ −K sinh √ −K(1 + β 0 ) β 0 ∥ ̄v∥ ̄v ∥ ̄v∥ (2) = 1 √ −K sinh √ −K∥v∥ v ∥v∥ . (B.288) The above comes from the following. (1) 0 is the gyro identity; (2) β 0 = 1 and d 0 π(v) = 1 2 v. Logarithmic Map at the Identity. As z =−0⊕ U y = y, we have Log 0 (y) = σ(0,y)z + τ (0,y)⟨0,z⟩0 = σ(0,y)y = 2 √ −K tanh −1 √ −K∥π(y)∥ ∥y∥ y. (B.289) From π(y) = β y 1+β y y and β y = 1 √ 1−K∥y∥ 2 , we obtain √ −K∥π(y)∥ = β y 1 + β y √ −K∥y∥.(B.290) Let t = √ −K∥y∥ and s = √ 1 + t 2 , so that β y = 1 √ 1−K∥y∥ 2 = 1 s . Define a = β y 1 + β y t = t s + 1 .(B.291) 348 Appendix B. Proofs Using the hyperbolic double-angle identity, we have tanh 2 tanh −1 (a) = 2a 1 + a 2 = 2t/(s + 1) 1 + t 2 /(s + 1) 2 = 2t(s + 1) (s + 1) 2 + t 2 = 2t(s + 1) s 2 + 2s + 1 + t 2 = 2t(s + 1) 2(1 + t 2 + s) = t(s + 1) 1 + t 2 + s = t(s + 1) s 2 + s = t s = t √ 1 + t 2 . (B.292) Denoting u = sinh −1 (t), we have cosh(u) = q 1 + sinh 2 (u) = √ 1 + t 2 .(B.293) Therefore, tanh sinh −1 (t) = tanh(u) = sinh(u) cosh(u) = t √ 1 + t 2 .(B.294) Since tanh is strictly increasing on R, this implies that 2 tanh −1 β y 1 + β y t = sinh −1 (t).(B.295) Substituting this identity back gives σ(0,y) = 1 √ −K sinh −1 √ −K∥y∥ ∥y∥ ,(B.296) and therefore Log 0 (y) = 1 √ −K sinh −1 √ −K∥y∥ y ∥y∥ .(B.297) Parallel Transport from the Identity. The gyration satisfies gyr M [0, ̄y] = gyr M [ ̄y,0] = id.(B.298) 349 B.6. Proper Velocity Neural Networks Substituting this into Eq. (B.284) gives PT 0→y (v) = 1 + β 0 β 0 d 0 π(v)− K (1 + β 0 )β y (1 + β y )β 0 ⟨y,d 0 π(v)⟩y = 2· 1 2 v− K 2β y 1 + β y · 1 2 ⟨y,v⟩y = v− K β y 1 + β y ⟨y,v⟩y. (B.299) Parallel Transport to the Identity. Taking y = 0 in Eq. (B.284) and using gyr M [0,− ̄x] = id yields PT x→0 (v) = 1 + β x β x d x π(v) = 1 + β x β x K β 3 x (1 + β x ) 2 ⟨x,v⟩x + β x 1 + β x v = v + K β 2 x 1 + β x ⟨x,v⟩x. (B.300) Distance from the Identity. This can be directly obtained by the gyro identity. d(0,y) = 2 √ −K tanh −1 √ −K∥π(y)∥ .(B.301) Using the same identity as above with t = √ −K∥y∥ yields d(0,y) = 1 √ −K sinh −1 √ −K∥y∥ .(B.302) B.6.5 Proof of Thm. 124 Proof. As shown in Thm. 88, the Möbius gyroaddition and gyromultiplication can be written by the Riemannian operators. Besides, the isometry π PV n K →P n K preserves the identity: π PV n K →P n K (0) = 0. By Nguyen and Yang [159, Lems. 2.1–2.2], one can directly obtain the results. B.6.6 Proof of Thm. 125 We first establish the PV hyperplane equivalence and then derive the distance formula. 350 Appendix B. Proofs B.6.6.1 Equivalent Characterization of the PV Hyperplane We first review a useful lemma from Chen et al. [55, Lem. J.1]. Lemma 196. We assume that the manifold M admits a gyrogroup defined by x⊕ y = Exp x (PT e→x (Log e (y))),∀x,y ∈M,(B.303) where e∈M is the origin of the manifold. Then, we have the following Log p (x),a p =⟨Log e (⊖p⊕ x), PT p→e (a)⟩ e , ∀x,p∈M and ∀a∈ T p M. (B.304) Now, we are ready to prove Thm. 125. Proof of PV hyperplane. Thm. 124 indicates that the assumption of Thm. 196 holds with M = PV n K , ⊕ =⊕ U and e = 0. Then, the PV hyperplane H a,p = n x∈ PV n K | Log p (x),a p = 0 o (B.305) can be rewritten as H a,p = x∈ PV n K |⟨Log 0 (−p⊕ U x), PT p→0 (a)⟩ 0 = 0 .(B.306) Using the explicit PV operators in Thm. 123 and the PV metric in Eq. (5.1), we have Log 0 (−p⊕ U x) = α(−p⊕ U x), for some scalar α≥ 0, PT p→0 (a) = βd p π(a), for some scalar β > 0, g 0 (u,v) =⟨u,v⟩. (B.307) As α = 0 is trivial, we only consider the case α > 0: ⟨Log 0 (−p⊕ U x), PT p→0 (a)⟩ 0 = 0 ⇐⇒ ⟨−p⊕ U x,d p π(a)⟩ = 0.(B.308) B.6.6.2 PV Point-to-Hyperplane Distance We first prove a lemma on the isometry and point-to-hyperplane distance, which will be used to derive the PV point-to-hyperplane distance. 351 B.6. Proper Velocity Neural Networks Lemma 197 (Isometry and point-to-hyperplane distance). Let (M,g) and ( ̄ M, ̄g) be Riemannian manifolds and let φ : M → ̄ M be a Riemannian isometry. For p∈M and a∈ T p M, define the hyperplane H a,p = x∈M| g p Log p (x),a = 0 .(B.309) Let ̄p = φ(p) and ̄a = d p φ(a) ∈ T ̄p ̄ M, and define the corresponding hyperplane on ̄ M by ̄ H ̄a, ̄p = ̄x∈ ̄ M| ̄g ̄p ̄ Log ̄p ( ̄x), ̄a = 0 .(B.310) Then φ maps H a,p onto ̄ H ̄a, ̄p , that is, φ (H a,p ) = ̄ H ̄a, ̄p .(B.311) Moreover, for every x∈M we have d M (x,H a,p ) = d ̄ M φ(x), ̄ H ̄a, ̄p ,(B.312) when the point-to-hyperplane distance exists. Here, d M and d ̄ M denote the Rie- mannian distances on M and ̄ M, respectively. Proof. Since φ is a Riemannian isometry, we have g p Log p (x),a = ̄g ̄p d p φ Log p (x) ,d p φ(a) = ̄g ̄p ̄ Log ̄p (φ(x)), ̄a .(B.313) Therefore, g p Log p (x),a = 0 ⇐⇒ ̄g ̄p ̄ Log ̄p (φ(x)), ̄a = 0,(B.314) which shows that x∈ H a,p if and only if φ(x)∈ ̄ H ̄a, ̄p , and hence φ (H a,p ) = ̄ H ̄a, ̄p . For the point-to-hyperplane distance, recall that for a subset S ⊂ M the distance from x to S is d M (x,S) = inf z∈S d M (x,z).(B.315) 352 Appendix B. Proofs For the point-to-hyperplane distance, we have d M (x,H a,p ) = inf z∈H a,p d M (x,z) = inf z∈H a,p d ̄ M (φ(x),φ(z)) = inf ̄z∈ ̄ H ̄a, ̄p d ̄ M (φ(x), ̄z) = d ̄ M φ(x), ̄ H ̄a, ̄p . (B.316) Next, we review the Poincaré hyperplane and point-to-hyperplane distance [76, Sec. 3.1]. Poincaré Point-to-Hyperplane Distance. For a point p ∈ P n K and a normal vector a ∈ T p P n K , the Poincaré point-to-hyperplane distance is given by Ganea et al. [76, Thm. 5]: H P a,p = n x∈ P n K | Log P p (x),a p = 0 o =x∈ P n K |⟨−p⊕ M x,a⟩ = 0, (B.317) d P (y,H P a,p ) = 1 √ −K sinh −1 2 √ −K|⟨−p⊕ M y,a⟩| 1 + K∥−p⊕ M y∥ 2 ∥a∥ ! .(B.318) Proof of the PV point-to-hyperplane distance. Let ̄p = π(p)∈ P n K , ̄a = d p π(a)∈ T ̄p P n K , ̄y = π(y)∈ P n K .(B.319) By Thm. 197, the point-to-hyperplane distances satisfy d PV (y,H a,p ) = d P ̄y, ̄ H ̄a, ̄p .(B.320) Applying the Poincaré distance formula in Eq. (B.318) with p = ̄p, a = ̄a, and y = ̄y gives d P ̄y, ̄ H ̄a, ̄p = 1 √ −K sinh −1 2 √ −K|⟨− ̄p⊕ M ̄y, ̄a⟩| 1 + K∥− ̄p⊕ M ̄y∥ 2 ∥ ̄a∥ ! .(B.321) The gyrovector isomorphism π implies − ̄p⊕ M ̄y = π(−p⊕ U y).(B.322) 353 B.6. Proper Velocity Neural Networks Denote z =−p⊕ U y. From Eq. (5.8), we have the explicit expression π(z) = β z 1 + β z z,(B.323) with β z > 0. Since β z = 1 √ 1−K∥z∥ 2 , we obtain 1 + K∥π(z)∥ 2 = 1 + K β z 1 + β z 2 ∥z∥ 2 = 1 + Kβ 2 z ∥z∥ 2 (1 + β z ) 2 = 1 + β 2 z (1− β −2 z ) (1 + β z ) 2 = 1 + β 2 z − 1 (1 + β z ) 2 = 1 + β z − 1 1 + β z = 2β z 1 + β z (B.324) The above yields 2 √ −K|⟨π(z), ̄a⟩| 1 + K∥π(z)∥ 2 ∥ ̄a∥ = √ −K|⟨z, ̄a⟩| ∥ ̄a∥ .(B.325) Therefore, d (y,H a,p ) = 1 √ −K sinh −1 √ −K|⟨−p⊕ U y,d p π(a)⟩| ∥d p π(a)∥ .(B.326) B.6.7 Proof of Thm. 126 Proof of PV MLR. For clarity, we fix a class index k and omit k in the notation when- ever possible. We denote π = π PV n K →P n K as in Thm. 125. Step 1: From Hyperplane Distance to a Signed Score. The PV MLR in 354 Appendix B. Proofs Eq. (5.24) associated with parameters (p,a) for x∈ PV n K is v k (x) = sign (⟨−p k ⊕ U x,d p k π(a k )⟩)∥a k ∥ p k d (x,H a k ,p k ) (1) = ∥a k ∥ p k √ −K sign (⟨−p k ⊕ U x,d p k π(a k )⟩) sinh −1 √ −K|⟨−p k ⊕ U x,d p k π(a k )⟩| ∥d p k π(a k )∥ (2) = ∥a k ∥ p k √ −K sinh −1 √ −K⟨−p k ⊕ U x,d p k π(a k )⟩ ∥d p k π(a k )∥ . (B.327) The above comes from the following. (1) Thm. 125; (2) sinh −1 is odd and strictly increasing. Step 2: Trivialization and Reduction to a Single Direction. We adopt the unidirectional parameterization in Sec. 5.2.4.1: p k = Exp 0 (r k [z k ]), a k = PT 0→p k (z k ),[z k ] = z k ∥z k ∥ ,(B.328) with z k ∈ T 0 PV n K ∼ = R n and r k ∈ R. As parallel transport is an isometry, we have ∥a k ∥ p k =∥z k ∥ 0 =∥z k ∥.(B.329) Moreover, p k and z k are collinear, because Exp 0 in Thm. 123 preserves directions at the origin. Using the explicit expression of PT 0→y at the origin in Thm. 123, we see that PT 0→p k maps z k to a linear combination of z k and p k . Therefore, a k is also collinear with z k . The differential d p k π in Thm. 121 has the form d p k π(v) = α k v + β k ⟨p k ,v⟩p k , α k > 0, β k ∈ R,(B.330) so d p k π maps any vector in spanz k into the same one-dimensional subspace. Conse- quently, there exists a scalar λ k > 0 such that d p k π(a k ) = λ k z k .(B.331) The sign of λ k can be absorbed into z k by redefining z k ← −z k if necessary. Without loss of generality we may assume λ k > 0. Putting Eq. (B.328), Eq. (B.329), Eq. (B.331) 355 B.6. Proper Velocity Neural Networks and ∥d p k π(a k )∥ = λ k ∥z k ∥ into Eq. (B.327) yields v k (x) = ∥z k ∥ √ −K sinh −1 √ −K ∥z k ∥ ⟨−p k ⊕ U x,z k ⟩ .(B.332) Step 3: Eliminating the Gyroaddition. The remaining task is to expand the gyro-additive term in Eq. (B.332). From Sec. 5.2.2, PV gyroaddition is given by u⊕ U v = u + v + 1− β v β v − K β u 1 + β u ⟨u,v⟩ u, β w = 1 p 1− K∥w∥ 2 . (B.333) Setting u =−p k and v = x yields −p k ⊕ U x =−p k + x + 1− β x β x − K β p k 1 + β p k ⟨−p k ,x⟩ (−p k ). (B.334) Taking the inner product with z k gives ⟨−p k ⊕ U x,z k ⟩ =⟨−p k ,z k ⟩ +⟨x,z k ⟩ + 1− β x β x − K β p k 1 + β p k ⟨−p k ,x⟩ ⟨−p k ,z k ⟩ =⟨x,z k ⟩ + 1 + 1− β x β x − K β p k 1 + β p k ⟨−p k ,x⟩ ⟨−p k ,z k ⟩. (B.335) Next, we rewrite the above expression using the unidirectional parameterization of p k . From Eq. (B.328) and the explicit PV exponential at the origin in Thm. 123, we have p k = Exp 0 (r k [z k ]) = 1 √ −K sinh √ −Kr k z k ∥z k ∥ .(B.336) Thus, ⟨−p k ,z k ⟩ =− 1 √ −K sinh √ −Kr k ∥z k ∥.(B.337) Moreover, since p k and z k share the same direction, any x admits the decomposition x = x ∥ + x ⊥ , x ∥ = ⟨x,z k ⟩ ∥z k ∥ 2 z k , ⟨x ⊥ ,z k ⟩ = 0,(B.338) which implies ⟨−p k ,x⟩ = −p k ,x ∥ = ⟨x,z k ⟩ ∥z k ∥ 2 ⟨−p k ,z k ⟩ =− 1 √ −K sinh √ −Kr k ⟨x,z k ⟩ ∥z k ∥ . (B.339) 356 Appendix B. Proofs The beta factor at p k is β p k = 1 p 1− K∥p k ∥ 2 = sech √ −Kr k ,(B.340) where we used ∥p k ∥ 2 =− 1 K sinh 2 √ −Kr k and the identity 1 + sinh 2 (t) = cosh 2 (t). Using ⟨−p k ,z k ⟩, ⟨−p k ,x⟩, and β p k , we have ⟨−p k ⊕ U x,z k ⟩ =⟨x,z k ⟩ + 1 + 1− β x β x − K β p k 1 + β p k ⟨−p k ,x⟩ ⟨−p k ,z k ⟩ =⟨x,z k ⟩ + 1 β x − K β p k 1 + β p k ⟨−p k ,x⟩ ⟨−p k ,z k ⟩ =⟨x,z k ⟩ + 1 β x − sinh √ −Kr k √ −K ∥z k ∥ ! − K β p k 1 + β p k − sinh √ −Kr k √ −K ⟨x,z k ⟩ ∥z k ∥ ! − sinh √ −Kr k √ −K ∥z k ∥ ! =⟨x,z k ⟩− sinh √ −Kr k √ −K ∥z k ∥ β x + β p k sinh 2 √ −Kr k 1 + β p k ⟨x,z k ⟩ = 1 + β p k sinh 2 √ −Kr k 1 + β p k ! ⟨x,z k ⟩− sinh √ −Kr k √ −K ∥z k ∥ β x . (B.341) Since β p k = sech √ −Kr k and 1 + sinh 2 √ −Kr k = cosh 2 √ −Kr k = 1/β 2 p k , we have 1 + β p k sinh 2 √ −Kr k 1 + β p k = 1 + β p k + β p k sinh 2 √ −Kr k 1 + β p k = 1 + β p k cosh 2 √ −Kr k 1 + β p k = 1 + 1/β p k 1 + β p k = 1 β p k = cosh √ −Kr k , (B.342) which implies ⟨−p k ⊕ U x,z k ⟩ = cosh √ −Kr k ⟨x,z k ⟩− sinh √ −Kr k √ −K ∥z k ∥ β x . (B.343) 357 B.6. Proper Velocity Neural Networks Recalling that β x = 1/ p 1− K∥x∥ 2 , we obtain ⟨−p k ⊕ U x,z k ⟩ = cosh √ −Kr k ⟨x,z k ⟩− sinh √ −Kr k √ −K ∥z k ∥ p 1− K∥x∥ 2 . (B.344) Substituting Eq. (B.344) into Eq. (B.332), we arrive at v k (x) = ∥z k ∥ √ −K sinh −1 cosh √ −Kr k √ −K ∥z k ∥ ⟨x,z k ⟩− sinh √ −Kr k p 1− K∥x∥ 2 . (B.345) Proof of PV MLR limits. By Taylor expansions, we have cosh √ −Kr k = 1− Kr 2 k 2 +O(K 2 ), sinh √ −Kr k = √ −Kr k +O (−K) 3/2 , p 1− K∥x∥ 2 = 1− K∥x∥ 2 2 +O(K 2 ). (B.346) The argument of sinh −1 (·) in Eq. (B.345) can be simplified as sinh −1 cosh √ −Kr k √ −K ∥z k ∥ ⟨x,z k ⟩− sinh √ −Kr k p 1− K∥x∥ 2 = sinh −1 1− Kr 2 k 2 +O(K 2 ) √ −K ∥z k ∥ ⟨x,z k ⟩ − √ −Kr k +O (−K) 3/2 1− K∥x∥ 2 2 +O(K 2 ) = sinh −1 √ −K ⟨x,z k ⟩ ∥z k ∥ − r k +O (−K) 3/2 = √ −K ⟨x,z k ⟩ ∥z k ∥ − r k +O (−K) 3/2 . (B.347) Substituting this into Eq. (B.345) gives v k (x) = ∥z k ∥ √ −K √ −K ⟨x,z k ⟩ ∥z k ∥ − r k +O (−K) 3/2 =∥z k ∥ ⟨x,z k ⟩ ∥z k ∥ − r k +O(−K) =⟨x,z k ⟩− r k ∥z k ∥ +O(−K), K→0 − −→⟨x,z k ⟩− r k ∥z k ∥. (B.348) 358 Appendix B. Proofs B.6.8 Proof of Thm. 127 Proof of PV FC layer. Specializing Thm. 125 to p = 0 and a = e k and using that −0⊕ U y = y gives the LHS sign (⟨d 0 π(e k ),−0⊕ U y⟩) d (y,H e k ,0 ) = 1 √ −K sinh −1 √ −Ky k ,(B.349) with y k =⟨y,e k ⟩. Then, we obtain y k = 1 √ −K sinh √ −Kv k (x) , k = 1,...,m.(B.350) Proof of PV FC limits. By Thm. 126, as K → 0 − we have v k (x)→⟨x,z k ⟩ + b k , b k =−r k ∥z k ∥.(B.351) For K < 0 and v k (x)̸= 0, we can rewrite y k as y k = v k (x) sinh √ −Kv k (x) √ −Kv k (x) ,(B.352) and we define the fraction to be 1 when v k (x) = 0. Since √ −K → 0 and v k (x) converges to a finite limit, we have √ −Kv k (x)→ 0. Using the standard limit lim u→0 sinh(u)/u = 1, it follows that sinh √ −Kv k (x) √ −Kv k (x) → 1 as K → 0 − .(B.353) Combining the above limits yields lim K→0 − y k = lim K→0 − v k (x) =⟨x,z k ⟩ + b k .(B.354) B.6.9 Proof of Thm. 128 Proof. The result is first established in the Poincaré ball model in Thms. 83 and 90. Since π = π PV n K →P n K : PV n K → P n K is a Riemannian isometry, for x,y ∈ PV n K , v ∈ T x PV n K , 359 B.6. Proper Velocity Neural Networks and t∈ R, it intertwines the key geometric operators used in the proof: π Exp PV x (v) = Exp P π(x) (d x π(v)),(B.355) d x π Log PV x (y) = Log P π(x) (π(y)),(B.356) π (x⊕ U y) = π(x)⊕ M π(y),(B.357) π (t⊗ U x) = t⊙ M π(x).(B.358) Together with the preservation of geodesic distances and Fréchet means under the Riemannian isometry π, these identities imply that both homogeneity identities are preserved under π. Therefore the same theorem holds for the PV model by the isometry π. B.6.10 Proof of Thm. 129 Proof. We first recall the isometries between the Poincaré ball and the hyperboloid [182, Sec. 2.1]: π H n K →P n K (x) = x s 1 + p |K|x t ,(B.359) π P n K →H n K (y) = 1 p |K| 1− K∥y∥ 2 1 + K∥y∥ 2 2y 1 + K∥y∥ 2 .(B.360) Hence, the following are Riemannian isometries: π H n K →PV n K = π P n K →PV n K ◦ π H n K →P n K , π PV n K →H n K = π P n K →H n K ◦ π PV n K →P n K .(B.361) It remains to derive the explicit formulas. For x = [x t ,x ⊤ s ] ⊤ ∈ H n K , we first map to the Poincaré ball: y = π H n K →P n K (x) = x s 1 + p |K|x t .(B.362) Applying π P n K →PV n K from Eq. (5.8) yields π H n K →PV n K (x) = π P n K →PV n K (y) = 2γ 2 y y, γ y = 1 q 1 + K∥y∥ 2 .(B.363) 360 Appendix B. Proofs Using y = x s /(1 + p |K|x t ), we compute 1 + K∥y∥ 2 = 1 + K ∥x s ∥ 2 1 + p |K|x t 2 = 1 + p |K|x t 2 + K∥x s ∥ 2 1 + p |K|x t 2 .(B.364) Since x∈ H n K satisfies ⟨x,x⟩ L = 1/K, we have ⟨x,x⟩ L =−x 2 t +∥x s ∥ 2 = 1 K ⇒ ∥x s ∥ 2 = x 2 t + 1 K .(B.365) Substituting this into the numerator gives 1 + p |K|x t 2 + K∥x s ∥ 2 = 1 + p |K|x t 2 + K x 2 t + 1 K = 1 + p |K|x t 2 + Kx 2 t + 1 = 1 + 2 p |K|x t +|K|x 2 t + Kx 2 t + 1 = 2 1 + p |K|x t . (B.366) Therefore, 1 + K∥y∥ 2 = 2 1 + p |K|x t 1 + p |K|x t 2 = 2 1 + p |K|x t ,(B.367) and hence γ 2 y = 1 1 + K∥y∥ 2 = 1 + p |K|x t 2 .(B.368) Finally, π H n K →PV n K (x) = 2γ 2 y y = 2· 1 + p |K|x t 2 · x s 1 + p |K|x t = x s .(B.369) For π PV n K →H n K , take x∈ PV n K and map to the Poincaré ball by Eq. (5.8): y = π PV n K →P n K (x) = β x 1 + β x x, β x = 1 q 1− K∥x∥ 2 .(B.370) 361 B.6. Proper Velocity Neural Networks Applying π P n K →H n K , we obtain π PV n K →H n K (x) = π P n K →H n K (y) = 1 p |K| 1− K∥y∥ 2 1 + K∥y∥ 2 2y 1 + K∥y∥ 2 .(B.371) We now simplify the spatial and temporal components separately. We write y = β x 1 + β x x, ∥y∥ 2 = β x 1 + β x 2 ∥x∥ 2 .(B.372) Using β 2 x = 1/(1− K∥x∥ 2 ), we obtain K∥y∥ 2 = K∥x∥ 2 β 2 x (1 + β x ) 2 = β 2 x − 1 (1 + β x ) 2 = β x − 1 (1 + β x ) ⇒ 1 + K∥y∥ 2 = 2β x 1 + β x ,1− K∥y∥ 2 = 2 1 + β x . (B.373) The spatial component of π PV n K →H n K (x) is 2y 1 + K∥y∥ 2 = 2 β x 1+β x x 2β x 1+β x = x,(B.374) and the temporal component is 1 p |K| 1− K∥y∥ 2 1 + K∥y∥ 2 = 1 p |K| · 2 1+β x 2β x 1+β x = 1 p |K|β x = q 1− K∥x∥ 2 p |K| = q ∥x∥ 2 − 1 K . (B.375) Thus, π PV n K →H n K (x) = q ∥x∥ 2 − 1 K x .(B.376) 362 Appendix B. Proofs B.7 Hyperbolic Busemann Neural Networks B.7.1 Proof of Thm. 130 Proof. Denoting κ 2 =−K with κ > 0, we rewrite the Busemann functions in Eqs. (5.40) and (5.41) as (Poincaré) B v (x) = 1 κ log ∥v− κx∥ 2 1− κ 2 ∥x∥ 2 ! ,(B.377) (Lorentz) B v (x) = 1 κ log (κx t − κ⟨x s ,v⟩).(B.378) For the Poincaré case, ∥v− κx∥ 2 1− κ 2 ∥x∥ 2 = 1− 2κ⟨v,x⟩ + κ 2 ∥x∥ 2 1− κ 2 ∥x∥ 2 = 1 + −2κ⟨v,x⟩ + 2κ 2 ∥x∥ 2 1− κ 2 ∥x∥ 2 .(B.379) Using log (1 + u) = u + O (u 2 ) as u→ 0 and 1− κ 2 ∥x∥ 2 −1 = 1 + O (κ 2 ), we obtain B v (x) = 1 κ −2κ⟨v,x⟩ + 2κ 2 ∥x∥ 2 1 + O κ 2 + O κ 2 =−2⟨v,x⟩ + 2κ∥x∥ 2 + O (κ) =−2⟨v,x⟩ + O (κ). (B.380) Therefore, B v (x) K→0 − −→−2⟨v,x⟩. For the Lorentz case, the hyperboloid constraint −x 2 t +∥x s ∥ 2 = −κ −2 and x t > 0 yield κx t = q 1 + κ 2 ∥x s ∥ 2 = 1 + 1 2 κ 2 ∥x s ∥ 2 + O κ 4 .(B.381) Set z =−κ⟨x s ,v⟩ + 1 2 κ 2 ∥x s ∥ 2 + O (κ 4 ). Then log (κx t − κ⟨x s ,v⟩) = log (1 + z) = z + O z 2 =−κ⟨x s ,v⟩ + O κ 2 ,(B.382) which implies B v (x) =−⟨x s ,v⟩ + O (κ) and therefore B v (x) K→0 − −→−⟨x s ,v⟩. Substituting the above two limits into u k (x) = −α k B v k (x) + b k gives the stated Euclidean limits of the logits. 363 B.7. Hyperbolic Busemann Neural Networks B.7.2 Proof of Thm. 132 We first review two useful characterizations of Busemann functions in Hadamard spaces. The class B below is given by Bridson and Haefliger [29, p. 271]. The equivalence follows from Bridson and Haefliger [29, Prop. I.8.22], whose formulation in terms of Busemann functions uses the identification of horofunctions with Busemann functions up to additive constants in Bridson and Haefliger [29, Cor. I.8.20]. Definition 198. Let (X, d) be a Hadamard space, and letB be the set of functions h :X → R on (X, d) satisfying: (1) h is convex; (2) 1-Lipschitz: |h(x)− h(y)|≤ d(x,y) for all x,y ∈X; (3) for any x 0 ∈ X and r > 0, the function h attains its minimum on the sphere S r (x 0 ) at a unique point y with h(y) = h(x 0 )− r. Proposition 199. For a function h :X → R, the following conditions are equiva- lent: (1) h is a Busemann function; (2) h∈B; (3) h is convex, and for every t ∈ R, the set h −1 (−∞,t] is nonempty; moreover, for each x∈X, the curve c x : [0,∞)→X defined by t7→ π h −1 (−∞,h(x)−t] (x) is a geodesic ray. Now, we are ready to prove Thm. 132. Proof. As τ 2 = τ 1 is trivial, we only consider τ 2 ̸= τ 1 . We assume τ 2 > τ 1 , and discuss the other direction last. Step 1: Symmetry. By definition, d H γ τ 1 ,H γ τ 2 =inf x∈H γ τ 1 ,y∈H γ τ 2 d(x,y) =inf y∈H γ τ 2 ,x∈H γ τ 1 d(y,x) = d H γ τ 2 ,H γ τ 1 .(B.383) Step 2: Lower Bound. For any x∈ H γ τ 2 and y ∈ H γ τ 1 , the 1-Lipschitz property of B γ gives |B γ (x)− B γ (y)|≤ d(x,y).(B.384) 364 Appendix B. Proofs With B γ (x) = τ 2 and B γ (y) = τ 1 this yields τ 2 − τ 1 ≤ d(x,y) ∀y ∈ H γ τ 1 .(B.385) Taking infimum in Eq. (B.385) over y ∈ H γ τ 1 gives τ 2 − τ 1 ≤ d x,H γ τ 1 .(B.386) Step 3: Upper Bound. For any x ∈ H γ τ 2 , by property (3) in Thm. 199, the projection map c x (t) = π B γ ≤B γ (x)−t (x) = π HB γ τ 2 −t (x), t∈ [0,∞),(B.387) is a unit-speed geodesic ray: d(x,c x (t)) = t. Let t = τ 2 −τ 1 > 0 and z = c x (t)∈ HB γ τ 1 . If B γ (z) < τ 1 , then z lies in the interior of the horoball HB γ τ 1 . We can move slightly from z toward x along the geodesic segment xz to obtain a point z ε with d(x,z ε ) < d(x,z), contradicting the minimality of the projection. Hence, the projected point z = c x (t) indeed lies on H γ τ 1 : B γ (z) = τ 1 . Then, we have the following: d(x,H γ τ 1 )≤ d(x,HB γ τ 1 ) = d(x,z) = d(x,c x (t)) = t = τ 2 − τ 1 .(B.388) Step 4: Sandwich Closure. Combining Eqs. (B.386) and (B.388) gives d(x,H γ τ 1 ) = τ 2 − τ 1 .(B.389) Combining Eq. (B.388) and Eq. (B.386), d x,H γ τ 1 = τ 2 − τ 1 ,for every x∈ H γ τ 2 .(B.390) The right-hand side does not depend on x, hence d H γ τ 2 ,H γ τ 1 = τ 2 − τ 1 .(B.391) Step 5: Opposite Direction. If τ 1 > τ 2 , geodesic completeness of the Hadamard space allows the geodesic ray c x from Step 3 to extend to a complete geodesic ec x : R→ X. For any x ∈ H γ τ 2 , let s = τ 1 − τ 2 and z = ec x (−s). Since B γ (c x (t)) = B γ (x)− t for t ≥ 0, convexity and the 1-Lipschitz property of B γ give B γ (z) = B γ (x) + s = τ 1 . 365 B.7. Hyperbolic Busemann Neural Networks Thus, z ∈ H γ τ 1 and d(x,z) = s. Combining this upper bound with the 1-Lipschitz lower bound from Step 2 gives d x,H γ τ 1 = τ 1 − τ 2 ,for every x∈ H γ τ 2 .(B.392) Therefore, in both directions, d (x,H γ τ ) =|B γ (x)− τ|, ∀x∈X,(B.393) and the symmetry of the distance between horospheres gives d H γ τ 2 ,H γ τ 1 =|τ 2 − τ 1 |.(B.394) B.7.3 Proof of Thm. 135 Proof. The hyperplane and point-to-hyperplane distance are presented in Tabs. A.16 and A.17. The origin of the Poincaré ball model is e = 0∈ P m K . The specific ones w.r.t. the origin [181, Def. 1 and Eq. (56)] are H e k ,0 =y ∈ P m K |⟨e k ,y⟩ = 0,(B.395) ̄ d (y,H e k ,0 ) = 1 √ −K sinh −1 2 √ −Ky k 1 + K∥y∥ 2 .(B.396) Equating ̄ d (y,H e k ,0 ) with u k (x) from Eq. (5.59) gives sinh −1 2 √ −Ky k 1 + K∥y∥ 2 = √ −Ku k (x), ∀k ∈1,...,m.(B.397) Note that Eq. (B.397) takes the same form as Shimizu et al. [181, Eq. (56)], except that their responses are different. The proof below is inspired by their derivation. Applying sinh(·) on both sides of Eq. (B.397) yields 2 √ −Ky k 1 + K∥y∥ 2 = sinh √ −Ku k (x) .(B.398) 366 Appendix B. Proofs Define ω k := 1 √ −K sinh √ −Ku k (x) and ω = [ω k ] m k=1 . Then, we have 2y k = 1 + K∥y∥ 2 ω k , ∀k,(B.399) which is equivalent to the vector identity 2y = 1 + K∥y∥ 2 ω.(B.400) Hence, y is collinear with ω. Write y = λω with λ ≥ 0. Substituting into Eq. (B.400) and taking norms gives a quadratic in λ: K∥ω∥ 2 λ 2 − 2λ + 1 = 0.(B.401) Solving and selecting the branch that satisfies y → 0 as ω → 0 yields λ = 1− q 1− K∥ω∥ 2 K∥ω∥ 2 = 1 1 + q 1− K∥ω∥ 2 .(B.402) Therefore, y = ω 1 + q 1− K∥ω∥ 2 , ω k = sinh √ −Ku k (x) √ −K ,(B.403) which proves the claim. One can check that y ∈ P m K . B.7.4 Proof of Thm. 136 Proof. Recalling Tab. A.16, a Lorentz hyperplane is H w,p =x∈ L m K |⟨w,x⟩ L = 0, with p∈ L m K , w ∈ T p L m K .(B.404) The canonical origin is 0∈ L m K . The tangent space at the origin is T 0 L m K =[0,v ⊤ ] ⊤ | v ∈ R m ,(B.405) where each tangent vector has a zero time component. Therefore, the coordinate hy- perplane through the origin and orthogonal to the k-th axis is H ̄e k ,e =y ∈ L m K |⟨ ̄e k ,y⟩ L = 0 =y = (y t ,y s )∈ L m K | (y s ) k = 0, (B.406) 367 B.7. Hyperbolic Busemann Neural Networks where ̄e k = [0,e ⊤ k ] ⊤ ∈ T 0 L m K . From Tab. A.17, the associated signed point-to-hyperplane distance is ̄ d (y,H ̄e k ,e ) = sign (⟨ ̄e k ,y⟩ L ) d (y,H ̄e k ,e ) = 1 √ −K sinh −1 √ −K(y s ) k . (B.407) Equating ̄ d (y,H ̄e k ,e ) with u k (x) from Eq. (5.59) gives sinh −1 √ −K(y s ) k = √ −Ku k (x),1≤ k ≤ m.(B.408) Applying sinh(·) to both sides of Eq. (B.408) yields (y s ) k = 1 √ −K sinh √ −Ku k (x) ,1≤ k ≤ m.(B.409) Stacking the coordinates gives y s = 1 √ −K sinh √ −Ku(x) , u(x) = (u 1 (x),...,u m (x)) ⊤ .(B.410) Since y ∈ L m K , the hyperboloid constraint⟨y,y⟩ L = 1/K implies−y 2 t +∥y s ∥ 2 = 1/K. Taking the positive time component yields y t = r 1 −K +∥y s ∥ 2 .(B.411) Combining the expressions for y t and y s proves the claim. B.7.5 Proof of Thm. 137 Proof. Set K =−κ 2 with κ > 0. Poincaré Case. Recall that y = ω 1 + q 1 + κ 2 ∥ω∥ 2 , ω k = sinh (κu k (x)) κ .(B.412) For any bounded scalar z, sinh (κz) = κz + κ 3 z 3 3! + O κ 5 ⇒ sinh (κz) κ = z + κ 2 z 3 6 + O κ 4 .(B.413) 368 Appendix B. Proofs Applying this to z = u k (x) yields ω k = u k (x) + κ 2 6 u k (x) 3 + O κ 4 = u k (x) + O κ 2 .(B.414) Thus, ω = u(x) + O (κ 2 ) and ∥ω∥ 2 =∥u(x)∥ 2 + O (κ 2 ). For the denominator, q 1 + κ 2 ∥ω∥ 2 = 1 + 1 2 κ 2 ∥ω∥ 2 + O κ 4 ,(B.415) which gives 1 + q 1 + κ 2 ∥ω∥ 2 = 2 + 1 2 κ 2 ∥ω∥ 2 + O κ 4 .(B.416) Taking the reciprocal produces 1 1 + q 1 + κ 2 ∥ω∥ 2 = 1 2 + O κ 2 .(B.417) Multiplying with ω = u(x) + O (κ 2 ) gives y = 1 2 u(x) + O κ 2 .(B.418) By Thm. 130, u k (x)→ 2α k ⟨v k ,x⟩ + b k , hence y k → α k ⟨v k ,x⟩ + 1 2 b k .(B.419) Lorentz Case. Recall that y s = 1 κ sinh (κu(x)), y t = r 1 κ 2 +∥y s ∥ 2 .(B.420) Using the same expansion as above, y s,k = 1 κ κu k (x) + κ 3 3! u k (x) 3 + O κ 5 = u k (x) + O κ 2 .(B.421) Therefore, y s = u(x) + O (κ 2 ) and ∥y s ∥ 2 =∥u(x)∥ 2 + O (κ 2 ). 369 B.8. Full-Rank Correlation Networks Factor out κ −1 and expand the square root: y t = 1 κ q 1 + κ 2 ∥y s ∥ 2 = 1 κ 1 + 1 2 κ 2 ∥y s ∥ 2 + O κ 4 = 1 κ + κ 2 ∥y s ∥ 2 + O κ 3 = 1 κ + O (κ)→∞. (B.422) By Thm. 130, u k (x)→ α k ⟨v k ,x s ⟩ + b k . Using the spatial expansion yields (y s ) k → α k ⟨v k ,x s ⟩ + b k .(B.423) This completes the proof. B.8 Full-Rank Correlation Networks B.8.1 Proof of Thm. 138 We first prove a lemma for MLRs on general isometric manifolds, of which this theorem is a specific case. Notably, the result and proof can be readily extended to the case where R m is endowed with an arbitrary inner product. Lemma 200 (Isometric Riemannian MLRs). Given m-dimensional Riemannian manifolds f M,g f M and M,g M with a Riemannian isometry φ : f M → M, their origins are E ∈ f M and φ(E) ∈ M. The Riemannian MLR over f M for the input X ∈ f M of each class k = 1,· ,C can be calculated by the one over M: v f M k (X;Z k ,γ k ) = v M k (φ(X);φ ∗,E (Z k ),γ k ),(B.424) with γ k ∈ R, Z k ∈ T E f M ∼ = R m , and φ ∗,E : T E f M→ T φ(E) M as the differential map. Here, v f M k and v M k are the specific realizations of the Riemannian MLR reviewed in Sec. 4.3.2.1 over f M and M, respectively. Proof. We omit the subscript k in A k and P k for simplicity. We denote e Γ, g Log, ⟨·,·⟩ P , ∥·∥ P , e d(X, e H A,P ), e H A,P as the parallel transport along the geodesic, Riemannian log- arithm, Riemannian metric, the induced norm, margin distance and hyperplane over 370 Appendix B. Proofs f M, while the counterparts over M are denoted as Γ, Log, ⟨·,·⟩ φ(P ) , ∥·∥ φ(P ) , d, and H, respectively. From the isometry, we have ∥A∥ P =∥φ ∗,P (A)∥ φ(P ) ,(B.425) D g Log P (X),A E P = Log φ(P ) (φ(X)),φ ∗,P (A) φ(P ) .(B.426) The above equations imply φ e H A,P = H φ ∗,P (A),φ(P ) .(B.427) Denoting H = H φ ∗,P (A),φ(P ) , we have the following for the margin distance e d(X, e H A,P ) = inf Q∈ e H A,P e d(X,Q) (1) =inf Q∈ e H A,P d(φ(X),φ(Q)) (2) = inf R∈H d(φ(X),R) (3) = d(φ(X),H). (B.428) The above comes from the following. (1) Isometry. (2) Eq. (B.427). (3) Definition of margin distance. Combining the above, we have v f M (X;P,A) = sign(⟨A, g Log P (X)⟩ P )∥A∥ P e d(X, e H A,P ) = sign Log φ(P ) (φ(X)),φ ∗,P (A) φ(P ) ∥φ ∗,P (A)∥ φ(P ) d(φ(X),H φ ∗,P (A),φ(P ) ) = v M (φ(X);φ(P ),φ ∗,P (A)). (B.429) Finally, let us further consider trivialization. By isometry, we have the following: A = e Γ E→P (Z) = φ −1 ∗,P PT φ(E)→φ(P ) (φ ∗,E (Z)) , (B.430) 371 B.8. Full-Rank Correlation Networks P = g Exp E (γ[Z]) = φ −1 Exp φ(E) (γ[φ ∗,E (Z)]) . (B.431) Then, we have φ ∗,P (A) = PT φ(E)→φ(P ) (φ ∗,E (Z)),(B.432) φ(P ) = Exp φ(E) (γ[φ ∗,E (Z)]).(B.433) Putting the above two equations into Eq. (B.429), we have v f M (X;Z,γ) = v f M (X;P,A) = v M (φ(X);φ(P ),φ ∗,P (A)) = v M φ(X); Exp φ(E) (γ[φ ∗,E (Z)]), PT φ(E)→φ(P ) (φ ∗,E (Z)) = v M k (φ(X);φ ∗,E (Z k ),γ k ). (B.434) Thm. 138 is a special case of Thm. 200 and can be readily proven accordingly. Proof of Thm. 138. MLR. In Euclidean space R m , simple computations show that the Riemannian MLR reviewed in Sec. 4.3.2.1 becomes Eq. (4.1), where the latter is equal to ⟨a k ,x− p k ⟩. Based on Thm. 200, we have v k (X;Z k ,γ k ) = v R m k (φ(X);φ ∗,E (Z k ),γ k ), =⟨φ(X)− γ k [φ ∗,E (Z k )],φ ∗,E (Z k )⟩ =⟨φ(X),φ ∗,E (Z k )⟩− γ k ∥φ ∗,E (Z k )∥, (B.435) Margin Hyperplane. In Euclidean space R m , the Riemannian margin hyperplane becomes the Euclidean one, which is parameterized by ⟨a k ,x− p k ⟩ = 0. Together with Eq. (B.435), the results can be easily obtained. B.8.2 Proof of Thm. 139 Proof. First, we have the following: Θ(I) = I,(B.436) Chol(I) = I,(B.437) log ∗,I (V ) = V, ∀V ∈ Hol(n),(B.438) 372 Appendix B. Proofs log ∗,I (V ) = V, ∀V ∈ LT 0 (n),(B.439) D ⋆ (I) = I.(B.440) Putting the above into the differential formulas collected in Sec. 2.9.2, one can directly get the result w.r.t. ECM, LECM, and OLM. For LSM, based on Sec. 2.9.2, we have Log ⋆ ∗,I (V ) = log ∗,Σ ∆V ∆ + 1 2 V 0 Σ + ΣV 0 (1) = V + 1 2 V 0 + V 0 (2) = V − diag (V 1). (B.441) The above comes from the following. (1) Σ = ∆ = I (2) V 0 =−2 diag (I n + Σ) −1 ∆V ∆1 =− diag (V 1) (B.442) B.8.3 Proof of Thm. 142 Let d n = n(n−1) 2 and d m = m(m−1) 2 be the manifold dimensions of Cor + (n) and Cor + (m), respectively. We have the following general results. Lemma 201. Let Cor + (n),g n be isometric to R d n by the diffeomorphism φ n : Cor + (n)→ R d n ,(B.443) and let Cor + (m),g m be isometric to R d m by the diffeomorphism φ m : Cor + (m)→ R d m .(B.444) The diffeomorphism satisfies I n = φ −1 n (0 d n ), I m = φ −1 m (0 d m ).(B.445) 373 B.8. Full-Rank Correlation Networks The correlation FC layer F : Cor + (n)→ Cor + (m) for the input X ∈ Cor + (n) is Y = φ −1 m d m X i=1 v i (X)e i ! ,(B.446) where e i d m i=1 is the canonical orthonormal basis over R d m with e i = (δ ik ) d m k=1 for each i. Here, v i (X) d m i=1 is given by Thm. 138: v i (X) =⟨φ n (X), (φ n ) ∗,I n (Z i )⟩− γ i ∥(φ n ) ∗,I n (Z i )∥,(B.447) with Z i ∈ T I n Cor + (n) ∼ = R d n and γ i ∈ R as the FC parameters. Proof of Thm. 201. LetO k = (φ m ) −1 ∗,I m (e k ) d m k=1 . ThenO k d m k=1 is an orthonormal basis over T I m Cor + (m). The LHS of Eq. (5.69) is sign Log I m (Y ),O k I m d(Y,H O k ,I m ) (1) = sign (φ m ) −1 ∗,I m φ m (Y ),O k I m d(Y,H O k ,I m ) (2) = sign (⟨φ m (Y ),e k ⟩) d(Y,H O k ,I m ) (3) = sign (⟨φ m (Y ),e k ⟩) d(φ m (Y ),H e k ,0 d m ) = (φ m (Y )) k , (B.448) where (1)–(2) come from the isometry, and (3) comes from Eq. (B.428). The RHS of Eq. (5.69) can be implied by Thm. 138. Thm. 201 can be naturally extended to the cases where the inner products of R d n and R d m are not canonical. Lemma 202. Following all the notation in Thm. 201, we further assume that the inner products Q n (·,·) over R d n and Q m (·,·) over R d m are not necessarily canonical. In addition, f : (R d m ,Q m (·,·))→ (R d m ,⟨·,·⟩) is a linear isometry to the canonical inner product. Then, we have Y = φ −1 m ◦ f −1 d m X i=1 v i (X)e i ! ,(B.449) v i (X) = Q n (φ n (X), (φ n ) ∗,I n (Z i ))− γ i ∥(φ n ) ∗,I n (Z i )∥ Q n ,(B.450) 374 Appendix B. Proofs Figure B.1: Illustration of the Euclidean spaces LT 0 (m), Hol(m) and Row 0 (m), where ⋆ can be obtained by symmetry. where ∥·∥ Q n is the norm induced by Q n . Proof of Thm. 202. First, we denote ψ m = f ◦ φ m : Cor + (m),g m → (R d m ,⟨·,·⟩).(B.451) Note that the differential of any linear map between vector spaces is itself. The rest of the proof is identical to that of Thm. 201. Now, we present the proof of Thm. 142. Proof of Thm. 142. As ECM, LECM, OLM, and LSM are pullback metrics from Eu- clidean spaces, we resort to Thm. 201 and its extension Thm. 202. Denoting the zero matrix as 0, we have the following: φ EC (I n ) = log◦Θ(I n ) = 0∈ LT 0 (n),(B.452) Log ◦ (I n ) = 0∈ Hol(n),(B.453) Log ⋆ (I n ) = 0∈ Row 0 (n).(B.454) Therefore, the identity matrix is indeed the origin defined in Thm. 201. Recalling Thm. 201, the prototype space is the vector space with the standard vector inner product. Obviously, LT 0 (m), Hol(m), and Row 0 (m) are linearly isomorphic to R m(m−1) /2 . As shown in Fig. B.1, each L ∈ LT 0 (m) can be identified with a vector of its lower triangular part. Besides, LT 0 (m) with the canonical matrix inner product is identified with R m(m−1) 2 with standard vector inner product. Therefore, the basis over LT 0 (m) corresponding to the canonical orthonormal basis over R m(m−1) /2 is (LT 0 (m),⟨·,·⟩) : U LT 0 (m) ij = E ij ,1≤ j < i≤ m,(B.455) 375 B.8. Full-Rank Correlation Networks where E ij ∈ R m×m is the standard basis matrix, with the (k,l)-th element defined as (E ij ) kl = 1 if k = i and l = j, 0 otherwise. (B.456) Without loss of generality, we identify (LT 0 (m),⟨·,·⟩) with (R m(m−1) 2 ,⟨·,·⟩), and refer to E ij 1≤j<i≤m as the canonical orthonormal basis. However, E ij is neither a canonical orthonormal basis nor even orthonormal for Hol(m) and Row 0 (m) under the standard matrix inner product. According to Thm. 202, we only need to find the linear isometry that maps these two spaces into (LT 0 (m),⟨·,·⟩). By Fig. B.1, we have the following linear isometries to pull back these two inner products to the standard ones over LT 0 (m): f Hol(m)→LT 0 (m) :(Hol(m),⟨·,·⟩)→ (LT 0 (m),⟨·,·⟩), Hol(m)∋ H 7−→ √ 2⌊H⌋∈ LT 0 (m), f Row 0 (m)→LT 0 (m) :(Row 0 (m),⟨·,·⟩)→ (LT 0 (m),⟨·,·⟩), Row 0 (m)∋ R7−→ √ 6⌊ e R⌋ + √ 3D( e R)∈ LT 0 (m), (B.457) where e R ∈ S m−1 is the leading principal submatrix of order m− 1 of R. The bases f −1 Hol(m)→LT 0 (m) (E ij ) and f −1 Row 0 (m)→LT 0 (m) (E ij ) are as follows: (Hol(m),⟨·,·⟩) : U Hol(m) ij = E ij + E ji √ 2 ,1≤ j < i≤ m(B.458) (Row 0 (m),⟨·,·⟩) : U Row 0 (m) ij = E i −E im −E mi √ 3 ,if 1≤ i < m E ij +E ji −E mi −E im −E mj −E jm √ 6 , if 1≤ j < i < m (B.459) Putting the required diffeomorphisms and v g ij in Thm. 140 into Thm. 202 for ECM, LECM, OLM, and LSM, the corresponding FC layers can be readily obtained. B.8.4 Proof of Thm. 143 Proof. First, we review the isometries between the open hemisphere and hyperboloid [195, Eqs. (4.1)–(4.2)], and the one between Poincaré ball and hyperboloid [182, Sec. 2.1]: ψ HS n →H n : (x 1 ,...,x n+1 ) ⊤ ∈ HS n 7−→ 1 x n+1 (1,x 1 ,...,x n ) ⊤ ∈ H n ,(B.460) 376 Appendix B. Proofs ψ H n →HS n : y t ,y ⊤ s ⊤ ∈ H n 7−→ 1 y t y ⊤ s , 1 ⊤ ∈ HS n ,(B.461) ψ H n →P n : x t ,x ⊤ s ⊤ ∈ H n 7−→ x s 1 + x t ∈ P n ,(B.462) ψ P n →H n : y ∈ P n 7−→ 1 +∥y∥ 2 1−∥y∥ 2 , 2y ⊤ 1−∥y∥ 2 ! ⊤ = 1 1−∥y∥ 2 1 +∥y∥ 2 2y ! ∈ H n . (B.463) For any (x ⊤ ,x n+1 ) ⊤ ∈ HS n and y ∈ P n , we have ψ HS n →P n x x n+1 !! = ψ H n →P n ◦ ψ HS n →H n x x n+1 !! = ψ H n →P n 1 x n+1 1 x !! = x x n+1 1 1 + 1 x n+1 = x 1 + x n+1 (B.464) ψ P n →HS n (y) = ψ H n →HS n ◦ ψ P n →H n (y) = ψ H n →HS n 1 1−∥y∥ 2 1 +∥y∥ 2 2y !! = 1 1 +∥y∥ 2 2y 1−∥y∥ 2 ! (B.465) B.8.5 Proof of Thm. 145 Proof. We denote D =D(H). By Sec. 2.9.2, we have dY = dD + dH dD =− diag H 0 −1 D exp ∗,Y (dH) 1 . (B.466) Following Ionescu et al. [111], we denote the inner product⟨·,·⟩ as· :· for simplicity. By the invariance of differential and properties of trace [111, Eqs. 67–72], we have the 377 B.8. Full-Rank Correlation Networks following: ∂l ∂Y : dY = ∂l ∂Y : dD + ∂l ∂Y : dH = ∂l ∂Y :− diag (H 0 ) −1 D exp ∗,Y (dH) 1 + ∂l ∂Y : dH (1) = tr −Dv ∂l ∂Y ⊤ (H 0 ) −1 D exp ∗,Y (dH) 1 ! + ∂l ∂Y : dH (2) = tr − " 1Dv ∂l ∂Y ⊤ (H 0 ) −1 # D exp ∗,Y (dH) ! + ∂l ∂Y : dH =−(H 0 ) −1 Dv ∂l ∂Y 1 ⊤ : D exp ∗,Y (dH) + ∂l ∂Y : dH =−D (H 0 ) −1 Dv ∂l ∂Y 1 ⊤ : exp ∗,Y (dH) + ∂l ∂Y : dH (3) = ∂l ∂Y − exp ∗,Y D (H 0 ) −1 Dv ∂l ∂Y 1 ⊤ : dH (4) = off ∂l ∂Y − exp ∗,Y D (H 0 ) −1 Dv ∂l ∂Y 1 ⊤ : dH (B.467) The above comes from the following. (1) A : diag(b) = Dv(A) : b, ∀A∈ R n×n ,b∈ R n ,(B.468) a : b = a ⊤ b = tr(a ⊤ b), ∀a,b∈ R n .(B.469) (2) Cyclic property of the trace for matrices A,B, and C of compatible dimensions: tr(ABC) = tr(CAB). (3) For any A ∈ S n and S ∈ S n , write Y = U ∆U ⊤ and let L = L exp be the Loewner matrix in Eq. (2.91). By the Daleckii–Krein formula in Eq. (2.90) and the properties of trace, we have A : exp ∗,Y (S) = A : U L⊛ U ⊤ SU U ⊤ = U L⊛ U ⊤ AU U ⊤ : S = exp ∗,Y (A) : S. (B.470) (4) H has zero diagonal elements. 378 Appendix B. Proofs The invariance of the first-order differential gives ∂l ∂Y : dY = ∂l ∂H : dH.(B.471) By the last equation in Eq. (B.467), we can obtain ∂l ∂H . B.8.6 Proof of Thm. 146 Proof. Denoting by f : Cor + (n) → Row + 1 (n) the map f (C) = D ⋆ (C)CD ⋆ (C) = Σ, we have Log ⋆ ∗,C = log ∗,Σ ◦f ∗,C .(B.472) Combining with the differential of Log ⋆ shown in Sec. 2.9.2, we have the following differential equation: dΣ = ∆dC∆− V 0 Σ + ΣV 0 ,(B.473) with V 0 = diag (I n + Σ) −1 ∆dC∆1 . Similarly to Thm. 145, we have the following: ∂l ∂Σ : dΣ = ∂l ∂Σ : ∆dC∆− V 0 Σ + ΣV 0 = ∆ ∂l ∂Σ ∆ : dC− ∂l ∂Σ : V 0 Σ + ΣV 0 = ∆ ∂l ∂Σ ∆ : dC− ∂l ∂Σ Σ + Σ ∂l ∂Σ : diag (I n + Σ) −1 ∆dC∆1 = ∆ ∂l ∂Σ ∆ : dC− Dv ∂l ∂Σ Σ + Σ ∂l ∂Σ : (I n + Σ) −1 ∆dC∆1 = ∆ ∂l ∂Σ ∆ : dC− tr ev ⊤ (I n + Σ) −1 ∆dC∆1 = ∆ ∂l ∂Σ ∆ : dC− tr ∆1ev ⊤ (I n + Σ) −1 ∆dC = ∆ ∂l ∂Σ ∆ : dC− ∆ (I n + Σ) −1 ev1 ⊤ ∆ : dC = ∆ ∂l ∂Σ ∆− ∆ (I n + Σ) −1 ev1 ⊤ ∆ : dC = ∆ ∂l ∂Σ − (I n + Σ) −1 ev1 ⊤ ∆ : dC. (B.474) By imposing symmetrization, we can obtain the results. 379 B.9. Adaptive Log-Euclidean Metrics B.8.7 Proof of Thm. 144 As β-splitting is the inverse of β-concatenation [181], we only need to show the case w.r.t. β-concatenation. Besides, it suffices to prove the 2D case, which is shown in the following lemma. Lemma 203. Given x ij ∈ P n j with i ∈ 1,...,N i and j ∈ 1,...,N j , applying the β-concatenation sequentially 2 times in the order j → i is equivalent to a single β-concatenation along all indices simultaneously. Proof. Denoting d = P N j j=1 n j and v ij = Log 0 (x ij ), we have the following Exp 0 concat N i i=1 β N i ×d β −1 d concat N j j=1 β d β −1 n j v ij = Exp 0 concat i=N i ,j=N j i=1,j=1 β N i ×d β −1 d β d β −1 n j v ij = Exp 0 concat i=N i ,j=N j i=1,j=1 β N i ×d β −1 n j v ij . (B.475) The last line implies the claim. A special case of the above lemma is where all n j are identical. Corollary 204. Given x ij ∈ P n with i∈1,...,N i and j ∈1,...,N j , applying the β-concatenation sequentially 2 times in the order j → i is equivalent to a single β-concatenation along all indices simultaneously. Thm. 144 can be obtained by Thms. 203 and 204. B.9 Adaptive Log-Euclidean Metrics B.9.1 Proof of Thm. 147 Proof of Thm. 147. Let us first deal with (α,β)-LEM. Substituting the differential of the matrix logarithm into Thm. 33 directly yields the result. Now, let us focus on LCM. Denote LCM, the standard Euclidean metric, and the metric on the Cholesky manifold [137] by g LC , g E , and g C , respectively. By Tab. 2.5, S n ++ ,g LC is isometric to L n ++ ,g C , with the Cholesky decomposition Chol as an isometry. This is exactly how Lin [137] derived LCM. So, the key point lies in the Cholesky metric g C . Let us reveal why it is defined in this way. In fact, g C is derived 380 Appendix B. Proofs from g E by φ ln . Simple computations show that (φ ln ) ∗,L (V ) =⌊V⌋ + D(L) −1 D(V ),(B.476) where V ∈ T L L n ++ . By Eq. (B.476), Tab. 2.5 can be rewritten as g C L (X,Y ) = g E (φ ln ) ∗,L (X), (φ ln ) ∗,L (Y ) .(B.477) Therefore, φ ln : L n ++ → LT n is an isometry. By transitivity, ψ LC : S n ++ → LT n is also an isometry. B.9.2 Proof of Thm. 148 Proof of Thm. 148. As R n(n+1)/2 ∼ = LT n ∼ = S n , LCM is therefore a pullback metric from the standard Euclidean space S n . Second, any two Euclidean spaces of the same finite dimension are naturally isometric; hence, (α,β)-LEM is also a pullback metric from the standard Euclidean space S n . B.9.3 Proof of Thm. 149 Proof of Thm. 149. By the definitions in Eqs. (6.2) to (6.5), the Hilbert-space and isomorphism claims follow directly. It remains to establish the geometric claims. As every Euclidean space is an abelian Lie group, S n ++ ,⊙ φ is an abelian Lie group. The geodesic distance in Eq. (6.6) also follows immediately because φ is a Riemannian isometry. We only need to prove Eqs. (6.7) to (6.9). Note that in the Euclidean space S n , for any x,y ∈S n and tangent vector v ∈ T x S n ∼ = S n , we have the following: Exp x v = x + v,(B.478) Log x y = y− x,(B.479) PT x→y v = v.(B.480) By the isometry of φ, we can readily obtain Eqs. (6.7) to (6.9). B.9.4 Proof of Thm. 150 Proof of Thm. 150. Obviously, log −1 α is the inverse of log α . What follows is to verify the smoothness of log α and its inverse. 381 B.9. Adaptive Log-Euclidean Metrics According to Magnus and Neudecker [144, Thm. 8.9], the map producing an eigen- value or an eigenvector from a real symmetric matrix is C ∞ . Recalling log α and its inverse map log −1 α , it is obvious that they comprise arithmetic calculations or composi- tions of smooth maps. Therefore, log α is a diffeomorphism with inverse log −1 α . B.9.5 Proof of Thm. 152 Proof of Thm. 152. This is a direct result of Thm. 149. B.9.6 Proof of Thm. 154 Proof of Thm. 154. The differentials of log −1 α and log α can be derived similarly. In the following, we only present the process of deriving the differential of log α . First, let us recall the differentials of eigenvalues and eigenvectors. Magnus and Neudecker [144, Thm. 8.9] offers their Euclidean differentials, which are the exact for- mulations for differentials under the canonical base on SPD manifolds. Thus, we can readily obtain the differentials of eigenvalues and eigenvectors as follows: σ ∗,S (V ) = u ⊤ V u,(B.481) u ∗,S (V ) = (σI n − S) + V u,(B.482) where Su = σu, u ⊤ u = 1, and (·) + is the Moore–Penrose inverse. By the RHS of Eq. (6.30), the differential map of log α is (log α ) ∗,S (V ) = U ∗,S (V ) log α (Σ)U ⊤ + U (log α ) ∗,Σ (Σ ∗,S (V ))U ⊤ + U log α (Σ)U ⊤ ∗,S (V ) = Q + Q ⊤ + U (log α ) ∗,Σ (Σ ∗,S (V ))U ⊤ , (B.483) where Q = U ∗,S (V ) log α (Σ)U ⊤ . For the differential of diagonal logarithm, it is (log α ) ∗,Σ (Σ ∗,S (V )) = A 1 Σ Σ ∗,S (V ),(B.484) where A is defined in Eq. (6.31). Denote the eigenvectors and eigenvalues of S = U ΣU ⊤ by U = (u 1 ,...,u n ) and Σ = diag(σ 1 ,...,σ n ). By Eqs. (B.481) to (B.484), the differential of log α can be 382 Appendix B. Proofs obtained. B.9.7 Proof of Thm. 155 Proof of Thm. 155. Following the notation in the proposition, we prove the result as follows. By abuse of notation, in the following, we omit the wide tildee. Now, we proceed to deal with the differential of log −1 α . We rewrite the formula of log −1 α as log −1 α (X)(B.485) = U diag (a σ 1 1 ,· ,a σ n n )U ⊤ ,(B.486) = U diag e log(a 1 )σ 1 ,· ,e log(a n )σ n U ⊤ ,(B.487) = U diag ∞ X k=0 (log(a 1 )σ 1 ) k k! ,· , ∞ X k=0 (log(a n )σ n ) k k! ! U ⊤ ,(B.488) = U ∞ X k=0 (BΣ) k k! ! U ⊤ ,(B.489) = ∞ X k=0 (PX) k k! (B.490) where P = UBU ⊤ , U is obtained from the eigendecomposition X = U ΣU ⊤ , and B = diag (log(a 1 ),· , log(a n )) is diagonal. By the properties of normed vector algebras [197, Prop. 15.14], we can obtain the last equation. Then, we can compute the differential of log −1 α by curves. Given a curve c on S n starting at X with initial velocity V ∈ T X S n , write c(t) = U (t)Σ(t)U (t) ⊤ and define P (t) = U (t)BU (t) ⊤ , so that P (0) = P. We have log −1 α ∗,X (V ) = d dt t=0 log −1 α (c(t)) = d dt t=0 ∞ X k=0 (P (t)c(t)) k k! . (B.491) Term-by-term differentiation gives log −1 α ∗,X (V ) = ∞ X k=1 1 k! ( k−1 X l=0 (PX) k−l−1 d dt t=0 (P (t)c(t))(PX) l ). (B.492) 383 B.9. Adaptive Log-Euclidean Metrics By the chain rule, we have d dt t=0 (P (t)c(t)) = P ′ (0)X + PV.(B.493) P ′ (0) is obtained by P ′ (0) = d dt t=0 U (t)BU (t) ⊤ = U ′ (0)BU ⊤ + UBU ′ (0) ⊤ = D U BU ⊤ + UBD ⊤ U , (B.494) where D U is derived from the differential of eigenvectors, D U = ( (σ 1 I n − X) + V u 1 · (σ n I n − X) + V u n ).(B.495) Substituting Eqs. (B.493) to (B.495) into Eq. (B.492) yields the differential of log −1 α . B.9.8 Proof of Thm. 156 Proof of Thm. 156. Obviously, the metric space S n ++ ,d ALE is isometric to the space S n endowed with the standard Euclidean distance. Therefore, the weighted Fréchet mean of S i in S n ++ corresponds to the weighted Fréchet mean of associated points log α (S i ) in S n . The weighted Fréchet means in Euclidean spaces are clearly the familiar weighted means. B.9.9 Proof of Thm. 157 Proof of Thm. 157. As log α is a Riemannian isometry andS n is bi-invariant, ALEM is therefore bi-invariant. 384 Appendix B. Proofs B.9.10 Proof of Thm. 158 Proof of Thm. 158. Following the notation in this proposition, we proceed as follows. The right-hand side can be rewritten as (FM(S β 1 ,·S β m )) = log −1 α m X i=1 1 m β log α (S i ) ! = log −1 α β m X i=1 1 m log α (S i ) ! = " log −1 α m X i=1 1 m log α (S i ) !# β = (FM(S 1 ,·S m )) β . (B.496) B.9.11 Proof of Thm. 159 Proof of Thm. 159. Recalling Eq. (6.24), Properties U1 and U2 obviously hold. When the SPD matrices A i i≤n commute, we have FM(A i ) = Y i A i ! 1 n .(B.497) With Eq. (B.497), Properties V1–V4 can be easily proved. B.9.12 Proof of Thm. 160 Proof of Thm. 160. Obviously, for a given SPD matrix S, log α (RSR ⊤ ) = R log α (S)R ⊤ ,(B.498) log α (s 2 S) = U log α (s 2 I n ) + log α (Σ) U ⊤ ,(B.499) where S = U ΣU ⊤ is the eigendecomposition. These identities yield the result. B.9.13 Proof of Thm. 161 Proof of Thm. 161. The three equations can be directly obtained. 385 B.9. Adaptive Log-Euclidean Metrics B.9.14 Proof of Thm. 163 Proof of Thm. 163. The input gradient follows from the Daleckii–Krein formula re- viewed in Eqs. (2.90) and (2.91). Now, let us focus on the gradient with respect to A. Differentiating both sides of Eq. (6.31) gives dX = (∗) + U (dA⊛ log(Σ))U ⊤ ,(B.500) where (∗) denotes other terms involving dU and d Σ. According to the invariance of the first-order differential form, we have ∇ X L : dX(B.501) =∇ S L : dS +∇ X L : U (dA⊛ log(Σ))U ⊤ (B.502) =∇ S L : dS + [U ⊤ (∇ X L)U ]⊛ log(Σ) : dA,(B.503) where A : B = tr(A ⊤ B) is the Euclidean Frobenius inner product. From the second term on the RHS of Eq. (B.503), we can obtain the gradient with respect to A. B.9.15 Proof of Thm. 164 Proof of Thm. 164. The derivation follows the same logic as Thm. 163. We only need to show the derivation of Eq. (6.35). Similarly to Thm. 163, we have the following: dX = (∗) + U dA⊛ diag a Σ 11 1 ,· ,a Σ n n −Σ A 2 U ⊤ ,(B.504) ∇ X L : dX(B.505) =∇ S L : dS + [U ⊤ (∇ X L)U ]⊛ diag a Σ 11 1 ,· ,a Σ n n −Σ A 2 : dA.(B.506) B.9.16 Proof of Thm. 167 Proof of Thm. 167. Following Nguyen [157], Nguyen and Yang [159], we first define gyrostructures under ALEM: P ⊕ ALE Q = Exp P PT I n →P Log I n (Q) ,(B.507) 386 Appendix B. Proofs gyr[P,Q]R = (⊖(P ⊕ ALE Q))⊕ ALE (P ⊕ ALE (Q⊕ ALE R)),(B.508) t⊙ ALE P = Exp I n t Log I n (P ) ,(B.509) ⊖P =−1⊙ ALE P = Exp I n − Log I n (P ) ,(B.510) ⟨P,Q⟩ gyr = Log I n (P ), Log I n (Q) I n ,(B.511) ∥P∥ gyr = q ⟨P,P⟩ gyr ,(B.512) d gyr (P,Q) = ⊖P ⊕ ALE Q gyr ,(B.513) where P,Q,R∈S n ++ , and I n is the identity matrix. The above operations are called gy- roaddition, gyroautomorphism, scalar gyromultiplication, gyroinverse, gyroinner prod- uct, gyronorm, and gyrodistance. Simple computations show that Eq. (B.507) and Eq. (B.509) are exactly ⊕ ALE and ⊙ ALE in Thm. 152. As indicated by Thm. 152, S n ++ ,⊕ ALE ,⊙ ALE forms a gyrovector space [158, Def. 1]. In the following proof, we use ⊙ ALE and ⊕ ALE . The gyro MLR [159] under ALEM is defined as p(y = k | S) ∝ exp sign(⟨ ̃ A k , Log P k (S)⟩ P k )∥ ̃ A k ∥ P k ̄ d(S,H ̃ A k ,P k ) , (B.514) where P k ∈ S n ++ and ̃ A k ∈ T P k S n ++ . ̄ d(S,H ̃ A k ,P k ) is the margin distance to the SPD hyperplane H ̃ A k ,P k , which is defined as ̄ d(S,H ̃ A k ,P k ) = sin(∠SP k Q ∗ )d gyr (S,P k ),(B.515) Q ∗ =argmax Q∈H ̃ A k ,P k \P k (cos(∠SP k Q)),(B.516) cos(∠SP k Q) = ⊖P k ⊕ ALE Q,⊖P k ⊕ ALE S gyr ∥⊖P k ⊕ ALE Q∥ gyr ∥⊖P k ⊕ ALE S∥ gyr ,(B.517) H ̃ A k ,P k =S ∈S n ++ |⟨Log P k S, ̃ A k ⟩ P k = 0.(B.518) Eqs. (B.515), (B.517) and (B.518) are called the SPD pseudo-gyrodistance, SPD gyro- cosine, and SPD gyrohyperplane. For simplicity, we further omit the subscript k in P k and ̃ A k . Eq. (B.518) can be 387 B.9. Adaptive Log-Euclidean Metrics simplified: ⟨Log P S, ̃ A⟩ P (1) = D log −1 α ∗,log α (P ) (log α (S)− log α (P )), ̃ A E P (2) = D (log α ) ∗,P ◦ log −1 α ∗,log α (P ) (log α (S)− log α (P )), (log α ) ∗,P ( ̃ A) E = D log α (S)− log α (P ), (log α ) ∗,P ( ̃ A) E . (B.519) The above derivation comes from the following. (1) Eq. (6.18). (2) The definition of ALEM. Similarly, a simple computation shows that Eq. (B.517) can also be simplified as ⟨− log α (P ) + log α (Q),− log α (P ) + log α (S)⟩ ∥− log α (P ) + log α (Q)∥ F ∥− log α (P ) + log α (S)∥ F .(B.520) Together with Eqs. (B.519) and (B.520), Eq. (B.515) is equivalent to the distance to the hyperplane in the Euclidean space. Therefore, Eq. (B.515) has a closed-form solution: ̄ d(S,H ̃ A,P ) = log α (S)− log α (P ), ̄ A ̄ A F = log α (S)− log α (P ), ̄ A ̃ A P , (B.521) where ̄ A = (log α ) ∗,P ( ̃ A). Substituting Eq. (B.521) into Eq. (B.514) yields the claimed result. B.9.17 Proof of Thm. 165 Proof of Thm. 165. To derive the ALEM-specific positive-scalar update, consider the general RSGD update reviewed in Sec. 2.7. For a minimization parameter w on an n-dimensional smooth connected Riemannian manifold M, we have w (t+1) = Exp w (t) (−γ (t) π w (t) (∇ w (t) L)),(B.522) 388 Appendix B. Proofs where Exp w (·) : T w M→M is the Riemannian exponential map, which maps a tangent vector at w back into the manifold M, and π w (·) : R n → T w M is the projection operator, projecting an ambient Euclidean vector into the tangent space at w. In the case of the SPD manifold, for all S ∈ S n ++ , X ∈ R n×n , and V ∈ S n , the exponential map and projection operator are formulated as follows: π S (X) = S X + X ⊤ 2 S,(B.523) Exp S (V ) = S 1/2 exp(S −1/2 V S −1/2 )S 1/2 ,(B.524) where exp(·) is the matrix exponential. For more details about Eq. (B.523) and Eq. (B.524), see Yger [225] and Amari [3]. Substituting Eq. (B.523) and Eq. (B.524) into Eq. (B.522) immediately yields Eq. (6.36). B.9.18 Proof of Thm. 166 Proof of Thm. 166. Without loss of generality, we focus on the equivalence between b = B 11 and a = a 1 . Note that b is essentially expressed as b = log(a). Suppose b (t) = log a (t) . Then, we have ∇ a (t) L =∇ b (t) L ∂ log(a) ∂a a (t) =∇ b (t) L 1 a (t) . (B.525) By Eq. (6.36), log a (t+1) is log a (t+1) = log a (t) e −γ (t) a (t) ∇ a (t) L = log a (t) − γ (t) a (t) ∇ a (t) L = log a (t) − γ (t) a (t) ∇ b (t) L/a (t) = log a (t) − γ (t) ∇ b (t) L = b (t) − γ (t) ∇ b (t) L. (B.526) The last row is the ESGD update formula for b. Therefore, if b (0) = log a (0) , the two optimization procedures yield equivalent iter- ates throughout training. 389 B.10. Product Cholesky Metrics B.10 Product Cholesky Metrics B.10.1 Proof of Thm. 169 Proof. Since θ-DPM is the product metric ofLT 0 (n),g E and n copies ofR ++ ,g θ-E , we first show the Riemannian operators on R ++ ,g θ-E . We can then readily obtain the Riemannian operators on L n ++ ,g θ-DE by the principles of product metrics. As shown by Thanwerdas and Pennec [194], g θ-E is the pullback metric of g E by the power function P θ (·) and scaled by 1 θ 2 , expressed as g θ-E = 1 θ 2 P ∗ θ g E . Besides, as constant scaling does not change the Christoffel symbols, the geodesic, Riemannian logarithm and exponential maps, and parallel transport along a geodesic remain the same under g θ-E and P ∗ θ g E . These Riemannian operators under P ∗ θ g E can be obtained by the properties of Riemannian isometries (Thm. 33). Specifically, given p,q ∈ R ++ and w,v ∈ T p R ++ , we have the following: (P θ ) ∗,p (v) = θp θ−1 v,(B.527) g θ-E p (v,w) = 1 θ 2 g E (P θ ) ∗,p (v), (P θ ) ∗,p (w) =⟨p θ−1 v,p θ−1 w⟩ = p 2(θ−1) vw, (B.528) γ (p,v) (t) = P −1 θ P θ (p) + t (P θ ) ∗,p (v) (B.529) = (p θ + tθp θ−1 v) 1 θ (B.530) = p(1 + tθp −1 v) 1 θ , with t∈t∈ R| 1 + tθp −1 v ∈ R ++ ,(B.531) Log p (q) = (P θ ) −1 ∗,p (P θ (q)− P θ (p))(B.532) = 1 θ p 1−θ (q θ − p θ ) = 1 θ p q p θ − 1 ! ,(B.533) PT p→q (v) = (P θ ) −1 ∗,q (P θ ) ∗,p (v) = q p 1−θ v.(B.534) The geodesic distance between p and q under g θ-E is given by d 2 (p,q) = g θ-E p (Log p (q), Log p (q)) = 1 θ 2 (q θ − p θ ) 2 .(B.535) The weighted Fréchet mean (WFM) of p i ∈ R ++ N i=1 with weights w i N i=1 satisfying w i > 0 for all i and P i w i = 1 under g θ-E is defined as WFM(w i ,p i ) = argmin p∈R ++ N X i=1 w i d 2 (p,p i ).(B.536) 390 Appendix B. Proofs Obviously, the WFM ofp i under g θ-E is the same as the one under P ∗ θ g E . Due to the isometry of P ∗ θ g E to g E , the WFM of p i under g θ-E can be calculated as WFM(w i ,p i ) = P −1 θ WFM E (w i ,P θ (p i )) = X N i=1 w i p θ i 1 θ ,(B.537) where WFM E in Eq. (B.537) is the Euclidean WFM, which is the familiar weighted average. So far, we have obtained all the necessary Riemannian operators on R ++ ,g θ-E . Combining these results with the Euclidean space LT 0 (n) yields the results in the the- orem. B.10.2 Proof of Thm. 170 Proof. As in the proof of Thm. 169, we only need to show the Riemannian operators on R ++ ,g m-BW with m∈ R ++ . Expressions for the Riemannian operators under GBWM can be found in Han et al. [93]. Here, we further simplify the associated expressions for the one-dimensional case. Specifically, given p,q ∈ R ++ and w,v ∈ T p R ++ , we have the following: L p,m (v) = v 2mp ,(B.538) g m-BW p (v,w) = 1 2 ⟨L p,m (v),w⟩ = vw 4mp ,(B.539) γ (p,v) (t) = p + tv +L p,m (tv) 2 m 2 p = p + tv + (tv) 2 4p = p(1 + t 2 v p ) 2 ,(B.540) Log p (q) = 2 m(m −2 pq) 1 2 − p = 2 (pq) 1 2 − p = 2p q p 1 2 − 1 ! , (B.541) d 2 (p,q) = g m-BW p (Log p (q), Log p (q)) = 1 4mp 4 (pq) 1 2 − p 2 = 1 mp pq− 2p 3 2 q 1 2 + p 2 = m −1 q− 2p 1 2 q 1 2 + p = m − 1 2 q 1 2 − p 1 2 2 . (B.542) As shown by Han et al. [93], GBWM on S n ++ is the pullback metric of BWM by π(S) = M − 1 2 SM − 1 2 for all S ∈ S n ++ with M ∈ S n ++ . For the specific R ++ ∼ = S 1 ++ , the 391 B.10. Product Cholesky Metrics isometry is simplified as π(p) = m −1 p,∀p∈ R ++ with m∈ R ++ .(B.543) The geodesic eγ (p,v) (t) under BWM on R ++ exists in the interval t ∈ R | 1 + tL p (v) ∈ R ++ [145], which can be simplified as n t∈ R| 1 + t v 2p ∈ R ++ o . Therefore, the geodesic γ (p,v) (t) under g m-BW exists in the interval: t∈ R| 1 + t π ∗,p (v) 2π(p) ∈ R ++ = t∈ R| 1 + t v 2p ∈ R ++ .(B.544) For BWM, the parallel transport is [196, Tab. 6] f PT p→q (v) = q p 1 2 v.(B.545) Therefore, the parallel transport on R ++ ,g m-BW is PT p→q (v) = π −1 ∗,q f PT π(p)→π(q) (π ∗,p (v)) = π −1 ∗,q π(q) π(p) 1 2 π ∗,p (v) ! = q p 1 2 v. (B.546) Finally, we show the WFM on R ++ ,g m-BW . Given p i ∈ R ++ N i=1 with weights w i N i=1 satisfying w i > 0 for all i and P N i=1 w i = 1, the WFM on R ++ ,g m-BW is WFM(w i ,p i ) = argmin p∈R ++ N X i=1 w i d 2 (p,p i ) = argmin p∈R ++ N X i=1 w i m −1 p 1 2 − p 1 2 i 2 = argmin p∈R ++ N X i=1 w i p 1 2 − p 1 2 i 2 = argmin p∈R ++ N X i=1 w i p− 2p 1 2 p 1 2 i = argmin p∈R ++ p− 2p 1 2 N X i=1 w i p 1 2 i . (B.547) 392 Appendix B. Proofs Let f (p) = p− 2p 1 2 P N i=1 w i p 1 2 i . Then, the first- and second-order derivatives are df dp = 1− p − 1 2 X i w i p 1 2 i ,(B.548) d 2 f dp 2 = 1 2 p − 3 2 X i w i p i 1 2 > 0, ∀p∈ R ++ .(B.549) Therefore, the optimal solution can be obtained by setting Eq. (B.548) equal to 0: df dp = 0⇒ p ∗ = X i w i p 1 2 i 2 .(B.550) Combining the above results with the Euclidean geometry on LT 0 (n), one can readily obtain the results. B.10.3 Proof of Thm. 172 Proof. The differential of DPow θ at L∈ Diag + (n) is (DPow θ ) ∗,L (V) = θL θ−1 V, ∀V∈ T L Diag + (n).(B.551) Substituting Eq. (B.551) into Thm. 171 yields the result. B.10.4 Proof of Thm. 173 We first present a useful lemma. Lemma 205. The Riemannian exponential and logarithmic maps and parallel transport along the geodesic are the same under (θ, M)-DBWM and θ /2-DPM. Proof. This follows directly from Thms. 169 and 170. Now we begin to prove Thm. 173. Proof. We first derive the expressions for DPM with a generic nonzero parameter θ. According to Thm. 205, the expressions for (θ, M)-DBWM then follow by replacing θ with θ /2. We omit the superscript C for simplicity. 393 B.10. Product Cholesky Metrics For X ∈ T I n L n ++ , we have the following: Log I n (L) =⌊L⌋ + 1 θ L θ − I n ,(B.552) PT I n →L (X) =⌊X⌋ + L 1−θ X,(B.553) Exp I n X =⌊X⌋ + (I n + θX) 1 θ .(B.554) For the binary operation, substituting Eqs. (B.552) and (B.553) into Eq. (2.77) gives L⊕ K = Exp L PT I n →L Log I n (K) = Exp L PT I n →L ⌊K⌋ + 1 θ K θ − I n = Exp L ⌊K⌋ + 1 θ L 1−θ K θ − I n =⌊L⌋ +⌊K⌋ + L I n + θL −1 1 θ L 1−θ K θ − I n 1 θ =⌊L⌋ +⌊K⌋ + L θ + K θ − I n 1 θ . (B.555) In the third row of Eq. (B.555), the well-definedness of the exponential map requires L + θ 1 θ L 1−θ K θ − I n ∈ Diag + (n)⇔ L θ + K θ − I n ∈ Diag + (n).(B.556) For the gyromultiplication, substituting Eqs. (B.552) and (B.554) into Eq. (2.78) gives t⊙ L = Exp I n (t Log I n (L)) = Exp I n t⌊L⌋ + t θ L θ − I n = t⌊L⌋ + I n + t L θ − I n 1 θ = t⌊L⌋ + tL θ + (1− t)I n 1 θ . (B.557) In the second row of Eq. (B.557), the exponential map requires I n + θ t θ L θ − I n ∈ Diag + (n)⇔ tL θ + (1− t)I n ∈ Diag + (n).(B.558) 394 Appendix B. Proofs B.10.5 Proof of Thm. 174 Proof. In this proof, we assume L,K,J ∈ L n ++ and s,t ∈ R, with all gyro operations satisfying the conditions in Thm. 173. By Thm. 205, we only need to prove the case of θ-DPM. For simplicity, we omit the superscript C. Axiom (G1). Eq. (2.77) implies that the identity element is the identity matrix. Axiom (G2). We define the inverse element of L as ⊖L =−1⊙ L =−⌊L⌋ + 2I n − L θ 1 θ .(B.559) Simple computations show that ⊖L⊕ L = I n . Axiom (G3). Gyroaddition in Thm. 173 indicates that L⊕ (K⊕ J ) = (L⊕ K)⊕ J =⌊L⌋ +⌊K⌋ +⌊J⌋ + L θ + K θ + J θ − 2I n 1 θ . (B.560) Therefore the gyroautomorphism is the identity map, i.e., gyr[L,K] = id. Axiom (G4). This is a direct corollary of (G3). Gyrocommutative law. Gyroaddition in Thm. 173 indicates that L⊕ K = K⊕ L.(B.561) Axiom (V1). This follows from gyromultiplication in Thm. 173 and Eq. (B.559). Axiom (V2). (s + t)⊙ L = (s + t)⌊L⌋ + (s + t)L θ + (1− (s + t))I n 1 θ = s⌊L⌋ + t⌊L⌋ + sL θ + (1− s)I n + tL θ + (1− t)I n − I n 1 θ = (s⊙ L)⊕ (t⊙ L). (B.562) Axiom (V3). (st)⊙ L = (st)⌊L⌋ + stL θ + (1− st)I n 1 θ = (st)⌊L⌋ + s tL θ + (1− t)I n + (1− s)I n 1 θ = s⊙ h t⌊L⌋ + tL θ + (1− t)I n 1 θ i = s⊙ (t⊙ L). (B.563) Axioms (V4) and (V5). These two axioms can be directly obtained, as gyroau- tomorphisms are all identity maps. 395 B.10. Product Cholesky Metrics B.10.6 Proof of Thm. 177 Proof. According to Nguyen and Yang [159, Thm. 2.4], gyrovector operations are pre- served under Riemannian isometries. Moreover, the Cholesky decomposition is a Rie- mannian isometry: Chol :S n ++ ,g S →L n ++ ,g C .(B.564) B.10.7 Proof of Thm. 179 Proof. Substituting the associated operators in Sec. 6.3.4 into Thm. 195 yields the results. 396