Paper deep dive
Small-scale photonic Kolmogorov-Arnold networks using standard telecom nonlinear modules
Luca Nogueira CalƧado, Sergei K. Turitsyn, Egor Manuylovich
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 4:31:50 AM
Summary
The paper introduces Small-scale Photonic Kolmogorov-Arnold Networks (SSP-KANs) implemented using standard telecommunications components, specifically Mach-Zehnder interferometers, semiconductor optical amplifiers, and variable optical attenuators. The architecture demonstrates strong nonlinear inference performance on classification and regression tasks, achieving high accuracy with significantly fewer parameters than traditional models, while maintaining robustness under realistic hardware impairments like optical noise and input quantization.
Entities (5)
Relation Signals (3)
SSP-KAN ā utilizes ā MZI-VOA-SOA-VOA
confidence 100% Ā· Each network edge employs a trainable nonlinear module composed of a Mach-Zehnder interferometer, semiconductor optical amplifier, and variable optical attenuators
SSP-KAN ā evaluatedon ā Two Moons
confidence 95% Ā· To establish that SSP-KAN can learn nonlinear decision boundaries, we evaluate on the Two Moons dataset
SSP-KAN ā evaluatedon ā Yacht Hydrodynamics
confidence 95% Ā· To evaluate SSP-KAN on a higher-dimensional task, we use the UCI Yacht Hydrodynamics dataset
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Photonic neural networks promise ultrafast inference, yet most architectures rely on linear optical meshes with electronic nonlinearities, reintroducing optical-electrical-optical bottlenecks. Here we introduce small-scale photonic Kolmogorov-Arnold networks (SSP-KANs) implemented entirely with standard telecommunications components. Each network edge employs a trainable nonlinear module composed of a Mach-Zehnder interferometer, semiconductor optical amplifier, and variable optical attenuators, providing a four-parameter transfer function derived from gain saturation and interferometric mixing. Despite this constrained expressivity, SSP-KANs comprising only a few optical modules achieve strong nonlinear inference performance across classification, regression, and image recognition tasks, approaching software baselines with significantly fewer parameters. A four-module network achieves 98.4\% accuracy on nonlinear classification benchmarks inaccessible to linear models. Performance remains robust under realistic hardware impairments, maintaining high accuracy down to 6-bit input resolution and 14 dB signal-to-noise ratio. By using a fully differentiable physics model for end-to-end optimisation of optical parameters, this work establishes a practical pathway from simulation to experimental demonstration of photonic KANs using commodity telecom hardware.
Tags
Links
- Source: https://arxiv.org/abs/2604.08432v1
- Canonical: https://arxiv.org/abs/2604.08432v1
Trouble viewing inline? Open PDF directly ā
Full Text
69,504 characters extracted from source content.
Expand or collapse full text
Small-scale photonic Kolmogorov-Arnold networks using standard telecom nonlinear modules Luca Nogueira CalƧado 1* , Sergei K. Turitsyn 1 and Egor Manuylovich 1 1* Aston Institute of Photonic Technologies (AiPT), Aston University, Birmingham, B4 7ET, United Kingdom. *Corresponding author(s). E-mail(s): l.cal_ado@aston.ac.uk; Contributing authors:s.k.turitsyn@aston.ac.uk; e.manuylovich@aston.ac.uk; Abstract Photonic neural networks promise ultrafast and energy-eļ¬icient inference, yet most existing architectures rely on linear optical meshes combined with elec- tronic nonlinearities, reintroducing optical-electrical-optical bottlenecks that limit speed and scalability. Here we introduce and demonstrate a concept of small-scale photonic Kolmogorov-Arnold networks (SSP-KANs) implemented entirely with standard telecommunications components. As an example of par- ticular implementation, we examine trainable nonlinear module composed of a Mach-Zehnder interferometer, semiconductor optical amplifier, and variable opti- cal attenuators, providing a compact four-parameter transfer function derived from gain saturation and interferometric mixing. Despite the limited functional expressivity of each module, restricted to only four tunable parameters, such small-scale networks comprising only few optical modules achieve strong nonlin- ear inference performance matching their software-equivalent counterparts across classification, regression, and image recognition tasks. A single-layer network with four-module architecture achieves 98.4% accuracy on nonlinear classifica- tion benchmarks inaccessible to linear models, approaching the software baseline of 99.7% with significantly fewer parameters. Performance remains robust under realistic hardware impairments, maintaining high accuracy down to 6-bit input resolution and signal-to-noise ratios of 14 dB. Scaling studies on image classi- fication further confirm that physically constrained nonlinear modules provide measurable advantages over linear photonic models on nonlinearly separable tasks. By implementing trainable nonlinear computation in commodity tele- com hardware and using a fully differentiable physics model for end-to-end 1 arXiv:2604.08432v1 [physics.optics] 9 Apr 2026 optimisation, this work establishes a practical pathway from simulation to exper- imental demonstration of photonic KANs. Our results show that small-scale, physically constrained networks are suļ¬icient for meaningful nonlinear inference, opening a realistic route toward deployable all-optical computing systems for latency-critical applications. Keywords:Photonic neural networks, Optical computing, Neuromorphic photonics, Kolmogorov-Arnold Networks 1 Introduction As neural networks become increasingly embedded within optical technologies - from high-capacity communications and information processing to advanced imaging and sensing (see e.g. [1ā9] and references therein) - there is rising demand for infer- ence hardware capable of matching the speed and bandwidth of light itself. Although photonic neural networks (PNNs) promise orders-of-magnitude improvements in latency and energy eļ¬iciency by performing computation directly in the optical domain, most existing implementations follow the conventional multilayer perceptron (MLP) paradigm. In these architectures, linear optical matrix-vector multiplication, implemented using interferometric meshes or wavelength multiplexing, is combined with electronic nonlinear activation functions. This hybrid approach reintroduces optical-electrical-optical (OEO) conversion bottlenecks, thereby limiting scalability and offsetting many of the anticipated advantages of all-optical processing [ 10,11]. Kolmogorov-Arnold Networks (KANs) offer a fundamentally different computation paradigm [12]. Inspired by the Kolmogorov-Arnold representation theorem, KANs depart from the conventional node-based activation model by placing trainable uni- variate nonlinear functions on network edges. In contrast to multilayer perceptrons (MLPs), where fixed activation functions follow linear transformations, KANs dis- tribute nonlinearity across connections, restructuring the computational graph itself. Recent work has demonstrated that this architecture can achieve accuracy compa- rable to, and in some cases exceeding, that of MLPs while requiring substantially fewer trainable parameters. Moreover, the learned edge functions admit symbolic approximation, enhancing interpretability and analytical transparency [12]. Recent results in short-reach IM/D optical communication systems show that KAN-based equalizers can achieve superior or comparable BER performance while using substantially fewer trainable parameters than conventional FNN and CNN architectures, demonstrating reductions of up toā¼66% [13]. These properties are par- ticularly attractive for photonic implementation, where each nonlinear element incurs optical loss, hardware complexity, and fabrication overhead. By concentrating expres- sive power into a reduced number of structured nonlinear modules, KANs provide a parameter-eļ¬icient and physically compatible framework for scalable optical neural computation. Several recent advances provide strong motivation for photonic implementations of KANs. Fischeret al.[ 14] applied software-based KANs to nonlinear equalization in 2 112 Gb/s passive optical networks, demonstrating that KAN equalizers outperform convolutional neural networks at comparable computational complexity. These results highlight the potential of KAN architectures for high-speed optical signal processing and motivate their direct realization in the optical domain. Penget al.[15] reported the first photonic KAN implementation using ring-assisted Mach-Zehnder interferometers (RAMZI), leveraging free-carrier dispersion in silicon microring resonators to create tunable nonlinear transfer functions. Their integrated silicon photonic architecture achieved 98% accuracy on MNIST handwritten digit recognition while reducing the energy-area product by approximately 65-fold compared to conventional MZI-based optical neural networks. More recently, Stroev and Berloff [16] introduced a theoretical framework for all-optical KANs implemented on spatial light modulator platforms, showing that trainable univariate nonlinear functions can arise from structural interference effects without relying on conventional nonlinear optical materials. Despite these advances, two fundamental challenges remain. First, it is unclear whether the constrained nonlinear transfer functions available from practical optical components (often defined by only a small number of tunable physical parameters) can achieve expressive capability comparable to software-based activation functions such as B-splines, which may involve hundreds of learnable coeļ¬icients per edge. Second, and more critically for real-world deployment, it remains an open question whether photonic KANs can be implemented using off-the-shelf telecommunications compo- nents that are mass-produced, well-characterized, and optimized for the spectral bands and power levels of fibre-optic systems. Previous photonic KAN demonstrations have relied on specialized resonator geometries or free space spatial light modulator platforms. In contrast, the ability to construct trainable optical neural networks from commodity telecom hardware would substantially lower the barriers to experimental realization, accelerate translation from simulation to laboratory validation, and provide a realistic pathway toward scalable deployment. Here we demonstrate that Mach-Zehnder interferometers combined with semicon- ductor optical amplifiers and variable optical attenuators (MZI-SOA-VOA modules), all standard and easily integrated components of telecommunications infrastructure, can function as compact, trainable nonlinear units for photonic KANs. Our approach differs from prior implementations based on specialized resonator geometries or cus- tom optical platforms, in two key aspects: (i) it leverages SOA gain saturation as a primary (native to telecom) nonlinearity that does not require bespoke fabrica- tion; (i) the nonlinear phase shift from SOA amplification is employed as a source of nonlinearity by placing the SOA inside an MZI. We have developed a fully differen- tiable physics-based model that enables end-to-end gradient optimization of all optical parameters, including injection current, attenuation coeļ¬icients, and interferometric phase. Despite the limited per-module parameter count, small-scale networks comprising only 4-8 optical modules achieve strong nonlinear inference performance. Using com- mercially specified SOA parameters (Thorlabs BOA1554P), we obtain 99.1% accuracy on nonlinear classification benchmarks and a coeļ¬icient of determinationR 2 = 0.977 3 on multivariate regression. Scaling studies on MNIST image classification further demonstrate competitive performance, reaching 92.7% accuracy. Systematic robust- ness analysis under optical noise and finite-resolution input quantization confirms stable operation across realistic impairment regimes. Overall, we introduce a small-scale photonic KAN design framework grounded entirely in standard telecommunications components. The approach is inherently mod- ular and can be extended to alternative nonlinear devices and photonic architectures, providing a foundation for telecom-grade implementations of trainable optical neu- ral networks. Rather than pursuing asymptotic universality, this work demonstrates that physically constrained, low-parameter photonic KANs are suļ¬icient for mean- ingful nonlinear inference in realistic settings. Such architectures are particularly well suited to structured computational tasks, including physics-informed modelling, low-dimensional regression, symbolic discovery, latency-critical edge inference, and applications where interpretability and hardware eļ¬iciency are paramount. 4 2 Results 2.1 SSP-KAN architecture Input Feature 1 Input Feature 2 MZI-VOA-SOA-VOA Module [2,2] SSP-KAN Schematic Decision Boundary 0,2,1 0,2,2 0,1,2 0,1,1 PD PD x 0,1 x 0,2 x 1,1 x 1,2 Transfer Function Family (a) (d) (b) (c) Fig. 1SSP-KAN architecture and optical module.(a)Schematic of a single MZI-VOA-SOA-VOA module. The input is split by a 50/50 coupler into a lower arm containing VOA 1 , SOA, and VOA 2 , and an upper reference arm with phase shifterĻ. The arms recombine at a second coupler, producing interference-dependent output power.(b)Family of transfer functions obtained by randomly sampling the four trainable parameters (injection currentI, attenuationsα 1 andα 2 , phaseĻ), illustrating the range of nonlinear inputāoutput mappings achievable from a single module.(c)A minimal [2,2] SSP-KAN network. Each input feature is encoded as optical power, split to dedicated MZI-VOA- SOA-VOA modulesĻ i,j on each edge, and the per-edge outputs are summed on photodetectors (PD). (d)Learned decision boundary on the concentric rings task. The nonlinear boundary emerges from the summation of four individually trained transfer functions, demonstrating that the network composes simple optical nonlinearities to separate classes that are not linearly separable. Following the KAN architecture, we place a dedicated optical module on each edge connecting input and output nodes. Figure1a shows a minimal [2,2] network with two input nodes and two output nodes, requiring four MZI-VOA-SOA-VOA modules. 5 Each output is computed as a sum of the nonlinearly-transformed inputs: y j = n in X i=1 Ļ i,j (x i )(1) whereĻ i,j represents the transfer function of the module on edge(i,j). This is the key structural difference from multilayer perceptrons: rather than applying a fixed non- linearity after a weighted sum, each edge independently transforms its input through a physically distinct optical transfer function. The summation itself can be performed through incoherent power combining on a photodetector, requiring no electronic computation within the network. Each module (Fig.1b) consists of a Mach-Zehnder interferometer with a semi- conductor optical amplifier and variable optical attenuators in one arm, providing a tunable nonlinear transfer function controlled by four trainable parameters. The SOA injection currentIsets both the gain magnitude and the degree of saturation nonlin- earity, ranging from approximately linear response at low currents to pronounced gain compression at high currents. The input attenuationα 1 controls the optical power entering the SOA, thereby setting the operating point on the gain saturation curve: low attenuation drives the SOA into saturation with strong nonlinearity, while high atten- uation keeps it in the linear amplification regime. The output attenuationα 2 scales the amplified arm contribution to the interference without affecting the saturation dynam- ics. The interferometer phaseĻselects which portion of the MZI cosine-squared fringe pattern reaches the output, choosing between constructive and destructive interfer- ence regimes. Together, these four parameters per module define a family of transfer functions that spans a rich space of nonlinear input-output relationships (Fig. 1c). All simulations use manufacturer-specified parameters for a commercially available SOA (Thorlabs BOA1554P; see Methods for the full physical model). 2.2 Nonlinear classification on Two Moons To establish that SSP-KAN can learn nonlinear decision boundaries, we evaluate on the Two Moons dataset, a standard binary classification benchmark where two interleaving half-circles cannot be separated by any linear classifier. The dataset com- prises 1,000 training and 1,000 test samples with Gaussian noise (Ļ= 0.1), scaled to the optical power range 10ā80 mW. We train a single-layer [2,2] network comprising 4 MZI-VOA-SOA-VOA modules (16 trainable parameters) and compare against a lin- ear baseline and a software KAN baseline implemented using pyKAN [17] with grid sizeG=1and B-spline activation functions. 6 Fig. 2Two Moons classification with SSP-KAN [2,2].(a)Learned decision boundary (98.4% test accuracy, 16 parameters). The network separates the two crescents using a smooth nonlinear boundary generated by four MZI-VOA-SOA-VOA modules. Inset: [2,2] network schematic showing 2 inputs, 4 modules, and 2 outputs.(b)Test accuracy and parameter count across models. The linear baseline (89.2%, 6 parameters) confirms the task is intrinsically nonlinear. pyKAN (G=1, 40 parameters) reaches 99.9%; SSP-KAN achieves comparable accuracy (98.4%) with fewer than half the parameters and using only physically realisable optical activation functions. Inset: [2,2] network schematic showing 2 inputs, 4 modules, and 2 outputs. Figure2a shows the learned decision boundary for the [2,2] architecture. The network achieves 98.4% test accuracy, separating the two crescents with a smooth nonlinear boundary that follows the data geometry. The linear baseline reaches only 89.2% (Fig.2b), confirming the intrinsically nonlinear nature of the task. The software KAN baseline (pyKAN,G=1) achieves 99.9% with 40 parameters; SSP-KAN closes to within 1.5 percentage points using 16 parameters and a transfer function family constrained to the physics of MZI-VOA-SOA-VOA modules. This small accuracy gap demonstrates that the nonlinear response of saturated semiconductor optical ampli- fiers, shaped by the four trainable parameters per edge, provides suļ¬icient functional diversity to approximate the B-spline activation functions used in software KANs. A deeper [2,2,2] architecture with 8 modules marginally improves no-noise accuracy (Supplementary SectionA); however, multi-layer operation introduces coherent phase accumulation at intermediate summation nodes, which is analysed in Supplementary SectionB. Having established classification performance under ideal conditions, we next eval- uate robustness to the two dominant hardware impairments in photonic systems: finite digital-to-analogue converter (DAC) resolution for input encoding and optical noise from amplified spontaneous emission (ASE). Figure3a shows classification accuracy for the [2,2] network across input quantization levels (12-bit down to 4-bit) and optical SNR conditions (30 dB down to 6 dB). Quantization has negligible effect on perfor- mance: at any given SNR, accuracy is nearly constant across all tested resolutions. This implies that inexpensive, low-resolution DACs are suļ¬icient for input encoding. 7 Fig. 3Hardware robustness of SSP-KAN [2,2] on Two Moons.(a)Test accuracy as a function of input DAC resolution (12-bit to 4-bit) and optical noise (SNR 30 dB to 6 dB). The smooth colour gradient confirms that quantization has negligible impact, while performance degrades vertically with decreasing SNR. The 90% and 95% contour lines indicate the practical operating boundary. Inset: [2,2] network schematic (16 parameters).(b)Decision boundaries at four corner conditions. Top row: 30 dB noise at 12-bit (95.4%) and 4-bit (94.5%) resolution. Bottom row: 6 dB noise at 12-bit (87.9%) and 4-bit (87.8%). The boundary shape is preserved across quantization levels but simplifies under severe noise. Optical noise, by contrast, produces measurable degradation. At SNR=30 dB accuracy remains above 95%, dropping to approximately 97% at 20 dB, 96% at 14 dB, and 94% at 10 dB. Under extreme noise (SNR=6 dB), the network retains approx- imately 87% accuracy, falling below the linear baseline (89.2%). At suļ¬iciently high noise levels the nonlinear decision boundary becomes a liability: the optimiser con- verges to a coarser separation that is less effective than a simple linear cut. The decision boundaries under extreme conditions (Fig.3b) illustrate this directly. Under 6 dB noise, the boundary approximates a broad diagonal rather than tracking the crescent geometry. This crossover defines a practical operating threshold for the [2,2] architecture: for SNRā„10 dB, the nonlinear model consistently outperforms the lin- ear baseline, while below 10 dB, simpler models may be preferable. Fig.3b confirms the quantization finding: the 4-bit boundary (87.8%) is visually and numerically indis- tinguishable from the 12-bit boundary (87.9%) at the same noise level. For realistic deployment conditions (8-bit DAC, SNRā„14 dB), the [2,2] architecture maintains accuracy above 96%, comfortably exceeding the linear baseline. 2.3 Regression performance on yacht hydrodynamics To evaluate SSP-KAN on a higher-dimensional task, we use the UCI Yacht Hydro- dynamics dataset [18], derived from Delft ship hydrodynamics experiments [19]. The dataset contains 308 samples, each described by six hull geometry features (longitu- dinal position of the centre of buoyancy, prismatic coeļ¬icient, length-displacement 8 ratio, beam-draught ratio, length-beam ratio, and Froude number) with the pre- diction target being residuary resistance per unit weight of displacement. This regression task tests whether SSP-KAN can learn complex multivariate functions from six-dimensional inputs using a small number of optical modules. We evaluate two architectures: a single-layer [6,1] network with 6 MZI-VOA- SOA-VOA modules (24 trainable parameters) and a two-layer [6,1,1] network with 7 modules (28 parameters). Figure4compares both against baseline models. A linear regression baseline achievesR 2 = 0.664, confirming the nonlinear nature of the hull- resistance relationship. A multilayer perceptron (MLP) with architecture [6ā4ā1] and 33 parameters achievesR 2 = 0.989, while a software KAN (pyKAN, [6,1] with grid sizeG=1, 60 parameters) reachesR 2 = 0.977. The single-layer SSP-KAN [6,1] achievesR 2 = 0.869: a substantial improvement over the linear baseline, but limited by the fact that each of its six modules computes a univariate function of a single input feature, with no capacity for feature interaction. Adding a second layer improves performance dramatically: SSP-KAN [6,1,1] reachesR 2 = 0.977, matching pyKAN exactly while using fewer than half the parameters (28 vs 60). The second-layer mod- ule composes the six first-layer outputs into a single nonlinear function, enabling the network to capture multivariate interactions that no sum of univariate functions can represent. 9 Fig. 4Regression performance on yacht hydrodynamics. TestR 2 scores and parameter counts for linear regression, MLP, software KAN (pyKAN,G=1), and SSP-KAN architectures [6,1] and [6,1,1]. SSP-KAN [6,1,1] matches pyKAN (R 2 = 0.977) with 28 parameters versus 60, demonstrating that compositional depth compensates for the constrained per-edge activation family. We systematically evaluate robustness to hardware impairments across SNR lev- els from 6 dB to 30 dB conditions and DAC resolutions from 4-bit to 12-bit (Fig.5). Figure5a,b showsR 2 as a function of input resolution at each noise level for the [6,1] and [6,1,1] architectures, respectively. As with Two Moons classification, input quan- tization has minimal effect: for the [6,1,1] architecture at SNR = 30 dB,R 2 remains between 0.975 and 0.977 from 12-bit down to 8-bit, with a modest decrease to approx- imately 0.95 at 4-bit. Optical noise produces more significant degradation, reducing [6,1,1] performance fromR 2 = 0.977at SNR = 30 dB conditions to approximately 0.95 at SNR=20 dB, 0.91 at 14 dB, 0.87 at 10 dB, and 0.78 at 6 dB. 10 Fig. 5Hardware robustness of SSP-KAN on yacht hydrodynamics regression. TestR 2 as a function of input DAC resolution (12-bit to 4-bit) and optical noise (SNR 30 dB to 6 dB) for(a)the single-layer [6,1] architecture (6 modules, 24 parameters) and(b)the two-layer [6,1,1] architecture (7 modules, 28 parameters). Contour lines markR 2 thresholds. The [6,1] panel shows a compressed performance range (R 2 ā0.78ā0.87), while [6,1,1] spans a wider range (R 2 ā0.78ā0.98), reflecting both higher peak performance and greater noise sensitivity. A notable feature of Fig.5is the contrast between the two panels. The [6,1] architecture (Fig.5a) shows performance compressed into a narrow range between R 2 ā0.78and 0.87 across the entire parameter space. The [6,1,1] architecture (Fig. 5b) spans a much wider range, fromR 2 ā0.98at 30 dB down to 0.78 at 6 dB, with contour lines indicating where each performance threshold is crossed. This contrast directly visualises the depth-robustness trade-off: the deeper network achieves substantially higher performance under mild impairment but is more sensitive to perturbation, as the second-layer module amplifies noise that has already propagated through the first layer. The shallower [6,1] network, by contrast, has less capacity to lose. Both archi- tectures converge to similar performance (R 2 ā0.78) under extreme noise, suggesting a noise floor set by the intrinsic diļ¬iculty of the regression task rather than by archi- tectural capacity. Under combined impairments representative of realistic deployment (8-bit DAC, SNR=14 dB), the [6,1,1] architecture maintainsR 2 ā0.91, well above the linear baseline (R 2 = 0.664). A mild increase in quantization sensitivity is visi- ble at 4-bit for the [6,1,1] network under lower SNR conditions, suggesting that 6-bit DAC resolution represents a practical lower bound for the deeper architecture. 2.4 Image classification on MNIST and Fashion-MNIST To evaluate SSP-KAN on high-dimensional inputs representative of practical inference tasks, we tested on MNIST handwritten digit classification (784-dimensional input, 10 classes) and Fashion-MNIST clothing classification (same dimensionality, 10 classes). MNIST serves as a direct comparison point with the D-RAMZI photonic KAN of Penget al.[ 15], while also testing whether the MZI-VOA-SOA-VOA transfer function can handle image-scale inputs. Fashion-MNIST is a substantially harder benchmark: 11 its classes share more visual similarity (e.g. shirts, coats, and pullovers occupy over- lapping regions of pixel space), making it a more demanding test of whether optical nonlinearity provides meaningful benefit over linear classification. We evaluated three SSP-KAN architectures of increasing capacity: a single-layer [784,10] network with 7,840 MZI-VOA-SOA-VOA modules, a two-layer [784,10,10] network with 7,940 modules, and a wider two-layer [784,20,10] network with 15,880 modules. All models were trained with identical hyperparameters (AdamW optimiser, weight decay10 ā3 , 150 epochs; see Methods). A linear classifier trained under the same conditions serves as the baseline, isolating the contribution of optical nonlinearity from the effect of increased model capacity. On MNIST, the linear baseline achieves 91.0% accuracy, consistent with the well- known near-linear separability of handwritten digits in pixel space. SSP-KAN [784,10] reaches 91.8%, SSP-KAN [784,10,10] reaches 92.0%, and SSP-KAN [784,20,10] reaches 92.7%, a gain of 1.7 percentage points over the linear model (Fig.6). On Fashion- MNIST, where the classification task is genuinely nonlinear, the advantage of optical nonlinearity becomes more pronounced. The linear baseline drops to 83.3%, while SSP- KAN [784,20,10] achieves 85.7%, widening the gap to 2.4 percentage points (Fig.6b). This pattern is consistent across all architectures: the SSP-KAN improvement over the linear model is systematically larger on Fashion-MNIST than on MNIST, confirm- ing that the MZI-VOA-SOA-VOA transfer functions provide greater benefit on tasks where nonlinear feature interactions are essential for classification. The confusion matrices (Fig.6c,d) reveal that SSP-KAN errors are concentrated among visually similar classes. On MNIST, the most common confusions occur between digits that share structural features (e.g. 4/9, 3/5, 7/9). On Fashion-MNIST, the dominant error mode is confusion among upper-body garments ā shirt, coat, pullover, and T-shirt/top ā which differ primarily in subtle textural features rather than gross shape. These error patterns are consistent with the limited expressivity of four-parameter optical transfer functions: the network captures coarse nonlinear features effectively but cannot resolve fine-grained distinctions that would require higher-order activation function complexity. 12 Fig. 6Image classification results. Training curves showing accuracy and loss for (a) MNIST and (b) Fashion-MNIST using the [784,20,10] architecture. Insets show the final training epochs at expanded scale. Confusion matrices for the [784,20,10] model on (c) MNIST and (d) Fashion-MNIST, showing that classification errors concentrate among visually similar classes. 3 Methods 3.1 Kolmogorov-Arnold Networks The Kolmogorov-Arnold representation theorem states that any multivariate con- tinuous functionf: [0,1] n āRcan be exactly represented as a superposition of continuous univariate functions and addition: f(x 1 ,...,x n ) = 2n X q =0 Φ q n X p =1 Ļ q,p (x p ) ! (2) whereĻ q,p : [0,1]āRandΦ q :RāRare continuous univariate functions [12]. This theorem implies that high-dimensional function approximation can be reduced 13 to learning one-dimensional functions, fundamentally differing from the multilayer perceptron (MLP) approach where fixed nonlinearities are applied at nodes. KANs leverage this representation by placing learnable activation functions on network edges rather than fixed activations at nodes. In a standard MLP, the transfor- mation between layers takes the formy=Ļ(Wx+b), whereĻis a fixed nonlinearity (e.g., ReLU or sigmoid) applied element-wise after a linear transformation. In contrast, a KAN layer computes: y j = n in X i=1 Ļ i,j (x i )(3) This architectural difference has significant implications for photonic implementa- tion. In conventional photonic neural networks, the linear matrix-vector multiplication Wxis naturally implemented using interference in Mach-Zehnder interferome- ter meshes or wavelength-division multiplexing, but the nonlinear activationĻ(Ā·) typically requires optical-to-electrical conversion, electronic processing, and electrical- to-optical conversion [10]. This OEO bottleneck limits throughput and increases power consumption. The KAN architecture inverts this challenge: rather than requiring a single power- ful nonlinearity after linear mixing, KANs require many weaker nonlinearities applied independently to each signal before linear summation. Optical summation (combining beams on a photodetector or through interference) is straightforward, while tunable optical nonlinearities, though constrained in functional form, are achievable through various physical mechanisms including saturable absorption, gain saturation, and nonlinear phase shifts [20]. In software implementations, the univariate functionsĻ i,j are typically parame- terized as B-splines with learnable control points, providing smooth, flexible function approximation withG+kparameters per edge, whereGis the grid size andkis the spline order [12]. A key question for photonic KANs is whether the more constrained transfer functions available from optical components, with only a handful of tunable parameters, can provide suļ¬icient expressivity for practical tasks. 3.2 MZI-VOA-SOA-VOA optical module Our photonic KAN implementation uses a composite optical module consisting of a Mach-Zehnder interferometer (MZI) with a semiconductor optical amplifier (SOA) and variable optical attenuators (VOAs) in one arm. This configuration provides a tunable nonlinear transfer function with four trainable parameters: the SOA injection currentI, two VOA attenuation coeļ¬icientsα 1 andα 2 , and the interferometer phase Ļ. 3.3 Module Architecture The MZI splits the input optical powerP 0 equally between two arms via a 50/50 coupler (Fig.1b). The lower arm contains VOA 1 , followed by the SOA, followed by VOA 2 . The upper arm provides a reference path with adjustable phaseĻ. The signals recombine at a second 50/50 coupler, producing interference that shapes the output. 14 The power entering the SOA is: P SOA,in =α 1 Ā· P 0 2 (4) whereα 1 = 10 āA 1 /10 is the linear transmission corresponding to VOA 1 attenuation A 1 in dB. Each component serves a distinct purpose: the injection currentIcontrols the SOA small-signal gainh 0 ; VOA 1 sets the operating point on the gain saturation curve by controlling the power entering the SOA; VOA 2 scales the amplitude of the SOA arm output independently of the saturation dynamics; and the phaseĻdetermines the interference condition between the two arms. 3.4 SOA gain dynamics We model the SOA using the AgrawalāOlsson formalism, where the integrated gain h= R L 0 g(z)dzevolves according to the rate equation: dh dt = h 0 āh Ļ c ā P in E sat (e h ā1)(5) whereh 0 is the small-signal integrated gain,Ļ c is the carrier lifetime, andE sat is the saturation energy. At steady state (dh/dt= 0), this becomes: h=h 0 ā P in P sat (e h ā1)(6) whereP sat =E sat /Ļ c is the saturation power. This transcendental equation is solved numerically by NewtonāRaphson iteration: h n+1 =h n ā h n āh 0 + (P in /P sat )(e h n ā1) 1 + (P in /P sat )e h n (7) initialised ath 0 , with three iterations suļ¬icient for convergence across the operating range. The power gain isG=e h . The small-signal gainh 0 depends on the injection currentI. Above the trans- parency currentI tr , we use a calibrated relationship: h 0 (I) =Īŗ I I tr ā1 (8) whereh 0,max = lnG max and the gain coeļ¬icientĪŗ=h 0,max /(I max /I tr ā1)is determined from device specifications. For our simulations, we use parameters corre- sponding to a commercially available SOA (Thorlabs BOA1554P): saturation power P sat = 18dBm, maximum small-signal gainG max = 35dB at maximum current I max = 1700mA, transparency currentI tr = 600mA, and linewidth enhancement factorα H = 5. 15 A critical feature of SOAs is the coupling between gain and phase through the linewidth enhancement factor. The SOA transforms the complex field envelope as: E out =E in exp 1āiα H 2 h (9) encapsulating both amplitude gain (G=e h ) and a nonlinear phase shift: āĻ SOA =ā α H 2 h(10) This gaināphase coupling means that intensity-dependent gain saturation simultane- ously produces intensity-dependent phase modulation, enriching the nonlinear transfer function beyond what either effect alone could provide. 3.5 Transfer Function Each 50/50 directional coupler is described by the unitary transfer matrix: C= 1 ā 2 1i i1 (11) For input fieldE 0 = ā P 0 entering port 1, the coupler directsE 0 / ā 2to the SOA arm andiE 0 / ā 2to the reference arm. After propagation through the VOA 1 āSOAāVOA 2 cascade (Eqs.4ā9) and the reference-arm phase shifter, the two arm fields are: E out 1 = ā α 1 α 2 G ā 2 e āiα H h/2 E 0 , (12) E out 2 = i ā 2 e iĻ E 0 . (13) Applying the output coupler (Eq. 11) and taking the transmitted port gives the output field: E out = E 0 2 h p α 1 α 2 Ge āiα H h/2 āe iĻ i (14) The detected output powerP out =|E out | 2 yields the complete module transfer function: P out = P 0 4 α 1 α 2 G+ 1ā2 p α 1 α 2 Gcos α H h 2 +Ļ (15) whereG=e h andh=h(P SOA,in , h 0 )is the solution to Eq. ( 6). The transfer function has the form of a MachāZehnder interference fringe whose visibility, fringe position, and input-dependent curvature are all controlled by the four trainable parameters. The key insight is thatα 1 affects not only the amplitude scaling but also the degree of gain saturation through Eq. (4), coupling the interference condition to the SOA operating point. 16 3.6 Role of the two VOAs The two attenuators provide qualitatively different control knobs over the MZIāSOA module. VOA 1 (pre-SOA) sets the SOA operating point viaP SOA,in =α 1 P 0 /2, and therefore tunes the degree of gain saturation and the associated gain-induced phase shift, changing theshapeof the nonlinear transfer function. In contrast, VOA 2 (post- SOA) does not alter the saturation dynamics because it is placed after the SOA; instead, it scales the relative amplitude of the SOA arm at recombination and thus mainly controls the interference contrast and output scaling. Figure7illustrates these distinct effects by sweeping VOA 1 at fixed VOA 2 and vice versa, for a representative operating point (I= 1200mA,Ļ=Ļ/2,P sat = 18dBm,α H = 5, withh 0 (I) calibrated fromG max = 35dB atI max = 1700mA andI tr = 600mA). 000 0 0 !P ! P ( " !P ! P ( " Fig. 7Distinct roles of the two VOAs in the MZIāSOA module transfer function. (a) Sweeping VOA 1 changes the SOA operating point and saturation strength, producing pronounced changes in the nonlinearity shape. (b) Sweeping VOA 2 primarily rescales the SOA-arm contribution at recom- bination, mainly adjusting interference contrast and overall scaling. Fixed parameters:I= 1200mA, Ļ=Ļ/2,P sat = 18dBm,α H = 5, withh 0 (I)calibrated fromG max = 35dB atI max = 1700mA andI tr = 600mA. In (a) VOA 2 is fixed atA 2 = 10dB; in (b) VOA 1 is fixed atA 1 = 10dB. 3.7 Trainable Parameters Each MZI-VOA-SOA-VOA module provides four independently trainable parame- ters. The injection currentI(600ā1700 mA) controls the small-signal gainh 0 through Eq. ( 8), with higher current producing stronger gain and more pronounced saturation nonlinearity. The input attenuationα 1 (0ā30 dB) sets the operating point on the sat- uration curve via Eq. (4): high attenuation keeps the SOA in its linear regime, while low attenuation drives it into saturation. The output attenuationα 2 (0ā30 dB) scales the SOA arm contribution to the interference without affecting the saturation dynam- ics. Finally, the phaseĻ(0ā2Ļ) determines whether the interference is constructive, destructive, or intermediate. 17 The entire forward pass ā from Eq. (4) through the NewtonāRaphson solution of Eq. (6) to the transfer function Eq. (15) ā is implemented in PyTorch with automatic differentiation. Gradients with respect to all four parameters are computed via back- propagation through the Newton iteration, enabling end-to-end training of networks composed of multiple MZI-VOA-SOA-VOA modules. 3.8 Noise and quantization models Practical deployment of photonic neural networks requires robustness to hardware impairments. We model two primary sources of signal degradation: optical noise from amplified spontaneous emission (ASE) in the SOA, and quantization noise from finite- resolution digital-to-analog converters (DACs) used to encode input signals as optical power levels. 3.8.1 Optical Noise Models SOAs introduce noise through amplified spontaneous emission (ASE), which adds random fluctuations to the optical signal. Our simulation framework implements two noise models with distinct physical assumptions. The amplitude-domain model captures the physics of ASE noise, where random optical fields add coherently to the signal field. For an input powerP in , we convert to field amplitudeA= ā P in , add complex Gaussian noise, and compute the output power: P noisy = A+Ļ n r +jn i ā 2 2 = (A+n ā² r ) 2 +n ā²2 i (16) wheren r andn i are independent random variables drawn from a standard normal distribution andĻ= p P max /SNR lin sets the noise strength relative to the maximum operating power and the linear-scale signal-to-noise ratio. This model produces signal- dependent variance Var[P noisy ]ā2P in Ļ 2 , meaning low-power signals experience worse relative noise, consistent with the physics of optical detection. The power-domain model provides a simplified alternative where Gaussian noise is added directly to the optical power: P noisy =P in +ĻĀ·n(17) wherenis again drawn from a standard normal distribution and the result is clamped to non-negative values. This model produces constant variance independent of signal level. For both models, the noise standard deviation is: Ļ= p P max Ā·10 āSNR dB /10 (18) whereP max is the maximum operating power. In our experiments, we use the power- domain model because it provides intuitive interpretation of SNR values: at 20 dB SNR, the relative noise at maximum power is approximately 1%, making it straight- forward to reason about system performance across different noise conditions. The amplitude-domain model remains available in our codebase for studies. 18 3.8.2 Input Quantization The input to each optical module must be encoded as an optical power level, typically using a laser source modulated by a DAC-controlled attenuator or modulator. Finite DAC resolution introduces quantization error. For ab-bit DAC spanning the power range[P min ,P max ], the quantized input is obtained by rounding to the nearest discrete level: P quant =round P in āP min P max āP min Ā·(2 b ā1) Ā· P max āP min 2 b ā1 +P min (19) where 2 b ā 1 is the number of discrete levels available (for example, 255 for an 8-bit converter). To enable gradient-based training despite the non-differentiable rounding operation, we employ the straight-through estimator (STE): the forward pass uses quantized values while the backward pass computes gradients as if no quantization occurred. We evaluate DAC resolutions from 4 to 16 bits. An 8-bit DAC provides 256 levels over the operating range, corresponding to a step size of approximately 0.3 mW for a 10ā80 mW range. Our results in Section2demonstrate that SSP-KAN maintains high accuracy down to 6-bit resolution, with gradual degradation at lower bit depths. 3.9 Training procedure All SSP-KAN models are trained end-to-end via backpropagation, with optical transfer-function gradients computed by automatic differentiation through the MZI- VOA-SOA-VOA forward model described above. Because no physical hardware is in the loop during training, the procedure reduces to standard gradient-based optimi- sation of the four trainable parameters(I,α 1 ,α 2 ,Ļ)per module, with the physics encoded entirely in the differentiable forward pass. 3.9.1 Datasets and preprocessing We evaluate SSP-KAN on four benchmarks spanning binary classification, multivari- ate regression, and high-dimensional image classification (Table1). Table 1Benchmark datasets and preprocessing details. DatasetSamples Input dim.TaskSplitPower range (mW) Two Moons2,0002Classification50/5010ā80 Yacht Hydro.3086Regression 70/15/1510ā80 MNIST70,000784Classification 72/8/201ā10 Fashion-MNIST 70,000784Classification 72/8/201ā10 For all datasets, inputs are linearly scaled to optical power:P=P min +xĀ·(P max ā P min ), wherexā[0,1]is the normalised input. Two Moons and yacht hydrodynamics use the range 10ā80 mW, representative of typical fibre-compatible operating pow- ers. For MNIST and Fashion-MNIST, a power-range sweep identified 1ā10 mW as 19 optimal, where the MZI-VOA-SOA-VOA transfer functions operate in a more linear regime better suited to the high-dimensional, low-contrast nature of pixel inputs. The Two Moons dataset comprises 1,000 training and 1,000 test samples with Gaussian noise (Ļ= 0.1), generated using scikit-learnāsmake_moons. The yacht hydrodynamics dataset [18,19] contains 308 samples split 70/15/15 into training, validation, and test sets. MNIST and Fashion-MNIST use the full 70,000-sample datasets with stratified 72/8/20 splits, ensuring balanced class representation in each partition. 3.9.2 Optimisation All models are trained using the AdamW optimiser [21]. For Two Moons and yacht hydrodynamics, which involve small datasets, we use full-batch training for 5,000 steps (Two Moons) or 5,000 epochs (yacht), with a OneCycleLR schedule (30% warmup, cosine annealing), weight decay10 ā5 , and learning rates of5Ć10 ā3 (Two Moons) and10 ā2 (yacht SSP-KAN) or5Ć10 ā3 (yacht baselines). For the yacht task, early stopping with a patience of 500 epochs monitors validation loss to prevent overfitting on the small dataset, and the best-validation-loss checkpoint is restored after training. For MNIST and Fashion-MNIST, we train for 150 epochs with batch size 256, learn- ing rate10 ā2 , weight decay10 ā3 , and a 5-epoch linear warmup followed by cosine annealing. The best model by validation accuracy is selected. Classification tasks use cross-entropy loss; the regression task uses mean squared error. No data augmentation, label smoothing, or stochastic weight averaging is applied in any experiment. 3.9.3 Baseline models To isolate the contribution of optical nonlinearity, we compare against several base- lines trained under identical optimisation conditions. A linear model with the same input-output dimensions measures the performance achievable without any nonlin- ear activation. For Two Moons and yacht hydrodynamics, we additionally compare against software KAN (pyKAN [17]) with B-spline grid sizesG= 1andG= 10, repre- senting low- and high-capacity software KAN baselines, and a multi-layer perceptron (MLP) with one hidden layer and ReLU activations, sized to approximately match the SSP-KAN parameter count. For image classification, we compare across SSP- KAN architectures of varying depth and width to characterise the effect of network topology. 3.9.4 Hardware robustness evaluation To assess the effect of realistic hardware impairments, we retrain each model from scratch under every combination of input quantization level and optical noise con- dition, using the same optimiser and schedule as the no-noise baseline. Input quantization is swept from Float32 to 4-bit DAC resolution (16 levels). Optical noise is applied in the power domain (Eq.17) at SNR levels from no noise to 6 dB. This train-under-impairment protocol tests whether the SSP-KAN architecture can learn effective representations when noise and quantization are present throughout training, representing the expected operating condition for physical hardware where impairments are always active. 20 4 Discussion The results demonstrate that standard telecommunications components can serve as trainable nonlinear activation functions within KolmogorovāArnold Networks. Each MZIāSOAāVOA module provides four tunable parameters ā injection currentI, attenuationsα 1 andα 2 , and phaseĻā yielding a transfer function family with suf- ficient diversity to support nonlinear inference across classification, regression, and image recognition tasks. Specifically, SSP-KAN achieves 98.4% accuracy on non- linear classification (Two Moons),R 2 = 0.977on six-input multivariate regression (yacht hydrodynamics), and 92.7% on 784-dimensional image classification (MNIST). All results are obtained from a fully differentiable physics-based model grounded in AgrawalāOlsson SOA gain dynamics [22] and commercially specified device param- eters (Thorlabs BOA1554P), enabling end-to-end optimisation via backpropagation. The Two Moons and yacht hydrodynamics results illustrate how compositional depth compensates for limited per-module expressivity. Each MZIāSOAāVOA module implements a cosine-squared nonlinearity modulated by SOA gain saturation (Eq.15); this four-parameter family is substantially more constrained than B-spline activations with tens or hundreds of degrees of freedom. A single-layer [2,2] SSP-KAN neverthe- less achieves 98.4% accuracy on Two Moons with only four modules, closing to within 1.5 percentage points of the software KAN baseline (pyKAN, 99.9%). On yacht hydro- dynamics, the single-layer [6,1] network (R 2 = 0.869) is limited by its inability to capture feature interactions: each module processes one input independently. Adding a second layer ([6,1,1],R 2 = 0.977) enables the network to compose six univariate functions, matching pyKAN with fewer than half the parameters (28 vs 60). This con- firms that cascading physically constrained nonlinearities produces richer functional representations than increasing parallel width alone. Multi-layer architectures with intermediate optical summation nodes (e.g. the [2,2,2] configuration) introduce addi- tional considerations related to coherent phase accumulation, which are analysed in the Supplementary Material. Robustness analyses confirm resilience to the two dominant hardware impairments: finite DAC resolution and optical noise from amplified spontaneous emission. Across both classification and regression tasks, input quantization has negligible impact ā accuracy andR 2 remain essentially unchanged from Float32 down to 6-bit resolution under all tested noise conditions. Optical noise is the primary limiting factor. On Two Moons, the [2,2] architecture retains>96% accuracy at SNR=14 dB; on yacht hydro- dynamics, the [6,1,1] network maintainsR 2 >0.91at the same SNR. Under extreme noise (SNR=6 dB), both tasks exhibit graceful degradation rather than catastrophic failure, with performance converging toward a noise floor that depends on the intrin- sic task diļ¬iculty rather than on architectural capacity. That quantization effects are negligible while noise effects are measurable indicates that, for practical implemen- tations, investment in low-noise optical amplification will yield greater returns than high-resolution digital-to-analogue conversion. On MNIST and Fashion-MNIST, SSP-KAN achieves 92.7% and 85.7% accuracy, respectively, compared with linear baselines of 91.0% and 83.3%. The modest mar- gin on MNIST reflects the datasetās near-linear separability rather than a limitation of the optical nonlinearity; consistent with this interpretation, the advantage widens 21 on Fashion-MNIST, where nonlinear structure is more pronounced (+2.4 percentage points versus +1.7 on MNIST). Compared with the dual ring-assisted MZI (D- RAMZI) architecture of Penget al.[15], which achieves 98% MNIST accuracy using microring-based nonlinearities with nine tunable parameters per edge and custom silicon photonic fabrication, SSP-KAN employs four parameters per edge using com- mercially available fibre-coupled components. The D-RAMZI approach offers richer per-edge nonlinearities through resonance-enhanced transmission and cascaded inter- ference; the remaining accuracy gap therefore quantifies the cost of reduced activation richness and suggests that enhancing per-module expressivity represents the most direct route toward closing this gap while retaining deployment simplicity. It is important to distinguish between architectures used to probe scaling behaviour and those intended for near-term experimental demonstration. The MNIST- scale [784,20,10] network, comprising 15,880 discrete MZIāSOAāVOA modules, is not proposed as a practical system; assembling such a configuration from fibre-coupled components would be prohibitive in footprint, cost, and thermal stability. Image classification serves instead to characterise SSP-KAN performance trends as input dimensionality and network width increase, informing future integrated implementa- tions. In contrast, the Two Moons and yacht hydrodynamics configurations require only 4ā7 modules and are directly realisable on an optical bench using commercially available components. The present study is based on numerical simulations employing calibrated device parameters. Experimental realisation is facilitated by the use of components operat- ing well within standard specifications: input powers of 1ā80 mW, injection currents of 600ā1700 mA, and attenuation ranges of 0ā30 dB. The primary practical challenges are (i) accurate characterisation of SOA transfer functions under varying bias conditions, and (i) maintaining thermal and mechanical phase stability across interferometric modules. The immediate experimental objective is the [2,2] Two Moons configuration, comprising four fibre-coupled modules initialised with parameters obtained from sim- ulation. Successful demonstration would validate both the fidelity of the physics-based model and the end-to-end training framework before scaling to the seven-module yacht hydrodynamics configuration. The SSP-KAN architecture is particularly relevant to signal equalisation in coher- ent optical communications, where inference latency and bandwidth increasingly challenge electronic digital signal processing at next-generation data rates [1]. The per- edge structure maps naturally to wavelength- or space-division multiplexed systems, enabling parallel channel-wise processing. Integration with reservoir computing could combine the temporal feature extraction of SOA-based dynamical systems with the structured nonlinear readout of SSP-KAN. More broadly, the approach of constructing differentiable forward models directly from component physics and optimising phys- ical parameters via backpropagation extends beyond SOA-based nonlinearities. Any photonic platform with analytically tractable transfer functions ā including Kerr- based or Brillouin-based mechanisms ā can in principle be incorporated into this physics-aware training paradigm. 22 5 Conclusion We have demonstrated SSP-KAN, a photonic implementation of small-scale KolmogorovāArnold Networks in which each network edge is realised by a dedi- cated MZIāSOAāVOA module acting as a trainable nonlinear activation function. Using only four physically grounded parameters per module ā injection current, two attenuation coeļ¬icients, and interferometric phase ā the architecture achieves 98.4% accuracy on nonlinear classification,R 2 = 0.977on six-input regression, and 92.7% on 784-dimensional image classification, all through end-to-end differentiable training of SOA gain dynamics using commercially specified device parameters. Compositional depth compensates for the constrained per-module expressivity: cascading physically limited nonlinearities yields richer functional representations than increasing parallel width alone. Robustness analyses confirm that optical noise, rather than DAC res- olution, is the dominant hardware impairment, with performance remaining strong down to 6-bit input quantization and SNR=14 dB across all tasks. Multi-layer archi- tectures with intermediate optical summation introduce coherent phase accumulation effects that are analysed in the Supplementary Material. The choice between coherent and incoherent source configurations has a depth- dependent effect on both expressivity and trainability. For the single-layer [2,2] network, incoherent power summation (99.4% accuracy, 16 parameters) substantially outperforms coherent field summation (91.8%, 20 parameters), despite the latter pro- viding additional phase degrees of freedom: the MZI phaseĻsimultaneously shapes the per-edge transfer function and controls the output interference condition, creating a multimodal loss landscape that impedes gradient-based training. For the two-layer [2,2,2] network, this balance reverses: coherent operation (99.9%, 40 parameters) slightly exceeds the incoherent baseline (99.5%, 32 parameters), as the interference cross-terms at intermediate nodes introduce input-dependent effective weights that enrich the networkās functional capacity. This depth-dependent crossover indicates that incoherent summation is preferable for shallow, experimentally accessible con- figurations, while coherent architectures may become advantageous as network depth increases and alternative optimisation strategies are employed. Clear limitations remain. The restricted four-parameter activation family produces a measurable accuracy gap relative to architectures employing resonant or highly engineered nonlinearities, such as the D-RAMZI design of Penget al.[15]. This gap quantifies the cost of prioritising experimental accessibility over per-edge expressivity, and indicates that enhancing activation richness ā through cascaded stages, resonant enhancement, or alternative mechanisms such as stimulated Brillouin scattering or Kerr nonlinearity ā represents the most direct route to closing it while retaining telecom compatibility. Furthermore, while the small-scale configurations studied here (4ā7 modules) are directly realisable on an optical bench, scaling to high-dimensional tasks will require transition to integrated photonic platforms for compactness, thermal stability, and cost. The fully differentiable physics model developed here provides a general framework for co-designing photonic neural network architectures with their physical substrates: any optical platform with analytically tractable transfer functions can in principle be incorporated into this optimisation paradigm. The immediate experimental objective 23 is the four-module [2,2] Two Moons configuration, initialised with simulation-trained parameters, to validate both model fidelity and the end-to-end training framework before scaling to the seven-module yacht hydrodynamics system. Beyond component- level validation, the per-edge SSP-KAN structure maps naturally to wavelength- or space-division multiplexed fibre systems, positioning it for latency-critical applica- tions such as optical channel equalisation where electronic digital signal processing increasingly struggles to keep pace with data rates. Supplementary information.Supplementary Information is available for this paper. It includes supplementary figures showing training evolution of decision boundaries (Supplementary Video 1), parameter sensitivity analysis for individual MZI-VOA-SOA-VOA module parameters (Supplementary Fig. 1). Acknowledgements.This work was supported by the Engineering and Physi- cal Sciences Research Council (EPSRC) under grant EP/Z534614/1, as part of the European Training Network on Post-Digital Computing+ (POSTDIGITAL+) funded through the Marie SkÅodowska-Curie Actions Horizon Europe programme. Declarations ⢠Funding:Engineering and Physical Sciences Research Council (EPSRC), grant EP/Z534614/1 (POSTDIGITAL+, Marie SkÅodowska-Curie Actions Horizon Europe). ⢠Conflict of interest:The authors declare no competing interests. ⢠Data availability:The Two Moons dataset is generated synthetically. The UCI Yacht Hydrodynamics dataset is publicly available [18]. MNIST and Fashion- MNIST are standard benchmarks available through PyTorch. ⢠Code availability:Source code for SSP-KAN, including the differentiable MZI- VOA-SOA-VOA model and all experiment scripts, will be made available on GitHub upon publication. Appendix A Architectural limitations of single-layer SSP-KAN A single-layer [2,2] SSP-KAN computes each output as a sum of univariate functions, y j =Ļ 1,j (x 1 ) +Ļ 2,j (x 2 ). This additive structure cannot represent decision bound- aries that depend on interactions between inputs. Two tasks illustrate the limitation (Supplementary Fig.A1). The XOR task assigns opposite labels to diagonally opposed quadrants, requiring a boundary that depends jointly on both features. The [2,2] network achieves only 51.8% accuracy, indistinguishable from random chance, because no sum of univariate functions can partition the plane into four quadrants. Adding a second layer recovers 89.4%: the [2,2,2] network composes first-layer outputs to approximate the diago- nal structure, though the task remains challenging owing to the sharp axis-aligned boundaries that smooth optical transfer functions approximate imperfectly. 24 A subtler failure arises when the Two Moons dataset is rotated by 45°. In its standard orientation the crescent boundary is approximately separable along the coor- dinate axes, and the [2,2] network reaches 98.4% (Section2.2). After rotation the boundary runs diagonally, breaking this separability: the [2,2] network drops to 89.0%, close to the linear baseline. The [2,2,2] network restores 99.9% by composing first- layer features to reconstruct the rotated separation. This pair of results, 98.4% at 0° versus 89.0% at 45°, directly quantifies the cost of the additive constraint and moti- vates multi-layer architectures, which in turn raise the coherent phase accumulation effects analysed in Supplementary SectionB. Fig. A1Tasks requiring feature interaction. Left column: single-layer [2,2]; right column: two-layer [2,2,2]. Top row: XOR. The [2,2] boundary is at chance (51.8%) while [2,2,2] recovers 89.4%. Bottom row: Two Moons rotated 45°. The [2,2] network (89.0%) falls to the linear baseline because the additive structure cannot represent diagonal boundaries, while [2,2,2] reaches 99.9%. 25 Appendix B Coherent versus incoherent operation Throughout the main text, module outputs combine as optical powers, which is valid when the recombining fields are mutually incoherent. If instead all paths originate from a single laser, the fields retain a fixed phase relationship and combine coher- ently, producing interference terms that alter the network computation. This section derives the incoherent and coherent summation regimes and compares them on the Two Moons benchmark. At each output or intermediate nodej, then in modules routed to that node contribute fields that recombine through a coupler tree. In the most general form, the detected power is y j = 1 ā n in n in X i=1 c ij E f ij ,out 2 ,(B1) wherec ij is the phase factor accumulated by the field travelling from inputito outputjthrough the coupler network. Expanding the square producesn in direct terms andn in (n in ā1)/2cross-terms, each depending on the relative phase between a pair of modules. The two regimes correspond to whether these cross-terms survive or vanish. B.1 Incoherent regime When the combining fields originate from separate lasers, or when path-length mis- match exceeds the coherence length, the cross-terms average to zero and only powers add. For a [2,2] layer the outputs reduce to y 1 = 1 2 f 1 (x 1 P 0 /4) +f 3 (x 2 P 0 /4) , (B2) y 2 = 1 2 f 2 (x 1 P 0 /4) +f 4 (x 2 P 0 /4) , (B3) wheref k is the power-domain transfer function of modulekand the factors1/4and 1/2arise from successive 50/50 couplers. The effective weight on every edge is fixed at1/2; only the four module parameters(I,α 1 ,α 2 ,Ļ)per edge are trainable. This is the regime assumed throughout the main text (16 trainable parameters for [2,2]). B.2 Coherent regime When all signals share the same source, the coupler phase factorsc ij in Eq. (B1) must be tracked. Each lossless 50/50 coupler follows the transfer matrix defined in Eq. (11): C= 1 ā 2 1i i1 ,(B4) where the bar port transmits unchanged and the cross port acquires a factori. A field traversingn Ć cross ports through a cascade of couplers therefore accumulates a phase factori n Ć . For a [2,2] layer fed by a coherent sourceE 0 = ā P 0 , inputs are encoded via amplitude modulators (Eā ā xE) and each arm is split once more to address two modules. Each field traverses two couplers: module 1 takes two bar ports (phase factor 26 i 0 = +1), modules 2 and 3 each take one cross port (i 1 = +i), and module 4 takes two cross ports (i 2 =ā1). In our simulation we label each path by a binary index kencoding bar (0) or cross (1) at each stage, so that the phase factor reduces to i popcount(k) . The resulting input fields are E f 1 ,in = + ā x 1 E 0 2 ,E f 2 ,in = + i ā x 1 E 0 2 ,(B5) E f 3 ,in = + i ā x 2 E 0 2 ,E f 4 ,in =ā ā x 2 E 0 2 .(B6) The sign pattern(+1,+i,+i,ā1)is fixed by coupler unitarity and enters every downstream interference term. Each module acts on the complex field as E f k ,out =t k |E f k ,in | 2 e iĪø k (|E f k ,in | 2 ) Ā·E f k ,in ,(B7) wheret k is the amplitude transmission andĪø k the total output phase, both set by the SOA gain dynamics and the module parameters. Expanding Eq. (B1) for the [2,2] case (modules 1 and 3 feeding output 1) gives y 1 = 1 2 |E 1 | 2 +|E 3 | 2 + 2|E 1 ||E 3 |sin(Īø 1 āĪø 3 ) .(B8) The cross-term has two consequences. First, becauseĪø k depends on input power through SOA gain saturation, the effective weights at recombination nodes vary with the signal. Second, each moduleās MZI phaseĻ k now controls both the transfer func- tion shape and the interference condition, coupling two roles that are independent in the power-domain formulation. A phase shifterĻ k placed after each module separates these roles: y 1 = 1 2 |E 1 | 2 +|E 3 | 2 + 2|E 1 ||E 3 |sin(Īø 1 āĪø 3 + āĻ 13 ) ,(B9) whereāĻ 13 =Ļ 1 āĻ 3 controls the interference independently of the MZIāVOAā SOAāVOA parameters that shape the activation. This adds one parameter per module (5 per edge versus 4 in the incoherent regime). B.3 Experimental comparison Supplementary Fig.B2compares both regimes on Two Moons classification for [2,2] and [2,2,2] architectures. For the single-layer [2,2] network, the incoherent regime (99.4%, 16 parameters) substantially outperforms the coherent regime (91.8%, 20 parameters). Although coherent operation provides the additionalĻdegrees of freedom, it also couplesĻ k to the interference condition at the output couplers. Since the MZI transfer function already depends onĻ k through a cosine-squared fringe, this coupling multiplies the 27 number of local minima in the loss landscape: the same parameter now sits inside two periodic functions with different periods. Gradient descent is more likely to become trapped, and in this case settles on an inferior boundary. The performance gap reflects the diļ¬iculty of gradient-based optimisation on the resulting multimodal landscape, not a limitation of the coherent model, which is strictly more expressive. Alternative training strategies which handle multimodal objectives more naturally may recover the missing performance. For the two-layer [2,2,2] network, the situation reverses: the coherent regime (99.9%, 40 parameters) slightly exceeds the incoherent regime (99.5%, 32 parame- ters). Here the interference cross-terms at intermediate nodes act as input-dependent effective weights that the incoherent formulation cannot access. The additional Ļparameters allow the optimiser to exploit this extra expressivity without the Ļ-interference coupling penalty dominating. The choice of source coherence therefore affects both expressivity and trainabil- ity, and the balance between them depends on network depth. For the small-scale configurations studied in the main text, the incoherent regime offers a simpler optimi- sation landscape at no cost in performance for [2,2] and [6,1,1]. The coherent regime becomes advantageous only when intermediate recombination nodes are present and the network is deep enough to exploit the additional interference degrees of freedom. 28 Fig. B2Coherent versus incoherent operation on Two Moons classification. Top row: single-layer [2,2]. The incoherent regime (99.4%, 16 parameters) outperforms the coherent regime (91.8%, 20 parameters), indicating that interference at the output couplers complicates optimisation without providing a net benefit at this depth. Bottom row: two-layer [2,2,2]. The coherent regime (99.9%, 40 parameters) slightly exceeds the incoherent regime (99.5%, 32 parameters), as the interference cross-terms at intermediate nodes introduce input-dependent effective weights that increase network expressivity. References [1]Shastri, B.J., Tait, A.N., Lima, T., Pernice, W.H.P., Bhaskaran, H., Wright, C.D., Prucnal, P.R.: Photonics for artificial intelligence and neuromorphic computing. Nature Photonics15(2), 102ā114 (2021)https://doi.org/10.1038/ s41566-020-00754-y 29 [2]Zibar, D., Wymeersch, H., Lyubomirsky, I.: Machine learning under the spotlight. Nature Photonics11(2017)https://doi.org/10.1038/s41566-017-0058-3 [3]Genty, G., Salmela, L., Dudley, J.M., Brunner, D., Kokhanovskiy, A., Kobtsev, S.M., Turitsyn, S.K.: Machine learning and applications in ultrafast photonics. Nature Photonics15, 91 (2021) [4]Freire, P.J., Napoli, A., Spinnler, B., Costa, N., Turitsyn, S.K., Prilepsky, J.E.: Neural networks-based equalizers for coherent optical transmission: Caveats and pitfalls. IEEE Journal of Selected Topics in Quantum Electronics28(4: Mach. Learn. in Photon. Commun. and Meas. Syst.), 7600223 (2022)https://doi.org/ 10.1109/JSTQE.2022.3174268 [5]Zibar, D., Piels, M., Jones, R., SchƤeffer, C.G.: Machine learning techniques in optical communication. Journal of Lightwave Technology34(6), 1442ā1452 (2016)https://doi.org/10.1109/JLT.2015.2508502 [6]Khan, F.N., Lu, C., Lau, A.P.T.: Machine learning methods for optical communication systems. In: Advanced Photonics 2017 (IPR, NOMA, Sensors, Networks, SPPCom, PS), p. 2ā3. Optica Publish- ing Group, ??? (2017).https://doi.org/10.1364/SPPCOM.2017.SpW2F.3. http://opg.optica.org/abstract.cfm?URI=SPPCom-2017-SpW2F.3 [7]Freire, P., Manuylovich, E., Prilepsky, J.E., Turitsyn, S.K.: Artificial neural net- works for photonic applicationsāfrom algorithms to implementation: tutorial. Adv. Opt. Photon.15(3), 739ā834 (2023)https://doi.org/10.1364/AOP.484119 [8]Piccinotti, D., MacDonald, K.F., Gregory, S.A., Youngs, I., Zheludev, N.I.: Arti- ficial intelligence for photonics and photonic materials. Reports on Progress in Physics84(1), 012401 (2020)https://doi.org/10.1088/1361-6633/abb4c7 [9]Abreu, S., Boikov, I., Goldmann, M., Jonuzi, T., Lupo, A., Masaad, S., Nguyen, L., Picco, E., Pourcel, G., Skalli, A., Talandier, L., Vettelschoss, B., Vlieg, E.A., Argyris, A., Bienstman, P., Brunner, D., Dambre, J., Daudet, L., Domenech, J.D., Fischer, I., Horst, F., Massar, S., Mirasso, C.R., Offrein, B.J., Rossi, A., Soriano, M.C., Sygletos, S., Turitsyn, S.K.: A photonics perspective on computing with physical substrates. Reviews in Physics12, 100093 (2024)https://doi.org/ 10.1016/j.revip.2024.100093 [10]Wetzstein, G., Ozcan, A., Gigan, S., Fan, S., Englund, D., SoljaÄiÄ, M., Denz, C., Miller, D.A.B., Psaltis, D.: Inference in artificial intelligence with deep optics and photonics. Nature588(7836), 39ā47 (2020)https://doi.org/10.1038/ s41586-020-2973-6 [11]McMahon, P.L.: The physics of optical computing. Nature Reviews Physics5(12), 717ā734 (2023) 30 [12]Ashtiani, F., Geers, A.J., Aflatouni, F.: An on-chip photonic deep neural network for image classification. Nature606(7914), 501ā506 (2022)https://doi.org/10. 1038/s41586-022-04714-0 [13]Chen, C., Xu, Z., Liu, Y., Wu, Q., Ji, T., Ji, H., Tang, J., Sun, Z., Fan, L., Liang, J.,et al.: Kolmogorov-arnold network for eļ¬icient equalization in short-reach im/d systems. Optics Express33(16), 33139ā33152 (2025) [14]Fischer, R., Matalla, P., Randel, S., Schmalen, L.: Non-linear equalization in 112 Gb/s PONs using KolmogorovāArnold networks (2024).https://doi.org/10. 48550/arXiv.2411.19631 [15]Peng, Y., Hooten, S., Yu, X., Van Vaerenbergh, T., Yuan, Y., Xiao, X., Tossoun, B., Cheung, S., Fiorentino, M., Beausoleil, R.: Photonic KAN: a Kolmogorovā Arnold network inspired eļ¬icient photonic neuromorphic architecture (2024). https://doi.org/10.48550/arXiv.2408.08407 [16]Stroev, N., Berloff, N.G.: Programmablek-local Ising machines and all-optical KolmogorovāArnold networks on photonic platforms (2025).https://doi.org/10. 48550/arXiv.2508.17440 [17]Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., SoljaÄiÄ, M., Hou, T.Y., Tegmark, M.: KAN: KolmogorovāArnold Networks (2025).https://doi.org/10. 48550/arXiv.2404.19756 [18]Dua, D., Graff, C.: UCI Machine Learning Repository.https://archive.ics.uci. edu/ml(2017) [19]Gerritsma, J., Onnink, R., Versluis, A.: Geometry, resistance and stability of the Delft systematic yacht hull series. International Shipbuilding Progress28(328), 276ā297 (1981)https://doi.org/10.3233/ISP-1981-2832801 [20]Sobhanan, A., Anthur, A., OāDuill, S., Pelusi, M., Namiki, S., Barry, L., Venkitesh, D., Agrawal, G.P.: Semiconductor optical amplifiers: recent advances and applications. Advances in Optics and Photonics14(3), 571 (2022)https: //doi.org/10.1364/AOP.451872 [21]Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: Proceedings of the International Conference on Learning Representations (ICLR) (2019) [22]Agrawal, G.P., Olsson, N.A.: Self-phase modulation and spectral broadening of optical pulses in semiconductor laser amplifiers. IEEE Journal of Quantum Electronics25(11), 2297ā2306 (1989)https://doi.org/10.1109/3.42059. Accessed 2025-11-18 31