Paper deep dive
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
Solomon Micheal Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat Taha, Muhammad Shafique, Mahmoud Rasras
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expensive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifically, MDTransformer performs complex matrix operations using spatial-mode interference, that leverages the inverse-designed multi-mode couplers, crossings, and Mach-Zehnder IQ modulators into a compact mode-division photonic tensor core (MPTC), capable of executing matrix multiplications in the optical domain. Its each guided mode (i.e., TE0-TE3) acts as an independent computational lane, enabling four-fold parallelism-per-waveguide without spectral filtering or free-spectral-range limitations. Moreover, its coherent detection and IQ modulation jointly encode amplitude and phase, realizing complex-valued arithmetic for full-range operations in transformers. MDTransformer offers analog multiplication with sub-4-bit effective precision and inter-modal crosstalk below -30 dB. Its inverse-designed approach also offers scalable and full compatibility with single-laser continuous-wave operation at 1550 nm. Experimental results show that MDTransformer achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT-Tiny/Small/Base and BERT-Base/Large). These results show that MDTransformer offers a practical solution for high-performance and energy-efficient transformer-based systems.
Tags
Links
- Source: https://arxiv.org/abs/2607.26016v1
- Canonical: https://arxiv.org/abs/2607.26016v1
Trouble viewing inline? Open PDF directly →
Full Text
53,203 characters extracted from source content.
Expand or collapse full text
1 MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar Solomon Micheal Serunjogi ∗ , Rachmad Vidya Wicaksana Putra ∗ , Member, IEEE, Ayat Taha, Muhammad Shafique, Senior Member, IEEE, and Mahmoud Rasras, Senior Member, IEEE Abstract—Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy effi- ciency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expen- sive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifi- cally, MDTransformer performs complex matrix operations using spatial-mode interference, that leverages the inverse-designed multi-mode couplers, crossings, and Mach-Zehnder IQ modula- tors into a compact mode-division photonic tensor core (MPTC), capable of executing matrix multiplications in the optical domain. Its each guided mode (i.e., TE 0 -TE 3 ) acts as an independent computational lane, enabling four-fold parallelism-per-waveguide without spectral filtering or free-spectral-range limitations. More- over, its coherent detection and IQ modulation jointly encode amplitude and phase, realizing complex-valued arithmetic for full-range operations in transformers. MDTransformer offers analog multiplication with sub-4-bit effective precision and inter- modal crosstalk below -30 dB. Its inverse-designed approach also offers scalable and full compatibility with single-laser continuous- wave operation at 1550 nm. Experimental results show that MDTransformer achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT- Tiny/Small/Base and BERT-Base/Large). These results show that MDTransformer offers a practical solution for high-performance and energy-efficient transformer-based systems. Index Terms—Silicon Photonics, Mode-Division Photonic Ac- celerator, Transformers, Hardware-Software Co-Design, Inverse Design, Coherent Crossbar. I. INTRODUCTION Transformer-based networks [1], such as Large Language Models (LLMs) and Vision Transformers (ViTs), have demon- strated state-of-the-art performance (e.g., accuracy) for solving Solomon Micheal Serunjogi and Ayat Taha are with Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: sms10215@nyu.edu, aat9458@nyu.edu). Rachmad Vidya Wicaksana Putra is with eBRAIN Lab, Division of En- gineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: rachmad.putra@nyu.edu). Muhammad Shafique is the Director of eBRAIN Lab, Division of Engineer- ing, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: muhammad.shafique@nyu.edu). Mahmoud Rasras is the Director of Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates (UAE); (e-mail: mrasras@nyu.edu). ∗ Equal contributions. 78 80 82 84 86 88 050100150200250300 Top -1 Accuracy [%] Number of Parameters [M] ImageNet-1K (a) accuracy improves as the model size increases (b) Mod Mod DDot DDot DDot Mod Mod DDot DDot DDot Mod DDot DDot DDot λ 0 λ 1 λ 2 λ 0 λ 1 λ 2 Mod Fig. 1. (a) Transformer networks typically improve their performance at the cost of larger memory footprint; based on data from [6]. (b) The state-of-the- art photonic tensor core for Transformer acceleration based on dynamically- operated dot-product (DDot) unit from the LT accelerator [18]. diverse machine learning (ML) tasks, e.g., natural language processing (NLP) and computer vision [2]–[4], thereby paving the way toward artificial general intelligence (AGI) [5]. This state-of-the-art performance comes at higher computational and memory costs as shown in Fig. 1(a), thereby leading to huge power/energy consumption [6]. This condition limits the wide adoption of Transformer models in diverse application use-cases. Toward this, specialized electronic accelerators for Transformer inference have been developed [7]–[9]. However, such conventional accelerators face challenges as transistor circuits hit the limits of Dennard scaling [10], leading to their diminishing return of performance efficiency (e.g., slower performance gains and increased power dissipation-per-unit area). Recent works have proposed optical-based integrated cir- cuits to expedite neural network (N) inference, exploiting ultra-high speed and low energy nature of optical-based com- putation, known as photonic accelerators [11]–[13]. They typically leverage optical components such as Micro-Ring Resonator (MRR) [14] [15], Mach-Zehnder Interferometer (MZI) [16], and Phase Change Material (PCM) [17] for designing a photonic tensor core (PTC). However, these works mainly target convolutional neural networks (CNN) acceleration [18], exposing the need for studies that target transformer acceleration. Therefore, the targeted problem in this work is how can we develop a high performance and energy-efficient photonic transformer accelerator (PTA)? A solution to this problem may enable a practical PTA design for diverse application use-cases. A. State-of-the-Art PTAs and Their Limitations Most of PTA designs employ PTC based on MZI, MRR banks [19]–[23], and PCM crossbars [24]. They statically store operands on the optical components for computation, hence arXiv:2607.26016v1 [cs.AR] 28 Jul 2026 2 LaserDAC MZMADC TIACore AdderMemory Micro_comb (a) (b) 12.6% 12.1% 20.1% 18.8% 14% 12.4% 14.8W excluding Micro_comb 28.1W excluding Micro_comb 24.4% 2.7% 1.2% 26.3% 1.1% 25.3% 2.9% 26% Fig. 2. Area breakdown of the state-of-the-art 4-bit LT accelerators: (a) LT- Base and (b) LT-Large, showing the contributions of different modules. they suffer from slow operand mapping and programming. Re- cently, the Lightening-Transformer (LT) accelerator [18] has been proposed. It inspires further studies in reconfigurability aspect [25] and digital-to-analog converter (DAC) optimiza- tion [26] [27]. LT improves performance efficiency of Trans- former inference over other PTA designs by employing PTC with dynamic operations of full-range input operands through dynamically-operated dot-product (DDot) unit; see Fig. 1(b). Hence, it eliminates slow operand mapping/programming and making it the state-of-the-art PTA design. Despite their bene- fits, all these works still have the following critical limitations. • They typically employ multiple wavelengths for opera- tions based on wavelength-division multiplexing (WDM), to achieve highly parallel multiply-accumulate (MAC) op- erations [28] [29]. However, their reliance on finely-spaced resonant filters and dispersion-limited channels imposes scalability bottlenecks, as generating multiple wavelengths consumes huge area and power/energy and is often done in a strongly nonlinear medium such as SiN. • The free-spectral-range (FSR) of MRRs restricts the number of usable wavelengths, while temperature-dependent reso- nance drift and fabrication non-uniformity demand active thermal control and calibration overhead [30], [31]. • Coherent operation across many wavelength lanes requires precise optical phase alignment and stabilization, thereby adding power and system complexity for a practical solu- tion [32] [33]. • State-of-the-art accelerators relies on Micro comb (MC)- based wavelength generator, Mach-Zehnder Modulator (MZM), and phase-shifter (PS)-based PTC, which are area and power hungry. Therefore, they impose scalability and ef- ficiency challenges when designing area- and power-efficient PTA architecture. To show the limitations of state-of-the-art and related research challenges, we conduct a case study in Section I-B. B. Case Study and Related Research Challenges We study the impact of different modules in the state-of- the-art 4-bit LT accelerators (i.e., LT-Base and LT-Large) [18] on area and power consumption using its open-source codes from the original authors. The experimental results are shown in Fig. 2, from which we draw the following key observations. • Micro comb, MZM, and PS-based PTC jointly occupy 45.5% area in LT-Base and 44.6% area in LT-Large, high- lighting their dominant area consumption. • Both LT accelerators incur high power consumption, even without Micro com. It is particularly inefficient for meeting diverse possible power-constrained computing systems. For instance, embedded AI systems typically require about 5W max. power envelope, which is difficult to meet with the existing solutions. • Employing a Micro comb module, which resides in a separate physical chip, comes with additional complexity and non-trivial challenges since it requires a highly precise control in different aspects, including optical stabilization and power distribution. These observations expose the following key research chal- lenges to address for providing a practical PTA design solution. • The use of Micro comb should be avoided to minimize design complexity and inter-chip communication challenges. Hence, its laser generation functionality should be replaced with an efficient alternative solution. • The area of laser generator, modulator, and PTC should be optimized to significantly reduce area, and hence minimiz- ing power and energy consumption. • PTA architecture and dataflow should be synergistically designed to exploit to maximize the performance and ef- ficiency benefits offered by the optical-based processing. C. Our Novel Contributions To address the targeted problem and related challenges, we propose MDTransformer, a novel hardware-software (HW- SW) co-design of photonic transformer accelerator (PTA) that leverages the spatial Mode-Division Multiplexing (MDM), inverse-designed coherent crossbar, and IQ modulation to enable a practical solution for high-performance and energy- efficient transformer-based systems. It is also the first work that leverages MDM and inverse design concepts for designing PTA. It employs the following key ideas. • Leveraging Spatial MDM for Photonic Computing (Sec- tion I-A). It employs MDM to enable on-chip photonic parallelism by employing efficient Mode-based Multiplexer or Demultiplexer (MUX/DEMUX), which distributes, en- codes, and routes a single optical source across multiple guided spatial channels. • MDOT: Mode-Division Dot-Product Unit (Section I-B). It aims to perform dot-product operation that represents multiplication of two full-range operands in multiple modes at sigle frequencies, thereby enabling dynamically-operated processing element (PE) for higher-level architecture hier- archy (i.e., PTC). • MPTC: Mode-Division Photonic Tensor Core (Sec- tion I-C). It aims to efficiently accelerate general matrix multiplication (GEMM) by leveraging MDOT, Mode-based MUX/DEMUX, IQ modulator, and crossings in a crossbar array fashion. • Architecture System Design (Section I-C). It aims to develop the architecture system of MDTransformer acceler- ator, by integrating multiple MPTCs and the supporting digi- tal circuits (e.g., on-chip memory). Furthermore, a dataflow pattern is also developed to maximize the benefits of the MDTransformer architecture. Key Results: We evaluate MDTransformer through func- tional simulation using Tidy3D [34], as well as hardware eval- 3 Fig. 3. Our proposed MDTransformer Accelerator: (a) MPTC with crossings, modulator, and MDOT; (b) architecture design. uation (e.g., area, power, energy, and latency) using the state- of-the-art PTA hardware simulator from [18]. Furthermore, our design is also under fabrication. Experimental results show that, MDTransformer offers 4-bit effective precision for multiplication, low inter-modal crosstalk (i.e., below -30 dB), and full compatibility with single-laser continuous-wave operation at 1550nm. MDTransformer also achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT-T/S/B 1 and BERT-B/L 2 ). I. PRELIMINARIES A. Transformer-based Network Models A transformer-based network is formed by multiple identical blocks: encoder and decoder blocks. Each block is formed by a multi-head self-attention (MA) module, a feed-forward network (FN) module, a layer normalization (LN) module, and shortcut connections [18]. The decoder block also has cross-attention and masked self-attention modules. The basic encoder block can be stated as Eq. 1-2, where X l is the input sequences of l-th layer. Multi-head self-attention (MA) module supports H self-attention heads, and each head makes the input vector into separate vectors, i.e., query (Q), key (K), and value (V) vectors. The attention function between these input vectors can be state as Eq. 3, where d k is the dimension of Q and K. X l+1 = FN(LN( ˆ X l+1 )) + ˆ X l+1 (1) ˆ X l+1 = MA(LN(X l )) + X l (2) Atten(Q, K, V) = softmax QK ⊺ √ d k V(3) B. Inverse Design in Photonic Circuits Inverse-designed components can implement complex trans- formations, such as filters and couplers, within an order of magnitude smaller than classical designs [35]–[37]. Such 1 DeiT-T/S/B denotes DeiT-Tiny, DeiT-Small, and DeiT-Base, respectively. 2 BERT-B/L denotes BERT-Base and BERT-Large, respectively. devices maintain low loss and high modal fidelity, making them ideally suited for large-scale photonic accelerators here thousands of operations must be packed into a small area. As photonics moves toward ultra-dense, domain-specific optical computing, inverse design has become a promising approach for building high-performance primitives that enable massive parallelism and energy-efficient linear algebra. I. OUR PROPOSED MDTransformer ACCELERATOR We develop our PTA design, called MDTransformer, based on MDM, inverse-designed coherent crossbar, and IQ modu- lation (overview in Fig. 3). Details of its design is discussed in Section I-A - Section I-D. A. Leveraging Spatial Mode-Division Multiplexing for Pho- tonic Computing We leverage spatial MDM to obtain an orthogonal degree of freedom for on-chip photonic parallelism [38], [39], thereby enabling light generation using limited number of on-chip lasers and reducing on the number of WDM channels. Instead of distributing computation across distinct wavelength like in the state-of-the-art works [18], [26], [27], MDM leverages multiple guided modes (e.g., TE 0 –TE 3 ) within a single multi- mode waveguide, enabling independent and simultaneous in- formation channels in the spatial domain. Here, operand pairs are mapped onto orthogonal modal channels (e.g., TE 0 –TE 3 ) of a multi-mode bus waveguide and processed in modular MDOT units. B. MDOT: Mode-Division Dot-Product Unit We propose a Mode-Division Dot-Product Unit (MDOT) to perform a signed multiplication between two full-range mode- encoded operands through optical interference; see Fig. 4. Fig. 4(a) shows the novel inverse-designed structure after optimization, which ensures that the coherent coupler occupies a small footprint and does not need an area- and power- hungry 90-degree phase-shifter. Meanwhile, Fig. 4(b) shows the intensity across the MDOT structure. Here, each spatial mode is routed into a compact 8× 8 μm 2 inverse-designed 4 x (μm) y (μm) E 2 (b) 6.00 4.00 2.00 0.00 -2.00 -4.00 -6.00 -6.00 -4.00 -2.00 0.00 2.00 4.00 6.00 8000 7000 6000 5000 4000 3000 2000 1000 x (μm) y (μm) (a) 6.00 4.00 2.00 0.00 -2.00 -4.00 -6.00 -6.00 -4.00 -2.00 0.00 2.00 4.00 6.00 12 10 8 6 4 2 휀 푟 Fig. 4. Inverse-designed MDOT design: (a) silicon projection of the design region. (b) Intensity distribution of light traveling through the structure from the left input. (a)(b) Fig. 5. MDOT properties: (a) simulated relative phase differences between the four output ports; and (b) CMRR. coherent mixer that produces four output interference states with fixed phase relationships. Each of the four mixer outputs can be described using a linear transformation of the two in- coming operands s A and s B ; see Eq. 4. It explicitly shows the 0 ◦ , 90 ◦ , 180 ◦ , −90 ◦ phase basis generated at the outputs. [E 1 E 2 E 3 E 4 ] ⊺ = 1 2 [(s A + js B ) (s A − js B ) (s A − s B ) (s A + s B )] ⊺ (4) Figure 5(a) shows the simulated relative phase differences between all four output ports, taken pairwise, when a single mode is launched from the input (left side). Adjacent port pairs (port 0–port 1 and port 2–port 3) maintain approximately 180 ◦ phase separation across the 1530–1560 nm band, while alternating port pairs maintain roughly 90 ◦ separation. The figure also shows a near zero port imbalance across the wavelength of interest as shown by the lightly shaded green region. Meanwhile, the corresponding common-mode rejection ratio (CMRR) is shown in Fig. 5(b), exceeding 30 dB at the operating wavelength of λ = 1550 nm. Bipolar Encoding and Signed Multiplication: The MDOT unit accepts two bipolar NRZ symbol streams s (m) x (k) and s (m) y (k) ∈ −1, +1 derived from input bits a k ,b k via s = 1− 2a. The fields applied to the coherent mixer are: E (m) x (k) = s (m) x e jφ (m) x , E (m) y (k) = s (m) y e jφ (m) y ,(5) with φ x = πp k and φ y = πq k and p k ,q k ∈ 0, 1 applied through a phase modulator. The balanced detection photocur- rent for mode m over N symbol periods yields the mode-wise dot product: I (m) PC = C Z Nτ 0 s (m) x (t)s (m) y (t) cos ∆φ m (t) dt ∝x m ·y m , (6) wherex m andy m denote the symbol vectors carried by mode m and C is an arbitrary constant. Summing over all supported spatial modes produces the full mode-division dot product: I (m) PC ∝x m ·y m , I PC = M−1 X m=0 I (m) PC ∝ M−1 X m=0 x m ·y m . (7) Therefore, the coherent MDOT unit directly performs signed multiplication and accumulation using both optical (high Q resonators) and electrical domain (capacitive dynamics) through time multiplexed integrators [40]–[42]. C. MPTC: Mode-Division Photonic Tensor Core We propose a novel photonic tensor-core architecture, re- ferred to as the Mode-Division Photonic Tensor Core (MPTC), for efficient general matrix multiplication (GEMM) using a single optical carrier and multiple orthogonal spatial modes. The MPTC combines four principal building blocks: (i) a mode-based multiplexer/demultiplexer (MUX/DEMUX), (i) the coherent mode-division dot-product unit (MDOT), (i) IQ- based complex modulation, and (iv) compact routing elements such as crossings and couplers, all assembled into a structured array, as shown in Fig. 3(a). Unlike prior photonic matrix engines that rely on wavelength-division parallelism or cascaded interferometric meshes, the proposed MPTC uses spatial modes as the pri- mary computational lanes. This choice reduces dependence on multiple laser wavelengths, resonance management, and spectral routing overhead, while enabling multiple operands to propagate within the same multimode waveguide. 1) Mode-based Multiplexer and Demultiplexer (MUX and DMUX): The MUX/DEMUX is the front-end modal interface of the MPTC and is responsible for converting single-mode input channels into a multimode computational bus, and con- versely extracting specific modes at later processing stages. In contrast to communication-only mode multiplexers, which are typically designed as standalone coupling elements, the MUX/DEMUX here is designed as a computational routing primitive whose role is to inject and recover operands inside a dense coherent dot-product array. Fig. 6 shows the operation of the four-mode MUX/DEMUX, designed using full-wave inverse design in Tidy3D. The fig- ure illustrates the field evolution for each input mode and the corresponding selective routing to the designated output port. Each panel should be interpreted as a mode-resolved demonstration of selective field transformation: for each input channel, the optical energy is redistributed so that only the target output port carries the desired mode, while leakage to the other ports is suppressed. The input interface consists of four single-mode waveguides of width 0.5 μm. This width is selected to ensure robust TE 0 operation at λ = 1550 nm, thereby providing a clean modal input state before multiplexing. The spacing between adjacent input waveguides is set to 1.5 μm. This spacing is large enough to suppress unwanted evanescent coupling 5 x (μm) 7.50 5.00 2.50 0.00 -2.50 -5.00 -7.50 y (μm) -5.00 -2.50 0.00 2.50 5.00 x (μm) -5.00 -2.50 0.00 2.50 5.00 (a) (b) x (μm) -5.00 -2.50 0.00 2.50 5.00 (c) x (μm) -5.00 -2.50 0.00 2.50 5.00 20 10 0 -10 -20 푅 퐸 (d) Fig. 6. Inverse-designed mode-based MUX/DEMUX: optical field distributions for the four input modes, showing selective routing to distinct single-mode outputs. Each panel illustrates a mode-resolved input-to-output field transformation rather than simple power splitting. between neighboring inputs, yet small enough to maintain dense layout compatibility with the surrounding tensor-core routing network. These four single-mode channels feed a multimode bus waveguide of width 2.5 μm. This width is chosen because it provides a practical trade-off between modal capacity and circuit density: it is sufficiently wide to support the first four guided TE modes (TE 0 –TE 3 ) at 1550 nm, while remaining narrow enough to avoid excessive crossing area, large bending penalties, and poor array density. In other words, the selected dimensions are not arbitrary; they arise from the joint require- ment of supporting four orthogonal modes and embedding them in a compact computational crossbar. To obtain the optimized freeform structure, the permittivity distribution ε(r) is solved through an adjoint-based gradient descent formulation that maximizes the transmission of each target mode into its assigned output port while penalizing leakage into all other ports. The optimization problem is expressed in Eq. 8, where T m→p m denotes the transmission from input mode m to its designated output port p m , and α is a penalty factor enforcing crosstalk suppression. This objective does not merely maximize throughput; it imposes a mode- selective field transformation that preserves the computational meaning of each channel. max ε(r) F = 3 X m=0 T m→p m − α 3 X n=0 n̸=m T m→p n (8) Physically, the optimized region acts as a compact dis- tributed scattering medium that directly maps one modal basis to another. This is an important distinction from conventional asymmetric directional couplers, microring assisted mode cou- plers, or MMI-based devices, where coupling is governed by predetermined geometric interference lengths and is often less flexible for simultaneously enforcing compactness, broad- band operation, and multi-port modal selectivity. Earlier on- chip mode-division multiplexers, such as microring-assisted designs, demonstrated selective mode coupling but remained tied to wavelength-sensitive routing concepts. Similarly, ultra- compact multimode routing work has focused on bends and crossings for dense integration. Here, by contrast, the inverse- designed MUX/DEMUX is integrated directly into a photonic tensor-core data path, where its purpose is not only multi- plexing, but controlled operand delivery to coherent compute nodes [39] [43]–[45]. Fabrication-Aware Inverse Design and Constraints: To ensure practical manufacturability, the inverse design process incorporates fabrication-aware constraints consistent with stan- dard electron-beam lithography in silicon photonics. Specifi- cally, a minimum feature size of 120 nm is enforced through spatial filtering and projection steps applied to the permittivity distribution during optimization. This avoids the formation of sub-resolution features and ensures that the final structure can be faithfully fabricated without requiring additional post- processing. Such feature-size-constrained inverse design has been widely adopted in recent nanophotonic devices to bridge the gap between idealized continuous permittivity optimization and binary fabrication-compatible layouts. Modal Superposition and Decomposition Strategy: Al- though the device operates on four orthogonal modes si- multaneously, the optimization is structured using a modal decomposition approach. Each mode transformation is treated as an independent objective, and the total cost function is constructed as a superposition of these modal targets, as shown in Eq. 8. This ensures that each input mode is selectively mapped to its corresponding output port without interfering with the routing of the remaining modes. This decomposition is physically justified by the orthogonality of the guided modes in the multimode waveguide, allowing independent control of each modal channel while maintaining a shared spatial structure. Reciprocity and Forward–Adjoint Consistency: The op- timization leverages electromagnetic reciprocity, whereby the adjoint simulation corresponds to exciting the device from the output ports and propagating fields backward. Consistency between forward and adjoint field distributions ensures that the optimized structure satisfies both excitation and collection conditions simultaneously. In practice, this guarantees that the device performs equivalently under forward multiplexing and reverse demultiplexing operation, which is essential for its dual role within the MPTC architecture. Output Mode Engineering and Power Capture Effi- ciency: In order to improve power transfer from the multimode region into the output waveguides, the single-mode output 6 6000 5000 4000 3000 2000 1000 30 20 10 0.0 -10 -20 -30 E 2 Re Ez (a)(b) (c) Fig. 7. Cross-sectional optical field distributions for (a) scissors crossing, (b) 50:50 3dB coupler, and (c) 90 ◦ waveguide crossing. ports are intentionally widened beyond the nominal 0.5 μm width. Specifically, the outputs are expanded to approximately 0.75 μm before being adiabatically tapered back to standard single-mode dimensions. This local widening improves mode overlap between the transformed field distribution and the guided mode of the output waveguide, thereby enhancing coupling efficiency and reducing scattering loss. Such taper- assisted mode matching is critical in inverse-designed struc- tures, where the output field profile may not perfectly match the fundamental mode of a narrow waveguide without addi- tional impedance matching. The field distributions in Fig. 6 also clarify why the inverse- designed approach is needed. For each launched mode, the structure does not simply split power; it redistributes phase and amplitude across a freeform subwavelength region so that the desired output field emerges at one specific port while the remaining ports are suppressed. This mode-resolved routing behavior is exactly what is required in the MPTC: at each downstream computational cell, one selected mode must be exposed to the coherent multiplier, while the remaining modes must continue propagating with minimal disturbance. In benchmarking terms, recent inverse-designed mode- division devices have demonstrated that compact mode mul- tiplexers can significantly outperform conventional mode- routing footprints, for example through five-mode inverse- designed MDM devices with a reported footprint of 16×7 μm 2 and measured crosstalk below approximately −11 dB, as well as recent scalable mode demultiplexers with sub-1 dB loss and crosstalk below approximately −13 dB at 1550 nm. Dense multimode routing elements such as 8 × 8 μm 2 crossings have also been reported for three-mode photonic circuits. Our design inherits the compactness philosophy of these works but targets a different system problem: rather than building a com- munication link, the present MUX/DEMUX is dimensioned and optimized as the operand-injection and mode-selection interface for a coherent photonic tensor core [39] [45] [46]. Additional supporting components, including the scissors crossing, 50:50 3 dB coupler, and 90 ◦ waveguide crossing, are presented in Fig. 7(a)–(c), respectively. These elements are used to construct the routing network surrounding the MPTC. The processing pipeline begins with a continuous-wave input field that is first split using integrated power splitters and 50:50 couplers. These peripheral components also form the basis of other circuit blocks such as IQ modulators and mode-scissors crossings, enabling flexible routing of optical data throughout the processor. After splitting, the optical field is expanded into the mul- timode bus waveguide and encoded into one of the four orthogonal TE modes used by the MD-Transformer. The inverse-designed MUX/DEMUX then maps each input modal profile onto a unique single-mode output port at 1550 nm. This provides clean modal separation, low inter-mode crosstalk, and a compact footprint suitable for dense dot-product arrays. Overall, the mode-based MUX/DEMUX forms the front- end interface of the MD-Transformer, enabling parallel spatial- mode encoding, selective demultiplexing, and physically struc- tured delivery of operands into the downstream coherent processing core. 2) IQ-Based Complex Modulator: To enable full complex- valued encoding of optical operands, each modal channel incorporates a compact IQ-modulator. Two complementary devices are used: (i) a single-input intensity I modulator for amplitude control, and (i) Q modulator for phase control. Each input channel employs an IQ modulator analogous to that in [47], but modified to operate at high speeds of 25Gb/s. The amplitude branch is defined by the normalized MZI power transfer in Eq. 9, and the quadrature branch applies an addi- tional phase-shift as in Eq. 10. Here, A,B,D,E are empirical calibration coefficients from device-level measurements. P MZI (I A ) = 1 2 + 1 2 cos AI 2 A + BI A + φ A ,(9) φ PS (I Q ) = DI 2 Q + EI Q + φ Q ,(10) Combining these responses, the complex field at the modulator output for mode m is defined as: E (m) x (I A ,I Q ) =|E (m) 0 | p P MZI (I A ) exp iφ PS (I Q ) . (11) D. Architecture System Design 1) Overall System: The architectural system of our pro- posed MDTransformer accelerator is shown in Fig. 3(b). A 7 ×M 1 = M 2 Tile 0 Tile 1 D v D h (a) ... ... × × × × ... ... Σ MPTC ... MPTC MPTCMPTC Tile 0 ADC Accum. Buffer out 0 out 1 (b) × crossing mode 3 mode 2 mode 1 mode 0 x 0 x 1 x 2 x 3 mode 3 mode 2 mode 1 mode 0 y 2 y 0 y 1 y 3 mode 3 mode 2 mode 1 mode 0 mode 0 mode 3 mode 2 mode 1 x 0 . y 0 MDOT x 1 . y 1 time ... B A a portion of data Fig. 8. Dataflow based on data tiling mechanism for MDTransformer. (a) It partitions data from M 1 across D v , and map them across tiles. (b) Its crossing routes operands based on their mode to the corresponding MDOT. single MDTransformer chip has N t tiles, and each tile consists of N c MPTCs. An MPTC contains an array of N h ×N v MDOTs. Furthermore, MDTransformer also employs on-chip global SRAM whose size should be at least meeting the minimum required size for storing the largest activations in a layer; following the LT design [18]. The global SRAM size should not be significantly smaller or larger than this minimum required size, because it can increase the costly off-chip data access (i.e., high access latency and energy) or aggravate the static power consumption, respectively [48]–[50]. 2) Dataflow: To maximize benefits of the MDTransformer architecture, a specialized dataflow is developed. Its key ideas are illustrated in Fig. 8 and described below. • Multiple data is processed in the same MDOT without any prior data programming; see Fig. 8(a). Multiple tiles can process multiple portions of data, which determines the parallelism level in a chip; see A. Then, multiple cores (MPTCs) can process a portion of data, which also determines the parallelism level in a tile; see B. Afterward, an N h ×N v MDOT array can perform multiplications in parallel in the core level. • Each MDOT performs multiplication between two operands from the same mode. Hence, a sequence of multiplications can be scheduled to be performed in the same MDOT, en- abling flexible scheduling for exploiting data reuse without expensive broadcast routing; see Fig. 8(b). IV. EVALUATION METHODOLOGY To evaluate our MDTransformer design, we employ: (1) functional simulation using Tidy3D [34], and (2) hardware evaluation using the state-of-the-art PTA hardware simulator from [18]; see Fig. 9(a). We use functional simulation to evaluate the functionality and characteristics of our proposed optical devices and circuits. The corresponding results are mainly presented in Section I to validate the functionality of MDTransformer. Meanwhile, we use PTA hardware simulator aims to evaluate area and power of the design as well as its energy consumption and latency when running the workload, while considering device parameters from measurements; see Table I. We select DeiT-T, DeiT-S, DeiT-B, BERT-B, and BERT-L as the workloads. As comparison partners, we use the state-of-the-art LT-Base, LT-Large, and LT-Custom with the following configurations. • MDTransformer employs N t =4 tiles, N c =2 cores-per-tile, N h =N v =4 input horizontal/vertical waveguides-per-core, N λ =1 wavelength, and 4 modes. • LT-Base employs N t =4, N c =2, and N h =N v =N λ =12. • LT-Large employs N t =8, N c =2, and N h =N v =N λ =12. • LT-Custom employs N t =4, N c =2, and N h =N v =N λ =4. LT-Large employs 4MB global SRAM, while the others use 2MB. Our design is under fabrication and its measurement setup is shown in Fig. 9(b). DC Probes RF Probes Optical Fiber Optical Chip Workloads (i.e., DeiT-T/S/B and BERT-B/L) Photonic Simulator (Tidy3D) (a)(b) MDTransformer Arch. Configuration (e.g., N t , N c , N h , N v ) Performance & Efficiency Results (i.e., area, power, energy, latency) Functionality & Characteristics Results (e.g., CMRR, Imbalance) P TA Hardware Simulator (.py) Measurement Setup Photonic Devices (Modulator, etc.) & Circuits (Crossing, etc.) Fig. 9. (a) Experimental setup and tools flow in this work. (b) Measurement setup for testing the fabricated chip. TABLE I SUMMARY OF DEVICE PARAMETERS DeviceParameterValue DAC [51] Precision8-bit Power42 mW (@28 GSPS) Area0.03 m 2 ADC [52] Precision32 lines, 6-bit Power410 mW (@12.8 GSPS) Area780 μm 2 TIA [53] Power30 mW Area< 50 μm 2 MZM Power50 mW Area∼2,260 μm 2 Crossing IL0.3 dB Area36 μm 2 Phase Shifter IL1 dB Area250 μm 2 Y-Splitter IL0.4 dB Area36 μm 2 4 Mode MUX/DEMUX IL6.2 dB Area64 μm 2 Coherent Hybrid IL6 dB Area144 μm 2 Photodetector Power1.1 mW Sensitivity-25 dBm Area4×10 μm 2 V. RESULTS AND DISCUSSION A. Reduction of Area and Power Consumption Experimental results for area and power consumption are presented in Fig. 10(a)-(d). The results show that MDTrans- 8 0 2 4 6 8 0 4 8 12 16 EmbedQKVAttenProjFFN1FFN2HeadOthersLatency 0 20 40 60 80 100 120 Micro_comb Memory Adder Core TIA ADC Modulator DAC Laser 0 5 10 15 20 25 30 Thousands (a)(b) Laser DAC Modulator ADC TIA PD Adder Memory (c) 88.4% 4.3%0.1% 2.6% 1.1% 2.8% 0.5% 16.6 m 2 697.9 mW 12.8% 0.6% 0.9% 5.9% 0.4% Area [m 2 ] (d) Area Breakdown Power Breakdown Power [W] LT -Large LT -Base LT -Custom MD Transformer Significant area savings compared to other state-of -the-art designs Significant power savings compared to other state-of -the-art designs 0 0.2 0.4 0.6 0.8 0 0.2 0.4 0.6 0.8 1 0.0 0.5 1.0 1.5 2.0 0 1 2 3 4 0.0 1.5 3.0 4.5 6.0 0 2 4 6 8 10 1234 0 15 30 45 60 0 20 40 60 80 1234 QKVAttenProjFFN1FFN2HeadOthersLatency LT -Large LT -Base LT -Custom MD Transformer Energy [mJ] (e.1) DeiT-T (e.2) DeiT-S (e.3) DeiT-B (e.4) BERT-B Latency [ms] 2 3 6 4 5 6 4 5 6 4 5 6 4 5 (e.5) BERT-L 6 4 5 1 Fig. 10. Experimental results for (a) area breakdown of MDTransformer; (b) power breakdown of MDTransformer; (c) comparison on area; (d) comparison on power; as well as energy consumption and latency for different workloads: (e.1) DeiT-T, (e.2) DeiT-S, (e.3) DeiT-B; (e.4) BERT-B, and (e.5) BERT-L. former occupies 16.6m 2 area and incurs 697.9mW power; see1. These profiles are dominated by on-chip memory as the impact of Microcomb, modulator, and phase-shifter is significantly decreased compared to state-of-the-art designs. The reason is that, our design strategy for developing MD- Transformer is to eliminate Microcomb, reduce modulator size, and remove phase-shifter in MDOT, hence leading to significantly small area and low power consumption. Area comparison: Our MDTransformer significantly saves area compared to all state-of-the-art designs, i.e., reducing area by 85.3% from LT-Large, 72.4% from LT-Base, and 40.3% from LT-Custom; see 2. MDTransformer occupies smaller area than LT-Custom despite having the same number of tiles, cores, and core size. The reason is that, MDTransformer employs smaller modulator as well as eliminates Micro comb and phase-shifter in MDOT. In addition to that, MDTrans- former also employs smaller number of tiles and smaller core size compared to LT-Large and LT-Base, thus leading to significantly smaller area. Power comparison: Our MDTransformer significantly de- creases power compared to all state-of-the-art designs, i.e., reducing power consumption by 97.5% from LT-Large, 95.3% from LT-Base, and 63.6% from LT-Custom; see 3. MD- Transformer incurs smaller power than LT-Custom despite having the same number of tiles, cores, and core size. The reason is that, MDTransformer employs efficient modulator design as well as completely removes power consumption from Micro comb and phase-shifter in MDOT. In addition to that, MDTransformer also employs smaller number of tiles and smaller core size compared to LT-Large and LT-Base, thus leading to significantly lower power consumption. B. Enabling High Performance and Energy Efficiency across Transformer Workloads Experimental results for energy consumption and latency of core processing on different workloads are provided in Fig. 10(e.1)-(e.5). Based on these results, we make the fol- lowing key observations. • MDTransformer consistently achieves lower energy con- sumption than LT-Custom across different workloads, de- spite having the same number of tiles, cores, and core size; as indicated by4. Specifically, MDTransformer saves en- ergy consumption by 43.1%-43.5% for DeiT-based models and by 40.6%-45.1% for BERT-based models as compared to LT-Custom. The reason is that, MDTransformer elimi- nates power requirement for phase-shifter in MDOT and reduces power cost for modulator, which in turns leading to lower energy consumption when processing the workload. • MDTransformer achieves comparable processing latency to LT-Custom across different workloads, as shown by 5. The reason is that, these two designs consider the same on- chip memory size as well as the same number of tiles, cores, and core size. Therefore, they have similar capabilities in storing data on-chip and performing computation based on their dataflow and scheduling. However, such a similar performance comes at the different cost of power consump- tion, as shown by 3. Therefore, their energy consumption profiles also differ significantly across different workloads, as indicated by4. LT-Large and LT-Base consume relatively low energy since they employ high parallelism to expedite the processing, hence leading to low latency. However, this comes at the cost of huge power consumption, as indicated in Fig. 10(d). This condition may limit the applicability of the LT-Large and LT-Base accelerators for diverse low- power application use-cases. In contrast, MDTransformer achieves competitive energy consumption compared to LT- Large and LT-Base, as shown by 6 , while incurring a significantly lower power consumption than LT-Large and LT-Base, as shown by 3 . The reason is that, MDTrans- former combines the benefits of low-power design through optimized optical devices/modules, selection of architecture configuration, and efficient dataflow for enabling high- performance and energy-efficient optical-based processing. C. Further Discussion In this work, we consider a configuration of N t =4, N c =2, N h =N v =4, N λ =1, and 4 modes for our MDTransformer. 9 However, this selection of configuration can be adjusted based on the requirements. For instance, if we need to increase the parallelism in the MDTransformer, then we can increase the number of tiles N t , number of cores-per-tile N c , and core size N h xN v . Conversely, if we have a targeted application that imposes tight design constraints, e.g., in terms of area, power, energy, and latency (or throughput), then the configuration should be selected carefully. All these adjustment choices are supported with our dynamically-operated architecture and dataflow design in MDTransformer, thereby providing a prac- tical, high-performance and energy-efficient PTA design. VI. CONCLUSION We propose a novel hardware-software co-design of MD- Transformer accelerator, which employs MDM-based compu- tation, inverse-designed coherent PTC, and IQ modulation. Ex- perimental results show that, our MDTransformer accelerator offers 4-bit effective precision for multiplication, low inter- modal crosstalk (i.e., less than -30dB), and full compatibility with single-laser continuous-wave operation at 1550nm. It also saves 40.4% area, 63.6% power, and 40.6% energy consump- tion, with comparable latency over the state-of-the-art across different transformer models. Therefore, our MDTransformer accelerator successfully provides a practical solution for high- performance and energy-efficient transformer-based systems. REFERENCES [1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al., “Attention is all you need,” Advances in Neural Information Processing Systems (NIPS), vol. 30, no. 1, p. 261–272, 2017. [2] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Trans- formers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021. [3] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ́ egou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML). PMLR, 2021, p. 10 347–10 357. [4] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Computing Surveys (CSUR), vol. 54, no. 10s, p. 1–41, 2022. [5] G. Yenduri, R. Murugan, P. Kumar Reddy Maddikunta, S. Bhattacharya, D. Sudheer, and B. Bhushan Savarala, “Artificial general intelligence: Advancements, challenges, and future directions in agi research,” IEEE Access, vol. 13, p. 134 325–134 356, 2025. [6] K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 45, no. 1, p. 87–110, 2023. [7] H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, p. 97–110. [8] M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory- based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, p. 1071–1085. [9] H. You, Z. Sun, H. Shi, Z. Yu, Y. Zhao, Y. Zhang, C. Li, B. Li, and Y. Lin, “Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA).IEEE, 2023, p. 273–286. [10] F. P. Sunny, E. Taheri, M. Nikdast, and S. Pasricha, “A survey on silicon photonics for deep learning,” ACM Journal of Emerging Technologies in Computing System (JETC), vol. 17, no. 4, p. 1–57, 2021. [11] K. Shiflett, A. Karanth, R. Bunescu, and A. Louri, “Albireo: Energy- efficient acceleration of convolutional neural networks via silicon pho- tonics,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2021, p. 860–873. [12] B. J. Shastri et al., “Photonics for artificial intelligence and neuromor- phic computing,” Nature Photonics, vol. 15, no. 2, 2021. [13] Z. Yin, M. Zhang, N. Gangi, R. Huang, J. Zhang, and J. Gu, “Simphony: A device-circuit-architecture cross-layer modeling and simulation frame- work for heterogeneous electronic-photonic ai system,” in 2025 62nd ACM/IEEE Design Automation Conference (DAC).IEEE, 2025, p. 1–7. [14] A. N. Tait, T. F. De Lima, E. Zhou, A. X. Wu, M. A. Nahmias, B. J. Shastri, and P. R. Prucnal, “Neuromorphic photonic networks using silicon photonic weight banks,” Scientific Reports, vol. 7, no. 1, p. 7430, 2017. [15] F. Sunny, A. Mirza, M. Nikdast, and S. Pasricha, “Crosslight: A cross- layer optimized silicon photonic neural network accelerator,” in 2021 58th ACM/IEEE design automation conference (DAC).IEEE, 2021, p. 1069–1074. [16] Y. Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund et al., “Deep learning with coherent nanophotonic circuits,” Nature photonics, vol. 11, no. 7, p. 441–446, 2017. [17] J. Feldmann, N. Youngblood, M. Karpov, H. Gehring, X. Li, M. Stap- pers, M. Le Gallo, X. Fu, A. Lukashchuk, A. S. Raja et al., “Parallel convolutional processing using an integrated photonic tensor core,” Nature, vol. 589, no. 7840, p. 52–58, 2021. [18] H. Zhu, J. Gu, H. Wang, Z. Jiang, Z. Zhang, R. Tang, C. Feng, S. Han, R. T. Chen, and D. Z. Pan, “Lightening-transformer: A dynamically- operated optically-interconnected photonic transformer accelerator,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024, p. 686–703. [19] Y. Li, A. Louri, and A. Karanth, “Sprint: A high-performance, energy- efficient, and scalable chiplet-based accelerator with photonic intercon- nects for cnn inference,” IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 10, p. 2332–2345, 2022. [20] —, “Spacx: Silicon photonics-based scalable chiplet accelerator for dnn inference,” in 2022 IEEE International Symposium on High- Performance Computer Architecture (HPCA), 2022, p. 831–845. [21] S. Afifi, F. Sunny, M. Nikdast, and S. Pasricha, “Tron: Transformer neural network acceleration with non-coherent silicon photonics,” in Great Lakes Symposium on VLSI (GSVLSI) 2023, 2023, p. 15–21. [22] S. Afifi, O. Alo, I. Thakkar, and S. Pasricha, “A light-speed large language model accelerator with optical stochastic computing,” in Great Lakes Symposium on VLSI (GLSVLSI) 2025, 2025. [23] —, “Astra: A stochastic transformer neural network accelerator with silicon photonics,” ACM Transactions on Embedded Computing Systems (TECS), 2025. [24] Y. Li, A. Louri, and A. Karanth, “Merit: A sustainable dnn accelerator design with photonic phase-change memory,” IEEE Transactions on Sustainable Computing (TSUSC), vol. 10, no. 4, p. 705–716, 2025. [25] H. Zhu, Z. Zhou, S. Ning, X. Wu, R. Chen, Y. Wan, and D. Pan, “En- lighten: Lighten the transformer, enable efficient optical acceleration,” arXiv preprint arXiv:2510.01673, 2025. [26] H. Li, D. Chen, and T. Mitra, “Hyatten: Hybrid photonic-digital architec- ture for accelerating attention mechanism,” in 2025 Design, Automation & Test in Europe Conference (DATE), 2025, p. 1–7. [27] W.-T. Chang, C.-F. Wu, and Y.-C. Lo, “P-dac: Power-efficient photonic accelerators for llm inference,” in 2025 62nd ACM/IEEE Design Au- tomation Conference (DAC), 2025, p. 1–7. [28] R. Hamerly, A. Sludds, S. Bandyopadhyay, Z. Chen, Z. Zhong, L. Bern- stein, and D. Englund, “Netcast: low-power edge computing with wdm- defined optical neural networks,” Journal of Lightwave Technology, vol. 42, no. 22, p. 7795–7806, 2024. [29] H. Li, D. Chen, and T. Mitra, “Hybrid photonic-digital accelerator for attention mechanism,” arXiv preprint arXiv:2501.11286, 2025. [30] S. Biasi, G. Donati, A. Lugnan, M. Mancinelli, E. Staffoli, and L. Pavesi, “Photonic neural networks based on integrated silicon microresonators,” Intelligent Computing, vol. 3, p. 0067, 2024. [31] A. N. Tait, A. X. Wu, T. F. De Lima, E. Zhou, B. J. Shastri, M. A. Nahmias, and P. R. Prucnal, “Microring weight banks,” IEEE Journal of Selected Topics in Quantum Electronics, vol. 22, no. 6, p. 312–325, 2016. [32] S. Banerjee, M. Nikdast, and K. Chakrabarty, “Characterizing coherent integrated photonic neural networks under imperfections,” Journal of lightwave technology, vol. 41, no. 5, p. 1464–1479, 2022. 10 [33] A. Totovic, G. Giamougiannis, A. Tsakyridis, D. Lazovsky, and N. Pleros, “Programmable photonic neural networks combining wdm with coherent linear optics,” Scientific reports, vol. 12, no. 1, p. 5605, 2022. [34] I. Flexcompute, “Tidy3D: Next-generation electromagnetic simula- tion tool,” https://w.flexcompute.com/tidy3d/solver/, 2024, accessed: 2025-01-01. [35] N. V. e. a. Sapra, “Inverse design of compact multimode multi-port photonic devices,” Nature Communications, vol. 11, p. 6361, 2020. [36] J. S. e. a. Jensen, “Adjoint-based inverse design of efficient, broadband mode conversion devices,” ACS Photonics, vol. 7, p. 1497–1506, 2020. [37] D. e. a. Vercruysse, “Compact broadband directional couplers using inverse design,” Optica, vol. 7, p. 179–185, 2020. [38] Y. Wang, Y. Wei, V. Dolores-Calzadilla, K. Williams, M. Smit, D. Dai, and Y. Jiao, “Mode division multiplexing on an inp membrane on silicon,” Optics Letters, vol. 47, no. 16, p. 4004–4007, 2022. [39] Y. Liu, K. Xu, S. Wang, W. Shen, H. Xie, Y. Wang, S. Xiao, Y. Yao, J. Du, Z. He et al., “Arbitrarily routed mode-division multiplexed photonic circuits for dense integration,” Nature communications, vol. 10, no. 1, p. 3263, 2019. [40] S. Lam, A. Khaled, S. Bilodeau, B. A. Marquez, P. R. Prucnal, L. Chrostowski, B. J. Shastri, and S. Shekhar, “Dynamic electro-optic analog memory for neuromorphic photonic computing,” arXiv preprint arXiv:2401.16515, 2024. [41] S. Ning, H. Zhu, C. Feng, J. Gu, Z. Jiang, Z. Ying, J. Midkiff, S. Jain, M. H. Hlaing, D. Z. Pan et al., “Photonic-electronic integrated circuits for high-performance computing and ai accelerators,” Journal of Lightwave Technology, 2024. [42] H. Babashah, Z. Kavehvash, A. Khavasi, and S. Koohi, “Temporal analog optical computing using an on-chip fully reconfigurable photonic signal processor,” Optics & Laser Technology, vol. 111, p. 66–74, 2019. [43] L.-W. Luo, N. Ophir, C. P. Chen, L. H. Gabrielli, C. B. Poitras, K. Bergmen, and M. Lipson, “Wdm-compatible mode-division multi- plexing on a silicon chip,” Nature communications, vol. 5, no. 1, p. 3069, 2014. [44] K. Y. Yang, C. Shirpurkar, A. D. White, J. Zang, L. Chang, F. Ashtiani, M. A. Guidry, D. M. Lukin, S. V. Pericherla, J. Yang et al., “Multi- dimensional data transmission using inverse-designed silicon photonics and microcombs,” Nature communications, vol. 13, no. 1, p. 7862, 2022. [45] J. L. Pita Ruiz, N. Dalvand, and M. M ́ enard, “Integrated silicon nitride devices via inverse design,” Nature Communications, vol. 16, no. 1, p. 9307, 2025. [46] J. Li, X. Li, L. Wu, M. Luo, Y. Li, Y. Wang, and Y. Qiu, “Ultra-compact scalable mode demultiplexers for high-speed optical interconnects via gpu-accelerated inverse design,” Optics Express, vol. 33, no. 21, p. 44 908–44 924, 2025. [47] S. Rahimi Kari, N. A. Nobile, D. Pantin, V. Shah, and N. Youngblood, “Realization of an integrated coherent photonic platform for scalable matrix operations,” Optica, vol. 11, no. 4, p. 542–551, 2024. [48] R. V. W. Putra, M. A. Hanif, and M. Shafique, “Drmap: A generic dram data mapping policy for energy-efficient processing of convolu- tional neural networks,” in 2020 57th ACM/IEEE Design Automation Conference (DAC), 2020, p. 1–6. [49] —, “Romanet: Fine-grained reuse-driven off-chip memory access management and data organization for deep neural network acceler- ators,” IEEE Transactions on Very Large Scale Integration Systems (TVLSI), vol. 29, no. 4, p. 702–715, 2021. [50] —, “Pendram: Enabling high-performance and energy-efficient pro- cessing of deep neural networks through a generalized dram data mapping policy,” arXiv preprint arXiv:2408.02412, 2024. [51] P. Caragiulo, O. E. Mattia, A. Arbabian, and B. Murmann, “A 2x time- interleaved 28-gs/s 8-bit 0.03-m 2 switched-capacitor dac in 16-nm finfet cmos,” IEEE Journal of Solid-State Circuits, vol. 56, no. 8, p. 2335–2346, 2021. [52] Y. Duan and E. Alon, “A 12.8 gs/s time-interleaved adc with 25 ghz effective resolution bandwidth and 4.6 enob,” IEEE Journal of Solid- State Circuits, vol. 49, no. 8, p. 1725–1738, 2014. [53] S. Serunjogi, M. Rasras, and M. Sanduleanu, “64gb/s nrz/pam4 burst- mode optical receiver frontend with gain control, offset correction and gain decoupled from bandwidth,” in 2021 IEEE International Sympo- sium on Circuits and Systems (ISCAS). IEEE, 2021, p. 1–4.