SearcharxivSearch

arXiv subjects

Siqi Yan

Publications and source records attributed to Siqi Yan.

18 recordsLinked to original sources

A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.

cs.CE

High-rate continuous-variable quantum key distribution coexisting with Tb/s coherent classical transmission in hollow-core fiber

Quantum key distribution (QKD) can provide secret keys with security rooted in quantum mechanics, but operation alongside high-capacity classical traffic remains limited by the excess-noise budget of weak quantum states in conventional solid-core fiber. Here, we combine ultralow-loss anti-resonant hollow-core fiber with residual-carrier-assisted discrete-modulation continuous-variable QKD (DM-CV-QKD) to address both propagation-induced coexistence noise and low-SNR phase recovery. Over a 24.3-km hollow-core link with 3.3-dB end-to-end loss, a dual-polarization 15-Gbaud DM-CV-QKD channel achieves an average asymptotic secret-key rate (SKR) of 153.22 Mb/s and a finite-size SKR of 149.99 Mb/s, while 39 coherent wavelength-division-multiplexed channels deliver an aggregate data rate of 7.6 Tb/s and a net data rate of 7.2 Tb/s. The system can even sustain a positive SKR under a high classical launch power of up to 15 dBm, without an optical bandpass filter (BPF). Finite-size analysis against collective attacks further yields a projected positive secret-key rate at a 100-km-equivalent condition. These results show that an anti-resonant hollow-core fiber, combined with carrier-assisted phase recovery, can greatly extend the operating regime of shared-fiber quantum-secured coherent links, pointing to a promising approach for integrating high-rate CV-QKD with high-capacity optical networks.

quant-ph

Shell-Supervised Gaussian Splatting for Urban Real-to-Sim Reconstruction

Real-to-sim reconstruction for embodied AI requires geometry that is useful for collision reasoning, navigation, and agent-environment interaction, not only photorealistic novel-view synthesis. However, close-range urban facades are difficult for video-to-3D reconstruction: glass, reflections, repeated windows, and weak texture can produce visually plausible renderings with unstable surface geometry. We introduce shell-supervised Gaussian Splatting, a reconstruction-stage framework that uses an external facade structural shell as lightweight geometric supervision for video-driven Gaussian reconstruction. The method aligns an exterior shell to the video reconstruction frame, renders per-view depth, camera-space normal, and valid-mask maps, and applies these cues through mask-gated losses during Gaussian optimization. This design preserves RGB-driven appearance while regularizing only visible shell-supported facade regions. Experiments on anonymized close-range urban facade scenes show improved facade orientation and visible-surface point-cloud consistency over photo-only, monocular-cue, and surface-oriented Gaussian baselines, while maintaining comparable held-out rendering quality.

cs.CV

A unified resource-pool architecture for high-dimensional direct-detection optical communication

Increasing optical communication capacity without proportionally increasing receiver complexity remains a key challenge for direct-detection links. Conventional systems typically assign wavelength, polarization and intensity to fixed, separately recovered functions, so that alphabet expansion is accompanied by additional demultiplexing, polarization handling, receiver branches and electronic processing. Here we introduce a unified resource-pool architecture for high-dimensional direct-detection optical communication, in which wavelength, polarization and intensity are jointly organized as a composite optical symbol space and recovered through optical-domain joint projection rather than dimension-by-dimension separation. The receiver is implemented with an integrated disordered photonic processor that transforms each composite optical state into a reproducible multi-output electrical fingerprint for single-shot direct recovery. In a dual-wavelength transmission experiment, the system resolves 4096 composite symbols, corresponding to 12 bits per symbol slot, with a bit error rate of 4.25e-4 after 10 km standard-fiber transmission. Additional experiments demonstrate dense polarization alphabets, wavelength-indexed state-space expansion and high-launch-power operation over hollow-core fiber. These results establish disorder-enabled joint projection in an integrated photonic processor as a route to hardware-efficient high-dimensional direct-detection communication beyond conventional dimension-partitioned receiver architecture.

physics.optics

FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost

Modern industrial Deep Learning Recommendation Models typically extract user preferences through the analysis of sequential interaction histories, subsequently generating predictions based on these derived interests. The inherent heterogeneity in data characteristics frequently result in substantial under-utilization of computational resources during large-scale training, primarily due to computational bubbles caused by severe stragglers and slow blocking communications. This paper introduces FreeScale, a solution designed to (1) mitigate the straggler problem through meticulously load balanced input samples (2) minimize the blocking communication by overlapping prioritized embedding communications with computations (3) resolve the GPU resource competition during computation and communication overlapping by communicating through SM-Free techniques. Empirical evaluation demonstrates that FreeScale achieves up to 90.3% reduction in computational bubbles when applied to real-world workloads running on 256 H100 GPUs.

cs.LG

Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

We introduce Sparse Forcing, a training-and-inference paradigm for autoregressive video diffusion models that improves long-horizon generation quality while reducing decoding latency. Sparse Forcing is motivated by an empirical observation in autoregressive diffusion rollouts: attention concentrates on a persistent subset of salient visual blocks, forming an implicit spatiotemporal memory in the KV cache, and exhibits a locally structured block-sparse pattern within sliding windows. Building on this observation, we propose a trainable native sparsity mechanism that learns to compress, preserve, and update these persistent blocks while restricting computation within each local window to a dynamically selected local neighborhood. To make the approach practical at scale for both training and inference, we further propose Persistent Block-Sparse Attention (PBSA), an efficient GPU kernel that accelerates sparse attention and memory updates for low-latency, memory-efficient decoding. Experiments show that Sparse Forcing improves the VBench score by +0.26 over Self-Forcing on 5-second text-to-video generation while delivering a 1.11-1.17x decoding speedup and 42% lower peak KV-cache footprint. The gains are more pronounced on longer-horizon rollouts, delivering improved visual quality with +0.68 and +2.74 VBench improvements, and 1.22x and 1.27x speedups on 20-second and 1-minute generations, respectively.

cs.CV

MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs

The pervasive "memory wall" bottleneck is significantly amplified in modern large-scale Mixture-of-Experts (MoE) architectures. MoE's inherent architectural sparsity leads to sparse arithmetic compute and also introduces substantial activation memory overheads -- driven by large token routing buffers and the need to materialize and buffer intermediate tensors. This memory pressure limits the maximum batch size and sequence length that can fit on GPUs, and also results in excessive data movements that hinders performance and efficient model scaling. We present MoEBlaze, a memory-efficient MoE training framework that addresses these issues through a co-designed system approach: (i) an end-to-end token dispatch and MoE training method with optimized data structures to eliminate intermediate buffers and activation materializing, and (ii) co-designed kernels with smart activation checkpoint to mitigate memory footprint while simultaneously achieving better performance. We demonstrate that MoEBlaze can achieve over 4x speedups and over 50% memory savings compared to existing MoE frameworks.

cs.LG

First Thin-Film Lithium Tantalate Polarization Controller Enabling Reset-Free Mrad/s Tracking for Optical Interconnects

The rapid escalation of computing power driven by large-scale artificial intelligence is placing unprecedented demands on the bandwidth, latency, and energy efficiency of data-center interconnects (DCIs). Self-homodyne coherent (SHC) transmission is a promising architecture because it preserves the spectral efficiency of coherent detection while greatly simplifying digital signal processing, but its practical deployment is critically limited by random and often ultrafast state-of-polarization (SOP) fluctuations that induce carrier fading and destabilize coherent reception. Here we report the first integrated polarization controller based on thin-film lithium tantalate (TFLT), enabling reset-free polarization tracking at Mrad/s speeds. The four-stage electro-optic device exhibits polarization-dependent loss (PDL) below 0.3 dB, a half-wave voltage below 2.5 V, high modulation bandwidth, and negligible DC drift. To accommodate the finite tuning range of integrated phase shifters, we develop a finite-boundary gradient-descent (FBGD) control algorithm that ensures reset-free SOP evolution with no phase jump. The implemented adaptive polarization controller (APC) is validated through both standalone polarization-tracking measurements and a dual-polarization 16-QAM SHC 400-Gbps transmission system. Transient polarization disturbances can be tracked at speeds up to 2 Mrad/s, while stable reset-free operation under continuous polarization disturbances is maintained up to 1 Mrad/s. This reset-free performance represents more than doubling the state of the art, while the pre-FEC bit-error rates remain below the HD-FEC threshold under realistic DCI conditions and lightning-scale polarization disturbances. These results establish TFLT as a new platform for ultrafast, low-power, reset-free, and drift-free polarization control in coherent optical interconnects and beyond.

physics.optics

Robust and Hyper-Efficient Multi-dimensional Optical Fiber Semantic Communication

The growing demands of artificial intelligence and immersive media require communication beyond bit-level accuracy to meaning awareness. Conventional optical systems that focused on syntactic precision suffer significant inefficiencies. Here, we introduce a multi-dimensional semantic communication framework that bridges this gap by directly mapping high-level semantic features onto the orthogonal physical dimensions of light, frequency, polarization, and intensity, within a multimode fiber. This synergistic co-design of semantic logic and the photonic channel achieve an unprecedented equivalent spectral efficiency approaching 1000 bit/s/Hz. Moreover, it demonstrates profound resilience, maintaining high-fidelity reconstruction even when the physical-layer symbol error rate exceeds 36%, a condition under which conventional communication systems fail completely. Crucially, this deeply integrated co-design of semantic encoding and physical-layer modulation enables full semantic demodulation with only single-ended intensity detection, therefore significantly reducing system complexity and cost. This work establishes a validated pathway toward hyper-efficient, error-resilient optical networks for the next generation of data-intensive computing.

physics.optics

Pushing DSP-Free Coherent Interconnect to the Last Inch by Optically Analog Signal Processing

To support the boosting interconnect capacity of the AI-related data centers, novel techniques enabled high-speed and low-cost optics are continuously emerging. When the baud rate approaches 200 GBaud per lane, the bottle-neck of traditional intensity modulation direct detection (IM-DD) architectures becomes increasingly evident. The simplified coherent solutions are widely discussed and considered as one of the most promising candidates. In this paper, a novel coherent architecture based on self-homodyne coherent detection and optically analog signal processing (OASP) is demonstrated. Proved by experiment, the first DSP-free baud-rate sampled 64-GBaud QPSK/16-QAM receptions are achieved, with BERs of 1e-6 and 2e-2, respectively. Even with 1-km fiber link propagation, the BER for QPSK reception remains at 3.6e-6. When an ultra-simple 1-sps SISO filter is utilized, the performance degradation of the proposed scheme is less than 1 dB compared to legacy DSP-based coherent reception. The proposed results pave the way for the ultra-high-speed coherent optical interconnections, offering high power and cost efficiency.

physics.optics

Quantum-Secured DSP-Lite Data Transmission Architectures for AI-Driven Data Centres

Artificial intelligence-driven (AI-driven) data centres, which require high-performance, scalable, energy-efficient, and secure infrastructure, have led to unprecedented data traffic demands. These demands involve low latency, high bandwidth connections, low power consumption, and data confidentiality. However, conventional optical interconnect solutions, such as intensity-modulated direct detection and traditional coherent systems, cannot address these requirements simultaneously. In particular, conventional encryption protocols that rely on complex algorithms are increasingly vulnerable to the rapid advancement of quantum computing. Here, we propose and demonstrate a quantum-secured digital signal processing-lite (DSP-Lite) data transmission architecture that meets all the stringent requirements for AI-driven data centre optical interconnects (AI-DCIs) scenarios. By integrating a self-homodyne coherent (SHC) system and quantum key distribution (QKD) through the multicore-fibre-based space division multiplexing (SDM) technology, our scheme enables secure, high-capacity, and energy-efficient data transmission while ensuring resilience against quantum computing threats. In our demonstration, we achieved an expandable transmission capacity of 2 Tbit per second (Tb/s) and a quantum secret key rate (SKR) of 229.2 kb/s, with a quantum bit error rate (QBER) of approximately 1.27% and with ultralow power consumption. Our work paves the way for constructing secure, scalable, and cost-efficient data transmission frameworks, thus enabling the next generation of intelligent, leak-proof optical interconnects for data centres.

quant-ph

Two-Stage Optimization for Efficient V2G Coordination in Distribution Power System

With the growing popularity of electric vehicles (EVs), maintaining power grid stability has become a significant challenge. To address this issue, EV scheduling control strategies have been developed to manage vehicle-to-grid (V2G) in coordination with the optimal power flow. In existing studies, such coordination optimization is formulated as a mixed-integer nonlinear programming (MINP), which is computationally challenging due to the binary EV charging and discharging variables. To address this challenge, we develop an efficient two-stage optimization method for this mixed-integer nonlinear coordination problem. This method first employs an efficient technique called the difference of convex (DC) to relax the integrality and reformulate MINP into a series of path-following continuous programming. Although the DC approach shows promising efficiency for solving MINP, it cannot guarantee the feasibility of the solutions. Consequently, we propose a trust region optimization method in stage two that constructs a trust region around DC's solution and then searches for the best feasible solution within this region. Our simulation results demonstrate that, compared to the open-source optimization solver SCIP, our proposed method significantly enhances computational efficiency while achieving near optimality.

cs.CE

Addressing Data Scarcity in Optical Matrix Multiplier Modeling Using Transfer Learning

We present and experimentally evaluate using transfer learning to address experimental data scarcity when training neural network (NN) models for Mach-Zehnder interferometer mesh-based optical matrix multipliers. Our approach involves pre-training the model using synthetic data generated from a less accurate analytical model and fine-tuning with experimental data. Our investigation demonstrates that this method yields significant reductions in modeling errors compared to using an analytical model, or a standalone NN model when training data is limited. Utilizing regularization techniques and ensemble averaging, we achieve < 1 dB root-mean-square error on the matrix weights implemented by a 3x3 photonic chip while using only 25% of the available data.

cs.LG

Data-driven Modeling of Mach-Zehnder Interferometer-based Optical Matrix Multipliers

Photonic integrated circuits are facilitating the development of optical neural networks, which have the potential to be both faster and more energy efficient than their electronic counterparts since optical signals are especially well-suited for implementing matrix multiplications. However, accurate programming of photonic chips for optical matrix multiplication remains a difficult challenge. Here, we describe both simple analytical models and data-driven models for offline training of optical matrix multipliers. We train and evaluate the models using experimental data obtained from a fabricated chip featuring a Mach-Zehnder interferometer mesh implementing 3-by-3 matrix multiplication. The neural network-based models outperform the simple physics-based models in terms of prediction error. Furthermore, the neural network models are also able to predict the spectral variations in the matrix weights for up to 100 frequency channels covering the C-band. The use of neural network models for programming the chip for optical matrix multiplication yields increased performance on multiple machine learning tasks.

cs.LG

Captum: A unified and generic model interpretability library for PyTorch

In this paper we introduce a novel, unified, open-source model interpretability library for PyTorch [12]. The library contains generic implementations of a number of gradient and perturbation-based attribution algorithms, also known as feature, neuron and layer importance algorithms, as well as a set of evaluation metrics for these algorithms. It can be used for both classification and non-classification models including graph-structured models built on Neural Networks (NN). In this paper we give a high-level overview of supported attribution algorithms and show how to perform memory-efficient and scalable computations. We emphasize that the three main characteristics of the library are multimodality, extensibility and ease of use. Multimodality supports different modality of inputs such as image, text, audio or video. Extensibility allows adding new algorithms and features. The library is also designed for easy understanding and use. Besides, we also introduce an interactive visualization tool called Captum Insights that is built on top of Captum library and allows sample-based model debugging and visualization using feature importance metrics.

cs.LG

Slow-light-enhanced energy efficiency for the graphene microheater on silicon photonic crystal waveguides

Slow light has been widely utilized to obtain enhanced nonlinearities, enhanced spontaneous emissions, and increased phase shifts owing to its ability to promote light-matter interactions. By incorporating a graphene microheater on a slow-light silicon photonic crystal waveguide, we experimentally demonstrated an energy-efficient graphene microheater with a tuning efficiency of 1.07 nm/mW and power consumption per free spectral range of 3.99 mW. The rise and decay times (10% to 90%) were only 750 ns and 525 ns, which, to the best of our knowledge, are the fastest reported response times for microheaters in silicon photonics. The corresponding record-low figure of merit of the device was 2.543 nW.s, which is one order of magnitude lower than results reported in previous studies. The influences of the graphene-photonic crystal waveguide interaction length and the shape of the graphene heater were also investigated, providing valuable guidelines for enhancing the graphene microheater tuning efficiency.

physics.optics