SearcharxivSearch

arXiv subjects

Ouhan Huang

Publications and source records attributed to Ouhan Huang.

7 recordsLinked to original sources

Reinforcement Learning-Enabled Agent for Transmitter Optimization in Digital-Analog Radio-over-Fiber Fronthaul

Digital-analog radio-over-fiber (DA-RoF) has emerged as a promising fronthaul solution that combines the high spectral efficiency of analog transmission with the robustness of digital transmission. However, the performance of DA-RoF critically depends on several tightly coupled parameters, including the rounding factor (RF), scaling factor (SF), geometric shaping (GS) factor, and pre-equalization taps coefficients, which jointly affect quantization noise, nonlinear distortion, and bandwidth-induced inter-symbol interference (ISI). Conventional grid search-based optimization is computationally prohibitive and impractical for optical communication. In this work, we propose a reinforcement-learning (RL)-enabled DA-RoF fronthaul agent architecture, capable of autonomously learning optimal transmitter parameters from end-to-end signal-to-noise ratio (SNR) feedback without a differentiable channel model. Experimental results demonstrate that the trained agent steadily improves SNR through sequential decision making and outperforms baseline, achieving ~2.7-dB SNR improvement for 1- to 4-order DA-RoF transmission, reaching final SNR of 35.8 dB, 42.9 dB, 53.8 dB, and 63.2 dB and supporting 1024-, 4096-, 16384-, 65536-quadrature amplitude modulation (QAM) format, respectively. These results validate that the proposed RL-enabled framework provides online, scalable, and hardware-efficient parameter optimization for DA-RoF fronthaul systems, paving the way toward high-order modulation format and intelligent next-generation radio access networks.

physics.optics

Fully multiplexed photonic tensor computing

Tensor operations dominate modern computational workloads, yet their further acceleration demands hardware platforms with greater parallelism. Although photonic computing provides a compelling route for parallel processing, fully exploiting all native multiplexing dimensions of optical fields is impeded by the challenges in routing and programming light in all dimensions simultaneously. Here we introduce FieldCore, a fully multiplexed photonic tensor core that jointly harnesses wavelength, radio-frequency, guided-mode, time and space dimensions, thereby enabling parallelism to scale multiplicatively within a single optical field. Enabled by inverse-designed silicon photonics, FieldCore preserves a uniform programmed computation across all multiplexed channels in parallel. Experimentally, we validate and benchmark its performance from ultra-high-baudrate arithmetic operations to high-fidelity image convolution and parallel handwritten-digit recognition. We further use FieldCore to unlock applications that naturally require high-dimensional data processing, such as high-dimensional hyperspectral classification and massively parallel mechanical fault diagnosis. Our FieldCore supports an estimated aggregate compute throughput of 69.12 tera operations per second (TOPS) and accommodates up to 1,800 parallel input streams within a single core, establishing a scalable paradigm for fully multiplexed photonic tensor computing and AI inference.

physics.optics

Seamless Optical Cloud Computing across Edge-Metro Network for Generative AI

The rapid advancement of generative artificial intelligence (AI) in recent years has profoundly reshaped modern lifestyles, necessitating a revolutionary architecture to support the growing demands for computational power. Cloud computing has become the driving force behind this transformation. However, it consumes significant power and faces computation security risks due to the reliance on extensive data centers and servers in the cloud. Reducing power consumption while enhancing computational scale remains persistent challenges in cloud computing. Here, we propose and experimentally demonstrate an optical cloud computing system that can be seamlessly deployed across edge-metro network. By modulating inputs and models into light, a wide range of edge nodes can directly access the optical computing center via the edge-metro network. The experimental validations show an energy efficiency of 118.6 mW/TOPs (tera operations per second), reducing energy consumption by two orders of magnitude compared to traditional electronic-based cloud computing solutions. Furthermore, it is experimentally validated that this architecture can perform various complex generative AI models through parallel computing to achieve image generation tasks.

cs.DC

Edge-guided inverse design of digital metamaterial-based mode multiplexers for high-capacity multi-dimensional interconnect

The escalating demands of compute-intensive applications urgently necessitate the adoption of optical interconnect technologies to overcome bottlenecks in scaling computing systems. This requires fully exploiting the inherent parallelism of light across scalable dimensions for data loading. Here we experimentally demonstrate a synergy of wavelength- and mode- multiplexing combined with high-order modulation formats to achieve multi-tens-of-terabits-per-second optical interconnects using foundry-compatible silicon photonic circuits. Implementing an edge-guided analog-and-digital optimization method that integrates high efficiency with fabrication robustness, we achieve the inverse design of mode multiplexers based on digital metamaterial waveguides. Furthermore, we employ a packaged five-mode multiplexing chip, achieving a single-wavelength interconnect capacity of 1.62 Tbit s-1 and a record-setting multi-dimensional interconnect capacity of 38.2 Tbit s-1 across 5 modes and 88 wavelength channels, with high-order formats up to 8-ary pulse-amplitude-modulation (PAM). This study highlights the transformative potential of optical interconnect technologies to surmount the constraints of electronic links, thus setting the stage for next-generation datacenter and optical compute interconnects.

physics.optics

Spatial Temporal Attention based Target Vehicle Trajectory Prediction for Internet of Vehicles

Forecasting vehicle behavior within complex traffic environments is pivotal within Intelligent Transportation Systems (ITS). Though this technology plays a significant role in alleviating the prevalent operational difficulties in logistics and transportation systems, the precise prediction of vehicle trajectories still poses a substantial challenge. To address this, our study introduces the Spatio Temporal Attention-based methodology for Target Vehicle Trajectory Prediction (STATVTPred). This approach integrates Global Positioning System(GPS) localization technology to track target movement and dynamically predict the vehicle's future path using comprehensive spatio-temporal trajectory data. We map the vehicle trajectory onto a directed graph, after which spatial attributes are extracted via a Graph Attention Networks(GATs). The Transformer technology is employed to yield temporal features from the sequence. These elements are then amalgamated with local road network structure maps to filter and deliver a smooth trajectory sequence, resulting in precise vehicle trajectory prediction.This study validates our proposed STATVTPred method on T-Drive and Chengdu taxi-trajectory datasets. The experimental results demonstrate that STATVTPred achieves 6.38% and 10.55% higher Average Match Rate (AMR) than the Transformer model on the Beijing and Chengdu datasets, respectively. Compared to the LSTM Encoder-Decoder model, STATVTPred boosts AMR by 37.45% and 36.06% on the same datasets. This is expected to establish STATVTPred as a new approach for handling trajectory prediction of targets in logistics and transportation scenarios, thereby enhancing prediction accuracy.

cs.RO

VTP: Volumetric Transformer for Multi-view Multi-person 3D Pose Estimation

This paper presents Volumetric Transformer Pose estimator (VTP), the first 3D volumetric transformer framework for multi-view multi-person 3D human pose estimation. VTP aggregates features from 2D keypoints in all camera views and directly learns the spatial relationships in the 3D voxel space in an end-to-end fashion. The aggregated 3D features are passed through 3D convolutions before being flattened into sequential embeddings and fed into a transformer. A residual structure is designed to further improve the performance. In addition, the sparse Sinkhorn attention is empowered to reduce the memory cost, which is a major bottleneck for volumetric representations, while also achieving excellent performance. The output of the transformer is again concatenated with 3D convolutional features by a residual design. The proposed VTP framework integrates the high performance of the transformer with volumetric representations, which can be used as a good alternative to the convolutional backbones. Experiments on the Shelf, Campus and CMU Panoptic benchmarks show promising results in terms of both Mean Per Joint Position Error (MPJPE) and Percentage of Correctly estimated Parts (PCP). Our code will be available.

cs.CV

Exploring Global Diversity and Local Context for Video Summarization

Video summarization aims to automatically generate a diverse and concise summary which is useful in large-scale video processing. Most of the methods tend to adopt self-attention mechanism across video frames, which fails to model the diversity of video frames. To alleviate this problem, we revisit the pairwise similarity measurement in self-attention mechanism and find that the existing inner-product affinity leads to discriminative features rather than diversified features. In light of this phenomenon, we propose global diverse attention which uses the squared Euclidean distance instead to compute the affinities. Moreover, we model the local contextual information by novel local contextual attention to remove the redundancy in the video. By combining these two attention mechanisms, a video SUMmarization model with Diversified Contextual Attention scheme is developed, namely SUM-DCA. Extensive experiments are conducted on benchmark data sets to verify the effectiveness and the superiority of SUM-DCA in terms of F-score and rank-based evaluation without any bells and whistles.

cs.CV