SearcharxivSearch

arXiv subjects

Shilin Xiao

Publications and source records attributed to Shilin Xiao.

12 recordsLinked to original sources

Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks

Vision-Language-Action (VLA) models revolutionize robotic systems by enabling end-to-end perception-to-action pipelines that integrate multiple sensory modalities, such as visual signals processed by cameras and auditory signals captured by microphones. This multi-modality integration allows VLA models to interpret complex, real-world environments using diverse sensor data streams. Given the fact that VLA-based systems heavily rely on the sensory input, the security of VLA models against physical-world sensor attacks remains critically underexplored. To address this gap, we present the first systematic study of physical sensor attacks against VLAs, quantifying the influence of sensor attacks and investigating the defenses for VLA models. We introduce a novel "Real-Sim-Real" framework that automatically simulates physics-based sensor attack vectors, including six attacks targeting cameras and two targeting microphones, and validates them on real robotic systems. Through large-scale evaluations across various VLA architectures and tasks under varying attack parameters, we demonstrate significant vulnerabilities, with susceptibility patterns that reveal critical dependencies on task types and model designs. We further develop an adversarial-training-based defense that enhances VLA robustness against out-of-distribution physical perturbations caused by sensor attacks while preserving model performance. Our findings expose an urgent need for standardized robustness benchmarks and mitigation strategies to secure VLA deployments in safety-critical environments.

cs.RO

SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities

Sensors are fundamental to cyber-physical systems (CPS), enabling perception and control by transducing physical stimuli into digital measurements. However, despite growing research on physical attacks on sensors, our understanding of sensor hardware vulnerabilities remains fragmented due to the ad-hoc nature of this field. Moreover, the infinite attack signal space further complicates threat abstraction and defense. To address this gap, we propose a systematization framework, termed sensor out-of-band (OOB) vulnerabilities, that for the first time provides a comprehensive abstraction for sensor attack surfaces based on underlying physical principles. We adopt a bottom-up systematization methodology that analyzes OOB vulnerabilities across three levels. At the component level, we identify the physical principles and limitations that contribute to OOB vulnerabilities. At the sensor level, we categorize known attacks and evaluate their practicality. At the system level, we analyze how CPS features such as sensor fusion, closed-loop control, and intelligent perception impact the exposure and mitigation of OOB threats. Our findings offer a foundational understanding of sensor hardware security and provide guidance and future directions for sensor designers, security researchers, and system developers aiming to build more secure sensors and CPS.

cs.CR

Throughput Maximization in Multi-Band Optical Networks with Column Generation

Multi-band transmission is a promising technical direction for spectrum and capacity expansion of existing optical networks. Due to the increase in the number of usable wavelengths in multi-band optical networks, the complexity of resource allocation problems becomes a major concern. Moreover, the transmission performance, spectrum width, and cost constraint across optical bands may be heterogeneous. Assuming a worst-case transmission margin in U, L, and C-bands, this paper investigates the problem of throughput maximization in multi-band optical networks, including the optimization of route, wavelength, and band assignment. We propose a low-complexity decomposition approach based on Column Generation (CG) to address the scalability issue faced by traditional methodologies. We numerically compare the results obtained by our CG-based approach to an integer linear programming model, confirming the near-optimal network throughput. Our results also demonstrate the scalability of the CG-based approach when the number of wavelengths increases, with the computation time in the magnitude order of 10 s for cases varying from 75 to 1200 wavelength channels per link in a 14-node network. Code of this publication is available at github.com/cchen000/CG-Multi-Band.

cs.NI

Private Eye: On the Limits of Textual Screen Peeking via Eyeglass Reflections in Video Conferencing

Using mathematical modeling and human subjects experiments, this research explores the extent to which emerging webcams might leak recognizable textual and graphical information gleaming from eyeglass reflections captured by webcams. The primary goal of our work is to measure, compute, and predict the factors, limits, and thresholds of recognizability as webcam technology evolves in the future. Our work explores and characterizes the viable threat models based on optical attacks using multi-frame super resolution techniques on sequences of video frames. Our models and experimental results in a controlled lab setting show it is possible to reconstruct and recognize with over 75% accuracy on-screen texts that have heights as small as 10 mm with a 720p webcam. We further apply this threat model to web textual contents with varying attacker capabilities to find thresholds at which text becomes recognizable. Our user study with 20 participants suggests present-day 720p webcams are sufficient for adversaries to reconstruct textual content on big-font websites. Our models further show that the evolution towards 4K cameras will tip the threshold of text leakage to reconstruction of most header texts on popular websites. Besides textual targets, a case study on recognizing a closed-world dataset of Alexa top 100 websites with 720p webcams shows a maximum recognition accuracy of 94% with 10 participants even without using machine-learning models. Our research proposes near-term mitigations including a software prototype that users can use to blur the eyeglass areas of their video streams. For possible long-term defenses, we advocate an individual reflection testing procedure to assess threats under various settings, and justify the importance of following the principle of least privilege for privacy-sensitive scenarios.

cs.CR

Low-complexity full-field ultrafast nonlinear dynamics prediction by a convolutional feature separation modeling method

The modeling and prediction of the ultrafast nonlinear dynamics in the optical fiber are essential for the studies of laser design, experimental optimization, and other fundamental applications. The traditional propagation modeling method based on the nonlinear Schrödinger equation (NLSE) has long been regarded as extremely time-consuming, especially for designing and optimizing experiments. The recurrent neural network (RNN) has been implemented as an accurate intensity prediction tool with reduced complexity and good generalization capability. However, the complexity of long grid input points and the flexibility of neural network structure should be further optimized for broader applications. Here, we propose a convolutional feature separation modeling method to predict full-field ultrafast nonlinear dynamics with low complexity and high flexibility, where the linear effects are firstly modeled by NLSE-derived methods, then a convolutional deep learning method is implemented for nonlinearity modeling. With this method, the temporal relevance of nonlinear effects is substantially shortened, and the parameters and scale of neural networks can be greatly reduced. The running time achieves a 94% reduction versus NLSE and an 87% reduction versus RNN without accuracy deterioration. In addition, the input pulse conditions, including grid point numbers, durations, peak powers, and propagation distance, can be flexibly changed during the predicting process. The results represent a remarkable improvement in the ultrafast nonlinear dynamics prediction and this work also provides novel perspectives of the feature separation modeling method for quickly and flexibly studying the nonlinear characteristics in other fields.

physics.optics

Fast and accurate waveform modeling of long-haul multi-channel optical fiber transmission using a hybrid model-data driven scheme

The modeling of optical wave propagation in optical fiber is a task of fast and accurate solving the nonlinear Schrödinger equation (NLSE), and can enable the optical system design, digital signal processing verification and fast waveform calculation. Traditional waveform modeling of full-time and full-frequency information is the split-step Fourier method (SSFM), which has long been regarded as challenging in long-haul wavelength division multiplexing (WDM) optical fiber communication systems because it is extremely time-consuming. Here we propose a linear-nonlinear feature decoupling distributed (FDD) waveform modeling scheme to model long-haul WDM fiber channel, where the channel linear effects are modelled by the NLSE-derived model-driven methods and the nonlinear effects are modelled by the data-driven deep learning methods. Meanwhile, the proposed scheme only focuses on one-span fiber distance fitting, and then recursively transmits the model to achieve the required transmission distance. The proposed modeling scheme is demonstrated to have high accuracy, high computing speeds, and robust generalization abilities for different optical launch powers, modulation formats, channel numbers and transmission distances. The total running time of FDD waveform modeling scheme for 41-channel 1040-km fiber transmission is only 3 minutes versus more than 2 hours using SSFM for each input condition, which achieves a 98% reduction in computing time. Considering the multi-round optimization by adjusting system parameters, the complexity reduction is significant. The results represent a remarkable improvement in nonlinear fiber modeling and open up novel perspectives for solution of NLSE-like partial differential equations and optical fiber physics problems.

eess.SP

Fast and Accurate Optical Fiber Channel Modeling Using Generative Adversarial Network

In this work, a new data-driven fiber channel modeling method, generative adversarial network (GAN) is investigated to learn the distribution of fiber channel transfer function. Our investigation focuses on joint channel effects of attenuation, chromic dispersion, self-phase modulation (SPM), and amplified spontaneous emission (ASE) noise. To achieve the success of GAN for channel modeling, we modify the loss function, design the condition vector of input and address the mode collapse for the long-haul transmission. The effective architecture, parameters, and training skills of GAN are also displayed in the paper. The results show that the proposed method can learn the accurate transfer function of the fiber channel. The transmission distance of modeling can be up to 1000 km and can be extended to arbitrary distance theoretically. Moreover, GAN shows robust generalization abilities under different optical launch powers, modulation formats, and input signal distributions. Comparing the complexity of GAN with the split-step Fourier method (SSFM), the total multiplication number is only 2% of SSFM and the running time is less than 0.1 seconds for 1000-km transmission, versus 400 seconds using the SSFM under the same hardware and software conditions, which highlights the remarkable reduction in complexity of the fiber channel modeling.

cs.IT

Throughput Maximization Leveraging Just-Enough SNR Margin and Channel Spacing Optimization

Flexible optical network is a promising technology to accommodate high-capacity demands in next-generation networks. To ensure uninterrupted communication, existing lightpath provisioning schemes are mainly done with the assumption of worst-case resource under-provisioning and fixed channel spacing, which preserves an excessive signal-to-noise ratio (SNR) margin. However, under a resource over-provisioning scenario, the excessive SNR margin restricts the transmission bit-rate or transmission reach, leading to physical layer resource waste and stranded transmission capacity. To tackle this challenging problem, we leverage an iterative feedback tuning algorithm to provide a just-enough SNR margin, so as to maximize the network throughput. Specifically, the proposed algorithm is implemented in three steps. First, starting from the high SNR margin setup, we establish an integer linear programming model as well as a heuristic algorithm to maximize the network throughput by solving the problem of routing, modulation format, forward error correction, baud-rate selection, and spectrum assignment. Second, we optimize the channel spacing of the lightpaths obtained from the previous step, thereby increasing the available physical layer resources. Finally, we iteratively reduce the SNR margin of each lightpath until the network throughput cannot be increased. Through numerical simulations, we confirm the throughput improvement in different networks and with different baud-rates. In particular, we find that our algorithm enables over 20\% relative gain when network resource is over-provisioned, compared to the traditional method preserving an excessive SNR margin.

cs.NI

Maximizing Revenue with Adaptive Modulation and Multiple FECs in Flexible Optical Networks

Flexible optical networks (FONs) are being adopted to accommodate the increasingly heterogeneous traffic in today's Internet. However, in presence of high traffic load, not all offered traffic can be satisfied at all time. As carried traffic load brings revenues to operators, traffic blocking due to limited spectrum resource leads to revenue losses. In this study, given a set of traffic requests to be provisioned, we consider the problem of maximizing operator's revenue, subject to limited spectrum resource and physical layer impairments (PLIs), namely amplified spontaneous emission noise (ASE), self-channel interference (SCI), cross-channel interference (XCI), and node crosstalk. In FONs, adaptive modulation, multiple FEC, and the tuning of power spectrum density (PSD) can be effectively employed to mitigate the impact of PLIs. Hence, in our study, we propose a universal bandwidth-related impairment evaluation model based on channel bandwidth, which allows a performance analysis for different PSD, FEC and modulations. Leveraging this PLI model and a piecewise linear fitting function, we succeed to formulate the revenue maximization problem as a mixed integer linear program. Then, to solve the problem on larger network instances, a fast two-phase heuristic algorithm is also proposed, which is shown to be near-optimal for revenue maximization. Through simulations, we demonstrate that using adaptive modulation enables to significantly increase revenues in the scenario of high signal-to-noise ratio (SNR), where the revenue can even be doubled for high traffic load, while using multiple FECs is more profitable for scenarios with low SNR.

cs.NI