SearcharxivSearch

arXiv subjects

Yuan Ma

Publications and source records attributed to Yuan Ma.

At least 19 recordsLinked to original sources

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.

cs.CV

Global well-posedness of strong solutions to a model for the morning glory cloud

In this paper, we investigate the global well-posedness of strong solution to a model recently derived by Constantin-Johnson \cite{CaJo} which describes the nonlinear wave propagation in the troposphere, especially for the morning glory cloud. Assuming that the initial velocity $v_0\in H^1$ and the thermodynamic forcing term $K\in L^2(0,T;L^2)$, we show that there exists a unique global strong solution to the initial boundary value problem of this system. Similar result was only known before for sufficiently small initial data in the existing literature.

math.AP

Twist Engineering for Reconfigurable Optical and Optoelectronic Devices

Reconfigurable optical and optoelectronic devices require compact tuning mechanisms capable of reshaping electronic, excitonic, polaritonic, and photonic responses without rebuilding the underlying nanostructure. Against this backdrop, twist has emerged as a powerful geometric degree of freedom that reconfigures interlayer coupling, momentum matching, symmetry, radiation channels, and chiral response by simply rotating adjacent two-dimensional layers or photonic lattices. In this Review, we survey twist-engineered optical and optoelectronic devices spanning van der Waals materials and photonic platforms. We first review the current landscape of twist-angle metrology, classifying existing characterization approaches into three complementary categories: direct structural imaging, methods based on moir\'e periodicity and morphological features, and techniques that infer the twist angle from spectroscopic or electronic responses. We then survey the principal technological routes for twist-angle control, including deterministic transfer and growth strategies, atomic force microscopy (AFM)-assisted manipulation, quantum twisting microscopy (QTM), microelectromechanical systems (MEMS), and emerging non-contact approaches, highlighting their respective capabilities, limitations, and prospects for programmable and scalable moir\'e photonic platforms. Finally, we discuss the future evolution of twist engineering from the fabrication of individual twisted structures toward dynamically reconfigurable, feedback-controlled, and manufacturable photonic systems. We further highlight MEMS-based rotation, piezoelectric actuation, and micro-LiDAR as representative enabling technologies and emerging applications within this broader landscape.

physics.optics

Detecting Avalanche Effect in Adversarial Settings: Spotting the Encryption Loops in Ransomware

Spotting encryption loops in binary-only ransomware is a critical reverse engineering task. Since the existence of avalanche effect, an intrinsic characteristic of any secure encryption algorithms, is unavoidable during a victim data encryption attack, it is a very promising direction to spot encryption loops through avalanche effect detection. Unfortunately, no existing work in this direction ensures that the being-checked effect is the avalanche effect itself. Although CipherXRay is inspired by avalanche effect, it only checks whether a "ripple effect" (i.e., a necessary but non-sufficient condition) of avalanche effect exists, allowing a straightforward counterattack to succeed. In this work, we present a new approach that checks the avalanche effect itself. Because the detection is conducted in adversarial settings (e.g., the ransomware author may obfuscate the code), a viable approach must tolerate inaccurate input \& output identification and must be resilient to adversarial evasion. These challenges are addressed by a novel record-and-replay detection mechanism that takes advantage of the statistical guarantees provided by the Shapiro-Wilk normality test. The experimental results show that our approach achieves 0.0\% false negative rate and 1.1\% false positive rate. When our tool is employed to reverse engineer real-world ransomware samples, it succeeds in analyzing all the ransomware samples selected from ten representative families.

cs.CR

Multibeam Phased Arrays with Spherical Gold Spatio-temporal Coding for Fading-Resilient and Delay Robust Beam Isolations

Future integrated sensing and communication (ISAC) systems require simultaneous multibeam operation with low-latency hardware and robust isolation under synchronization error and fading. Conventional code-division multiplexing using Walsh-Hadamard codes is extremely time-sensitive. This paper demonstrates that conventional temporal-only coded multibeam arrays suffer from inter-beam sidelobe level (SLL) collapse to within a few dB of the main lobe, with variations exceeding 10-20 dB over delay. By embedding moderate-length Gold sequences into a spherical spatial codebook, the proposed Spherical-Gold scheme leverages both temporal and spatial correlation bounds, achieving effective inter-beam isolation without increasing RF complexity. Measurement results and verifications are performed using an Analog Devices ADAR3002 Ka-band 256-element receiver with four simultaneous beams. The proposed scheme demonstrates at least 15 dB rejection with less than 2.5 dB variation in SLL under time error and fading, whereas temporal-only CDMA degrades to approximately -5 to -7 dB SLL with nearly 8 dB variation under time delay.

eess.SP

Probing mesoscopic nonlocal screening in van der Waals heterostructures with polaritons

Predictive optical modelling of van der Waals (vdW) heterostructures is critical for meta-optics, near-field photonics and quantum technologies. At their buried interfaces, charge transfer and spatially extended screening challenge local descriptions based on layer-by-layer stacking of fixed permittivity tensors. However, such nonlocal corrections have been established mainly for plasmonic systems at {\aa}ngstr\"om-nanometre scales and are often assumed negligible on optical-wavelength scales. Here we challenge this view by uncovering a mesoscopic nonlocal screening regime, extending up to ~140 nm, at buried charge-transfer interfaces in transition-metal dichalcogenide/{\alpha}-molybdenum trioxide (TMDC/{\alpha}-MoO3) phonon-polaritonic heterostructures. Using phonon polaritons as an ultrasensitive probe, we quantify charge transfer from polariton-wavelength shifts and find a thickness-independent saturated response as {\alpha}-MoO3 is thinned. Rather than merely complicating optical modelling, this nonlocal saturation turns a design-level correction into an opportunity by yielding a transferable cross-material metric. Across more than 120 devices, this metric scales linearly with the work-function difference between the TMDC and {\alpha}-MoO3. We further identify a lattice-mismatch-set energy threshold for charge transfer, revising Anderson-type band alignment for vdW interfaces.

physics.optics

Isotropic Superconductivity in Room-temperature Superconductor LaSc$_{2}$H$_{24}$

The discovery of LaSc$_{2}$H$_{24}$ represents a milestone in the quest for room-temperature superconductivity, yet the microscopic mechanism underlying its superior performance remains unclear. Through a comprehensive revisit of theoretical calculations, we uncover a pivotal transition from the anisotropic two-gap superconductivity of LaH$_{10}$ to the isotropic single-gap superconductivity in LaSc$_{2}$H$_{24}$ upon the introduction of scandium, thereby enhancing the superconducting critical temperature ($T_\mathrm{c}$). This enhancement is rooted in a critical dual role of Sc $3d$ electrons: i) the Sc-derived Jahn-Teller effect promotes hydrogen metallization via the elongation of specific interlayer H-H bonds and enhances electron-phonon coupling (EPC) through the softening of associated phonon modes; ii) Sc $3d$ electrons reconstruct the electronic structure into an MgB$_{2}$-like configuration, generating novel Sc-H-Sc $\sigma$- and $\pi$-bonding states with EPC strengths comparable to LaH$_{10}$. Crucially, the pronounced hybridization between Sc and the hydrogen cages effectively unifies these two contributions on the Fermi surface. This Sc-induced gap unification bridges the high-EPC H-H states with widespread Sc-H states, establishing an isotropic single-gap nature with a large overall EPC strength. Our findings identify this Sc-induced gap unification as the fundamental mechanism for achieving room-temperature superconductivity in LaSc$_{2}$H$_{24}$, offering a theoretical blueprint for the future design of superior superconducting hydrides.

cond-mat.supr-con

Benchmarking Real-World Medical Image Classification with Noisy Labels: Challenges, Practice, and Outlook

Learning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research on learning with noisy labels (LNL), the robustness of existing methods in medical imaging has not been systematically assessed. To address this gap, we introduce LNMBench, a comprehensive benchmark for Label Noise in Medical imaging. LNMBench encompasses \textbf{10} representative methods evaluated across 7 datasets, 6 imaging modalities, and 3 noise patterns, establishing a unified and reproducible framework for robustness evaluation under realistic conditions. Comprehensive experiments reveal that the performance of existing LNL methods degrades substantially under high and real-world noise, highlighting the persistent challenges of class imbalance and domain variability in medical data. Motivated by these findings, we further propose a simple yet effective improvement to enhance model robustness under such conditions. The LNMBench codebase is publicly released to facilitate standardized evaluation, promote reproducible research, and provide practical insights for developing noise-resilient algorithms in both research and real-world medical applications.The codebase is publicly available on https://github.com/myyy777/LNMBench.

cs.CV

Intense and Tunable Multi-color Terahertz Radiation from Laser-Shaped Electron Beams

High-power multi-color terahertz (THz) radiation exhibits extraordinary scientific application prospects at various scientific frontiers, for its capacity to deliver THz excitation at multiple frequencies simultaneously. However, the generation of high-power multi-color THz radiation with tunable frequencies remains a challenge for existing techniques. Here, a technique by combining the multi-laser pulses frequency beating and coherent undulator amplification is proposed for generating high-power multi-color THz radiation with tunable frequency. Numerical simulations indicate that the proposed technique can produce multi-color THz radiation with three to six distinguished colors and a peak power up to hundreds of MW, and the temporally separated two-color pulses can also be produced by employing undulators with different resonance. Due to the intrinsic properties of the proposed technique, the THz frequencies, the color number and the frequency interval can be effectively controlled by simply adjusting the beating laser. This method paves the way for advanced application of THz pump-THz probe experiments for selective excitation of atomic multi-level systems and molecular fingerprint recognition.

physics.acc-ph

Towards Spectrally Efficient and Physically Reconfigurable Architectures for Multibeam-Waveform Co-Design in Joint Communication and Sensing

Joint Communication and Sensing (JCAS) platforms are emerging as a foundation of next-generation mmWave (MMW) and sub-THz systems, enabling both high-throughput data transfer and angular localization within a shared signal path. This paper investigates multibeam architectures for JCAS that simultaneously optimize waveform shaping and beamforming across the time, frequency, code, and direct analog/ radio frequency (RF) domains. The paper compares Orthogonal Frequency-Division Multiplexing (OFDM), Frequency Modulated Arrays (FMA), Time-Modulated Arrays (TMA), direct RF/MMW modulation, and Code-Division Multiple Access (CDMA)-based systems with respect to spectral efficiency, beam orthogonality, latency, and Angle-of-Arrival (AoA) estimation accuracy. The results highlight architecture-specific tradeoffs among beam agility, efficiency, accuracy and resolution, and complexity. It also provides a framework for selecting JCAS front ends optimized for power, latency, inter-beam and multi-user interference, and rapid system reconfiguration

eess.SP

The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning

We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in executing real-world robotic tasks, their deployment on resource-constrained platforms is often bottlenecked by the heavy attention-based computation over large sets of visual tokens. LightVLA addresses this challenge through adaptive, performance-driven pruning of visual tokens: It generates dynamic queries to evaluate visual token importance, and adopts Gumbel softmax to enable differentiable token selection. Through fine-tuning, LightVLA learns to preserve the most informative visual tokens while pruning tokens which do not contribute to task execution, thereby improving efficiency and performance simultaneously. Notably, LightVLA requires no heuristic magic numbers and introduces no additional trainable parameters, making it compatible with modern inference frameworks. Experimental results demonstrate that LightVLA outperforms different VLA models and existing token pruning methods across diverse tasks on the LIBERO benchmark, achieving higher success rates with substantially reduced computational overhead. Specifically, LightVLA reduces FLOPs and latency by 59.1% and 38.2% respectively, with a 2.6% improvement in task success rate. Meanwhile, we also investigate the learnable query-based token pruning method LightVLA* with additional trainable parameters, which also achieves satisfactory performance. Our work reveals that as VLA pursues optimal performance, LightVLA spontaneously learns to prune tokens from a performance-driven perspective. To the best of our knowledge, LightVLA is the first work to apply adaptive visual token pruning to VLA tasks with the collateral goals of efficiency and performance, marking a significant step toward more efficient, powerful and practical real-time robotic systems.

cs.RO

Role of settling inertial particles in modulating flow structures and drag in Taylor-Couette turbulence

The modulation of drag through dispersed phases in wall turbulence has been a longstanding focus. This study examines the effects of particle Stokes number ($St$) and Froude number ($Fr$) on drag modulation in turbulent Taylor-Couette (TC) flow, using a two-way coupled Eulerian-Lagrangian approach with Reynolds number $Re_i = r_i \omega_i d/\nu$ fixed at 3500. For light particles (small $St$), drag reduction is observed in the TC system, exhibiting a non-monotonic dependence on $Fr$. In specific, drag reduction initially increases and then decreases with stronger influence of gravitational settling (characterized by inverse of $Fr$), indicating the presence of an optimal $Fr$ for maximum drag reduction. For heavy particles, similar non-monotonic trend can also be observed, but significant drag enhancement is resulted at large $Fr^{-1}$. We further elucidate the role of settling particles in modulating the flow structure in TC by decomposing the advective flux into contributions from coherent Taylor vortices and background turbulent fluctuations. At moderate effects of particle inertia and gravitational settling, particles suppress the coherence of Taylor vortices which remarkably reduces angular velocity transport and thus leads to drag reduction. However, with increasing influence of particle inertia and gravitational settling, the flow undergoes abrupt change. Rapidly settling particles disrupt the Taylor vortices, shifting the bulk flow from a vortex-dominated regime to one characterized by particle-induced turbulence. With the dominance by particle-induced turbulence, velocity plumes -- initially transported by small-scale G{\"{o}}rtler vortices near the cylinder wall and large-scale Taylor vortices in bulk region -- are instead carried into the bulk by turbulent fluctuations driven by the settling particles.

physics.flu-dyn

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving

In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabilities to modern end-to-end autonomous driving systems has also emerged as a promising direction. However, existing diffusion-based trajectory generative models often exhibit mode collapse where different random noises converge to similar trajectories after the denoising process.Therefore, state-of-the-art models often rely on anchored trajectories from pre-defined trajectory vocabulary or scene priors in the training set to mitigate collapse and enrich the diversity of generated trajectories, but such inductive bias are not available in real-world deployment, which can be challenged when generalizing to unseen scenarios. In this work, we investigate the possibility of effectively tackling the mode collapse challenge without the assumption of pre-defined trajectory vocabulary or pre-computed scene priors. Specifically, we propose TransDiffuser, an encoder-decoder based generative trajectory planning model, where the encoded scene information and motion states serve as the multi-modal conditional input of the denoising decoder. Different from existing approaches, we exploit a simple yet effective multi-modal representation decorrelation optimization mechanism during the denoising process to enrich the latent representation space which better guides the downstream generation. Without any predefined trajectory anchors or pre-computed scene priors, TransDiffuser achieves the PDMS of 94.85 on the closed-loop planning-oriented benchmark NAVSIM, surpassing previous state-of-the-art methods. Qualitative evaluation further showcases TransDiffuser generates more diverse and plausible trajectories which explore more drivable area.

cs.RO

Implementing Trust in Non-Small Cell Lung Cancer Diagnosis with a Conformalized Uncertainty-Aware AI Framework in Whole-Slide Images

Ensuring trustworthiness is fundamental to the development of artificial intelligence (AI) that is considered societally responsible, particularly in cancer diagnostics, where a misdiagnosis can have dire consequences. Current digital pathology AI models lack systematic solutions to address trustworthiness concerns arising from model limitations and data discrepancies between model deployment and development environments. To address this issue, we developed TRUECAM, a framework designed to ensure both data and model trustworthiness in non-small cell lung cancer subtyping with whole-slide images. TRUECAM integrates 1) a spectral-normalized neural Gaussian process for identifying out-of-scope inputs and 2) an ambiguity-guided elimination of tiles to filter out highly ambiguous regions, addressing data trustworthiness, as well as 3) conformal prediction to ensure controlled error rates. We systematically evaluated the framework across multiple large-scale cancer datasets, leveraging both task-specific and foundation models, illustrate that an AI model wrapped with TRUECAM significantly outperforms models that lack such guidance, in terms of classification accuracy, robustness, interpretability, and data efficiency, while also achieving improvements in fairness. These findings highlight TRUECAM as a versatile wrapper framework for digital pathology AI models with diverse architectural designs, promoting their responsible and effective applications in real-world settings.

eess.IV

Backdoor Attack with Invisible Triggers Based on Model Architecture Modification

Machine learning systems are vulnerable to backdoor attacks, where attackers manipulate model behavior through data tampering or architectural modifications. Traditional backdoor attacks involve injecting malicious samples with specific triggers into the training data, causing the model to produce targeted incorrect outputs in the presence of the corresponding triggers. More sophisticated attacks modify the model's architecture directly, embedding backdoors that are harder to detect as they evade traditional data-based detection methods. However, the drawback of the architectural modification based backdoor attacks is that the trigger must be visible in order to activate the backdoor. To further strengthen the invisibility of the backdoor attacks, a novel backdoor attack method is presented in the paper. To be more specific, this method embeds the backdoor within the model's architecture and has the capability to generate inconspicuous and stealthy triggers. The attack is implemented by modifying pre-trained models, which are then redistributed, thereby posing a potential threat to unsuspecting users. Comprehensive experiments conducted on standard computer vision benchmarks validate the effectiveness of this attack and highlight the stealthiness of its triggers, which remain undetectable through both manual visual inspection and advanced detection tools.

cs.CR

Deep learning and random light structuring ensure robust free-space communications

Having shown early promise, free-space optical communications (FSO) face formidable challenges in the age of information explosion. The ever-growing demand for greater channel communication capacity is one of the challenges. The inter-channel crosstalk, which severely degrades the quality of transmitted information, creates another roadblock in the way of efficient FSO implementation. Here we advance theoretically and realize experimentally a potentially high-capacity FSO protocol that enables high-fidelity transfer of an image, or set of images through a complex environment. In our protocol, we complement random light structuring at the transmitter with a deep learning image classification platform at the receiver. Multiplexing novel, independent, mutually orthogonal degrees of freedom available to structured random light can potentially significantly boost the channel communication capacity of our protocol without introducing any deleterious crosstalk. Specifically, we show how one can multiplex the degrees of freedom associated with the source coherence radius and a spatial position of a beamlet within an array of structured random beams to greatly enhance the capacity of our communication link. The superb resilience of structured random light to environmental noise, as well as extreme efficiency of deep learning networks at classifying images guarantees high-fidelity image transfer within the framework of our protocol.

physics.optics

Self-supervised Learning for Electroencephalogram: A Systematic Survey

Electroencephalogram (EEG) is a non-invasive technique to record bioelectrical signals. Integrating supervised deep learning techniques with EEG signals has recently facilitated automatic analysis across diverse EEG-based tasks. However, the label issues of EEG signals have constrained the development of EEG-based deep models. Obtaining EEG annotations is difficult that requires domain experts to guide collection and labeling, and the variability of EEG signals among different subjects causes significant label shifts. To solve the above challenges, self-supervised learning (SSL) has been proposed to extract representations from unlabeled samples through well-designed pretext tasks. This paper concentrates on integrating SSL frameworks with temporal EEG signals to achieve efficient representation and proposes a systematic review of the SSL for EEG signals. In this paper, 1) we introduce the concept and theory of self-supervised learning and typical SSL frameworks. 2) We provide a comprehensive review of SSL for EEG analysis, including taxonomy, methodology, and technique details of the existing EEG-based SSL frameworks, and discuss the difference between these methods. 3) We investigate the adaptation of the SSL approach to various downstream tasks, including the task description and related benchmark datasets. 4) Finally, we discuss the potential directions for future SSL-EEG research.

eess.SP

Guidelines in Wastewater-based Epidemiology of SARS-CoV-2 with Diagnosis

With the global spread and increasing transmission rate of SARS-CoV-2, more and more laboratories and researchers are turning their attention to wastewater-based epidemiology (WBE), hoping it can become an effective tool for large-scale testing and provide more ac-curate predictions of the number of infected individuals. Based on the cases of sewage sampling and testing in some regions such as Hong Kong, Brazil, and the United States, the feasibility of detecting the novel coronavirus in sewage is extremely high. This study re-views domestic and international achievements in detecting SARS-CoV-2 through WBE and summarizes four aspects of COVID-19, including sampling methods, virus decay rate cal-culation, standardized population coverage of the watershed, algorithm prediction, and provides ideas for combining field modeling with epidemic prevention and control. Moreover, we highlighted some diagnostic techniques for detection of the virus from sew-age sample. Our review is a new approach in identification of the research gaps in waste water-based epidemiology and diagnosis and we also predict the future prospect of our analysis.

q-bio.QM