SearcharxivSearch

arXiv subjects

Cong Zhao

Publications and source records attributed to Cong Zhao.

At least 19 recordsLinked to original sources

NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.

cs.RO

Driver2Map: Imitating Human Driving for Online High-Definition Map Construction

High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective differences among these data sources, existing methods often struggle to effectively align and fuse them, making online HD map construction still challenging. To address these issues, we propose Driver2Map, an online HD map construction model inspired by human drivers. Unlike existing HD map construction models that utilize only two modalities, our Driver2Map can simultaneously exploit three modalities. Specifically, we propose a "two-stage alignment" strategy to reduce spatial misalignment across different modalities. Additionally, we introduce "Pose-Guided BEV Fusion", a BEV (bird's-eye-view) generation module that leverages camera pose information to adaptively weight multi-view features, thereby effectively suppressing cross-view feature overlap during BEV generation. Also, we design a "Pretrained Prior for Map Refinement" module to refine the initial prediction by learning map structure priors, thus improving the HD map prediction under dynamic occlusions. Extensive experiments demonstrate that Driver2Map outperforms existing methods on both IoU and AP metrics.

cs.CV

ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance

Reflection symmetry detection remains challenging due to interference from asymmetric regions and arbitrary orientations of symmetric patterns. Asymmetric regions introduce background clutter that disrupts symmetric pattern matching, whereas conventional convolutional neural networks lack rotation equivariance, leading to inconsistent feature representations under rotational transformations. To address these issues, we propose an Asymmetric Region Denoising (ARD) module and a Rotation Equivariant Feature Similarity Matching (REFSM) module. The ARD module suppresses asymmetric interference to refine symmetric patterns, while the REFSM module enhances rotation equivariance through feature similarity matching between original and rotated images. Specifically, our dual-input REFSM framework leverages rotation loss to maximize consistency between the score maps of original and rotated images, thereby enabling precise prediction of rotation-equivariant symmetry axes. Furthermore, we introduce GMSYM, a new benchmark dataset that categorizes images into diverse scenarios and incorporates various interferences to address the limitations of existing reflection symmetry detection benchmarks. Extensive experiments on four standard datasets (DENDI, NYU, LDRS, SDRW) and our proposed GMSYM dataset demonstrate that our method achieves state-of-the-art performance in both accuracy and robustness.

cs.CV

Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data

Robotics faces a fundamental challenge of data scarcity. Unlike language or vision research, there is no internet-scale dataset for robotic manipulation. A promising path forward is to leverage egocentric human data, which can be collected more easily, with greater breadth, and at a larger scale. Towards this end, we investigate key design choices for learning across human and humanoid embodiments equipped with dexterous five-finger hands, using the $\pi_{0.5}$ model as a foundation. Our results show that human data enables robots to learn new task semantics and compose existing skills into novel behaviors without corresponding robot data. The paper website is here: https://egopipaper.github.io/

cs.RO

Discovering Multiscale Deep Formulas in Complex Systems via Neural-Guided Lambda Calculus

A fundamental problem in science is identifying underlying patterns of complex systems in the form of concise mathematical formulas. Current Artificial Intelligence (AI)-based methods have shown strong performance in single-scale systems, yet remain limited in identifying scale-specific formulas in multiscale complex systems. We present Deflex, an end-to-end AI method to automatically extract multiscale formulas with potentially different forms, including invariants and distributions, from complex systems. Deflex consists of two subsystems named Deflexformer and Deflexpressor. Deflexpressor is a lambda-calculus symbolic regression model for higher-order formulas. Deflexformer is a decomposable deep energy model for learning unified representations across scales. Deflexpressor generates synthetic data to pre-train Deflexformer, which then guides formula discovery by decoupling multiscale latent relationships. Across six representative complex systems with diverse behaviors, Deflex achieves up to 7-fold higher efficiency than the state-of-the-art methods while enabling automated multiscale discovery. Our work could be a useful tool for scientific discovery across disciplines.

cs.LG

2D quantum-path interference in high-harmonic generation driven by highly-bichromatic fields

We experimentally observe a new type of quantum-path interference, in two-dimensions (2D-QPI), in high-harmonic generation (HHG) driven by an orthogonally-polarised highly-bichromatic field. This regime is marked by comparable intensities of the two orthogonal colours. In this highly-bichromatic regime, we demonstrate that 2D-QPI is encoded in the measured harmonic intensity modulations with respect to the relative phase of the two-colour field. The modulations of the odd-order harmonics show a monomodal behaviour, whereas the even harmonics are modulated in a bimodal structure. Our calculations using the strong-field approximation and saddle-point method disentangle contributions from multiple quantum orbits in this HHG regime, revealing that the dipole response for both odd and even harmonics inherits the dynamic symmetry of the orthogonally-polarised driving field. This new type of 2D-QPI offers a novel route to HHG spectroscopy of attosecond electron dynamics by lifting up the dimensionality of the quantum paths involved in the interference.

quant-ph

Self-attention enabled quantum path analysis of high-harmonic generation in solids

High-harmonic generation (HHG) in solids provides a powerful platform to probe ultrafast electron dynamics and interband--intraband coupling. However, disentangling the complex many-body contributions in the HHG spectrum remains challenging. Here we introduce a machine-learning approach based on a Transformer encoder to analyze and reconstruct HHG signals computed from a one-dimensional Kronig--Penney model. The self-attention mechanism inherently highlights correlations between temporal dipole dynamics and high-frequency spectral components, allowing us to identify signatures of nonadiabatic band coupling that are otherwise obscured in standard Fourier analysis. By combining attention maps with Gabor time--frequency analysis, we extract and amplify weak coupling channels that contribute to even-order harmonics and anomalous spectral features. Our results demonstrate that multi-head self-attention acts as a selective filter for strong-coupling events in the time domain, enabling a physics-informed interpretation of high-dimensional quantum dynamics. This work establishes Transformer-based attention as a versatile tool for solid-state strong-field physics, opening new possibilities for interpretable machine learning in attosecond spectroscopy and nonlinear photonics.

cond-mat.mtrl-sci

Determination of the absolute energy scale of the DAMPE calorimeter with the geomagnetic rigidity cutoff method

The Dark Matter Particle Explorer (DAMPE) is a satellite-borne detector designed to detect high-energy cosmic ray particles with its core component being a BGO calorimeter capable of measuring energies from $\sim$GeV to $O(100)$ TeV. The 32 radiation lengths thickness of the calorimeter is designed to ensure full containment of showers produced by cosmic ray electrons and positrons (CREs) and $\gamma$-rays at energies below tens of TeV, providing high resolution in energy measurements. The absolute energy scale therefore becomes a crucial parameter for precise measurements of the CRE energy spectrum. The geomagnetic field induces a rapid drop in the low energy spectrum of electrons and positrons, a phenomenon that provides a method to determine the calorimeter's absolute energy scale. By comparing the cutoff energies of the measured spectra of CREs with those expected from the International Geomagnetic Reference Field model across 4 McIlwain $L$ bins - which cover most regions of the DAMPE orbit - we find that the calorimeter's absolute energy scale exceeds the calibration based on Geant4 simulation by $1.013\pm0.012_{\rm stat}\pm0.026_{\rm sys}$ for energies between 7 GeV and 16 GeV. The absolute energy scale should be taken into account when comparing the absolute CREs fluxes among different detectors.

hep-ex

Trace-anomaly-subtracted $\sigma$-mass for Heavy Quarks and Matching from Lattice QCD

We demonstrate that the leading IR-renormalon divergence in the perturbative pole mass of a massive quark resides entirely in the contribution from the trace anomaly of the energy-momentum tensor in QCD. Consequently, the recently proposed trace-anomaly-subtracted $\sigma$-mass definition for heavy quarks is not only scheme- and scale-invariant, hence unique for each flavor, but also free from the leading IR-renormalon ambiguity. We further derive a formula connecting this $\sigma$-mass to the perturbative pole mass, solely in terms of the QCD $\beta$-function, quark-mass anomalous dimension $\gamma_m$ and a proper rewritten form of the pole-to-$\overline{\mathrm{MS}}$ mass conversion factor. Utilizing this formula together with the ingredients available in the literature, we present the explicit five-loop result for the perturbative relationship between the $\sigma$-mass and the perturbative pole mass in QCD under the approximation of keeping only a single quark massive. Furthermore, we outline how $\sigma$-mass for a heavy quark can in principle be determined from lattice QCD through an intermediate regularization-independent renormalization and continuum perturbative matching, for which we analyze the matching factor through four loops with appealing features revealed. Given the aforementioned theoretical merits of this mass definition and the availability of high-precision conversion relations as well as the feasibility of direct extraction from lattice QCD, it is well suited for acting as a promising candidate reference scheme for comparing or combining quark masses of the same flavor obtained in different schemes.

hep-ph

Generative Model-Based Feature Attention Module for Video Action Analysis

Video action analysis is a foundational technology within the realm of intelligent video comprehension, particularly concerning its application in Internet of Things(IoT). However, existing methodologies overlook feature semantics in feature extraction and focus on optimizing action proposals, thus these solutions are unsuitable for widespread adoption in high-performance IoT applications due to the limitations in precision, such as autonomous driving, which necessitate robust and scalable intelligent video analytics analysis. To address this issue, we propose a novel generative attention-based model to learn the relation of feature semantics. Specifically, by leveraging the differences of actions' foreground and background, our model simultaneously learns the frame- and segment-dependencies of temporal action feature semantics, which takes advantage of feature semantics in the feature extraction effectively. To evaluate the effectiveness of our model, we conduct extensive experiments on two benchmark video task, action recognition and action detection. In the context of action detection tasks, we substantiate the superiority of our approach through comprehensive validation on widely recognized datasets. Moreover, we extend the validation of the effectiveness of our proposed method to a broader task, video action recognition. Our code is available at https://github.com/Generative-Feature-Model/GAF.

cs.CV

Floquet-engineering unveiled by high-harmonic generation

Ultrafast optical control of solids has uncovered new phenomena and advanced non-equilibrium condensed matter physics, where photon dressed electronic states - Floquet Bloch states (FBSs) - emerge under a strong oscillating laser field, also known as Floquet engineering. Although FBSs have been extensively investigated using time and angle resolved photoemission spectroscopy, direct evidence of their role in high-harmonic generation spectroscopy (HHGS) has remained elusive. Here, we present combined experimental and theoretical evidence that FBSs can be probed by HHG emission in the wide-bandgap solid magnesium oxide (MgO) driven by few cycle near infrared pulses. Experimentally, we observe clear evidence of FBSs in the HHG yield dependence on the crystal orientation. This specific feature is attributed to nonadiabatic coupling between FBSs and conduction bands near the Brillouin zone edge, where the strong laser field transiently breaks time reversal symmetry. We have confronted the experimental findings with numerical solutions of the time dependent Schr\"odinger equation, which reproduce the new feature and confirm its Floquet origin. The theoretical results show a coupling inducing a local band structure renormalization and Floquet like hybridization under strong field excitation. It also shows that FBS nonadiabatic dynamics persist in the strong field regime, establishing HHGS as a powerful probe of ultrafast light induced band hybridization in solids.

quant-ph

Critical assessment of contact resistance and mobility in tin perovskite semiconductors

Recent reports highlight the potential of tin-based perovskite semiconductors for high-performance p-type field-effect transistors (FETs) with mobilities exceeding 20 cm2V-1s-1. However, these high mobilities--often obtained via two-probe (2P) methods on devices with small channel length-to-width ratios (L/W < 0.5) operating in the saturation regime at high drain-source currents--raise concerns about overestimation due to contact resistance and non-ideal FET characteristics. Here, we performed gated four-point probe (4PP) FET measurements on Hall bar devices (L/W = 6) of Cs0.15FA0.85SnI3, obtaining a consistent mobility of 3.3 cm2V-1s-1. Upon comparing these with gated 2P measurements of narrow-channel FETs (L/W = 0.1) on the same chip, we resolved the contact resistance (R_C). The 2P linear mobility is underestimated due to voltage drops across R_C, while the 2P saturation mobility is overestimated because of high (dR_C)/(dV_G) near the threshold. Contact resistance effects become more pronounced at lower temperatures. Contact-corrected four-point-probe (4PP) mobilities are independent of bias conditions and are observed to flatten at temperatures lower than 180 K. Future reports of perovskite FET mobilities should include gated 4PP measurements and use devices with larger L/W ratios to minimize nonidealities arising from contact resistance effects.

physics.app-ph

TelOps: AI-driven Operations and Maintenance for Telecommunication Networks

Telecommunication Networks (TNs) have become the most important infrastructure for data communications over the last century. Operations and maintenance (O&M) is extremely important to ensure the availability, effectiveness, and efficiency of TN communications. Different from the popular O&M technique for IT systems (e.g., the cloud), artificial intelligence for IT Operations (AIOps), O&M for TNs meets the following three fundamental challenges: topological dependence of network components, highly heterogeneous software, and restricted failure data. This article presents TelOps, the first AI-driven O&M framework for TNs, systematically enhanced with mechanism, data, and empirical knowledge. We provide a comprehensive comparison between TelOps and AIOps, and conduct a proof-of-concept case study on a typical O&M task (failure diagnosis) for a real industrial TN. As the first systematic AI-driven O&M framework for TNs, TelOps opens a new door to applying AI techniques to TN automation.

cs.AI

EdgeSync: Faster Edge-model Updating via Adaptive Continuous Learning for Video Data Drift

Real-time video analytics systems typically place models with fewer weights on edge devices to reduce latency. The distribution of video content features may change over time for various reasons (i.e. light and weather change) , leading to accuracy degradation of existing models, to solve this problem, recent work proposes a framework that uses a remote server to continually train and adapt the lightweight model at edge with the help of complex model. However, existing analytics approaches leave two challenges untouched: firstly, retraining task is compute-intensive, resulting in large model update delays; secondly, new model may not fit well enough with the data distribution of the current video stream. To address these challenges, in this paper, we present EdgeSync, EdgeSync filters the samples by considering both timeliness and inference results to make training samples more relevant to the current video content as well as reduce the update delay, to improve the quality of training, EdgeSync also designs a training management module that can efficiently adjusts the model training time and training order on the runtime. By evaluating real datasets with complex scenes, our method improves about 3.4% compared to existing methods and about 10% compared to traditional means.

cs.CV

Belt and Braces: When Federated Learning Meets Differential Privacy

Federated learning (FL) has great potential for large-scale machine learning (ML) without exposing raw data.Differential privacy (DP) is the de facto standard of privacy protection with provable guarantees.Advances in ML suggest that DP would be a perfect fit for FL with comprehensive privacy preservation. Hence, extensive efforts have been devoted to achieving practically usable FL with DP, which however is still challenging.Practitioners often not only are not fully aware of its development and categorization, but also face a hard choice between privacy and utility. Therefore, it calls for a holistic review of current advances and an investigation on the challenges and opportunities for highly usable FL systems with a DP guarantee. In this article, we first introduce the primary concepts of FL and DP, and highlight the benefits of integration. We then review the current developments by categorizing different paradigms and notions. Aiming at usable FL with DP, we present the optimization principles to seek a better tradeoff between model utility and privacy loss. Finally, we discuss future challenges in the emergent areas and relevant research topics.

cs.CR

FedLED: Label-Free Equipment Fault Diagnosis with Vertical Federated Transfer Learning

Intelligent equipment fault diagnosis based on Federated Transfer Learning (FTL) attracts considerable attention from both academia and industry. It allows real-world industrial agents with limited samples to construct a fault diagnosis model without jeopardizing their raw data privacy. Existing approaches, however, can neither address the intense sample heterogeneity caused by different working conditions of practical agents, nor the extreme fault label scarcity, even zero, of newly deployed equipment. To address these issues, we present FedLED, the first unsupervised vertical FTL equipment fault diagnosis method, where knowledge of the unlabeled target domain is further exploited for effective unsupervised model transfer. Results of extensive experiments using data of real equipment monitoring demonstrate that FedLED obviously outperforms SOTA approaches in terms of both diagnosis accuracy (up to 4.13 times) and generality. We expect our work to inspire further study on label-free equipment fault diagnosis systematically enhanced by target domain knowledge.

cs.LG

Generative Model-based Feature Knowledge Distillation for Action Recognition

Knowledge distillation (KD), a technique widely employed in computer vision, has emerged as a de facto standard for improving the performance of small neural networks. However, prevailing KD-based approaches in video tasks primarily focus on designing loss functions and fusing cross-modal information. This overlooks the spatial-temporal feature semantics, resulting in limited advancements in model compression. Addressing this gap, our paper introduces an innovative knowledge distillation framework, with the generative model for training a lightweight student model. In particular, the framework is organized into two steps: the initial phase is Feature Representation, wherein a generative model-based attention module is trained to represent feature semantics; Subsequently, the Generative-based Feature Distillation phase encompasses both Generative Distillation and Attention Distillation, with the objective of transferring attention-based feature semantics with the generative model. The efficacy of our approach is demonstrated through comprehensive experiments on diverse popular datasets, proving considerable enhancements in video action recognition task. Moreover, the effectiveness of our proposed framework is validated in the context of more intricate video action detection task. Our code is available at https://github.com/aaai-24/Generative-based-KD.

cs.CV

Controlled Randomness Improves the Performance of Transformer Models

During the pre-training step of natural language models, the main objective is to learn a general representation of the pre-training dataset, usually requiring large amounts of textual data to capture the complexity and diversity of natural language. Contrasting this, in most cases, the size of the data available to solve the specific downstream task is often dwarfed by the aforementioned pre-training dataset, especially in domains where data is scarce. We introduce controlled randomness, i.e. noise, into the training process to improve fine-tuning language models and explore the performance of targeted noise in addition to the parameters of these models. We find that adding such noise can improve the performance in our two downstream tasks of joint named entity recognition and relation extraction and text summarization.

cs.CL