SearcharxivSearch

arXiv subjects

Juntao Gao

Publications and source records attributed to Juntao Gao.

14 recordsLinked to original sources

Toward Quantum Advantage in Learning Parities with Structured Noise via Lower Bound Optimization of the Condition Number

Learning Parities with Structured Noise (LPSN) can be reduced to solving nonlinear Boolean systems. In quantum computing, such systems are typically transformed into Macaulay linear systems and solved via quantum linear system algorithms, a process severely limited by the condition number. To address this, we propose a novel reduction method for Macaulay linear systems. Under the assumptions of Ding et al., we derive a condition number lower bound incorporating a scaling factor. This reduction not only guarantees efficient quantum state preparation but also exhibits a distinct advantage regarding the condition number interval relative to the reduced right-hand side vector, thereby reducing the lower bound of the condition number and ultimately optimizing the upper bound on the time complexity of the quantum algorithm for solving Boolean systems. Furthermore, applying this improved quantum algorithm to LPSN significantly reduces sample complexity by exploiting the Macaulay system's solution structure. We further provide a concrete logical-level quantum resource estimate, demonstrating that the optimized condition number translates directly into a reduction in circuit width, depth, and gate count. Finally, we establish an algorithm selection strategy by systematically comparing quantum and classical approaches across noise pattern adaptability, sample complexity, and time complexity. Results demonstrate that our quantum algorithm exhibits the potential to outperform classical counterparts under specific parameter regimes.

cs.CR

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov Decision Process, even though real-world robotic control is inherently partially observable and requires reasoning over past interactions. To address this mismatch, we reformulate VLA policy learning from a Partially Observable Markov Decision Process perspective and propose AVA-VLA, a framework that conditions action generation on a recurrent state that serves as a neural approximation to the agent's belief over task history. Built on this recurrent state, we introduce Active Visual Attention (AVA), which dynamically reweights visual tokens in the current observation to focus on regions most relevant given both the instruction and execution history. Extensive experiments show that AVA-VLA achieves state-of-the-art performance on standard robotic benchmarks, including LIBERO and CALVIN, and transfers effectively to real-world dual-arm manipulation tasks. These results demonstrate the effectiveness of temporally grounded active visual processing for improving VLA performance in robotic sequential decision-making. The project page is available at https://liauto-dsr.github.io/AVA-VLA-Page.

cs.LG

NeuroSymb-MRG: Differentiable Abductive Reasoning with Active Uncertainty Minimization for Radiology Report Generation

Automatic generation of radiology reports seeks to reduce clinician workload while improving documentation consistency. Existing methods that adopt encoder-decoder or retrieval-augmented pipelines achieve progress in fluency but remain vulnerable to visual-linguistic biases, factual inconsistency, and lack of explicit multi-hop clinical reasoning. We present NeuroSymb-MRG, a unified framework that integrates NeuroSymbolic abductive reasoning with active uncertainty minimization to produce structured, clinically grounded reports. The system maps image features to probabilistic clinical concepts, composes differentiable logic-based reasoning chains, decodes those chains into templated clauses, and refines the textual output via retrieval and constrained language-model editing. An active sampling loop driven by rule-level uncertainty and diversity guides clinician-in-the-loop adjudication and promptbook refinement. Experiments on standard benchmarks demonstrate consistent improvements in factual consistency and standard language metrics compared to representative baselines.

cs.CV

Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation

Vision-Language-Action (VLA) models have emerged as a powerful paradigm in Embodied AI. However, the significant computational overhead of processing redundant visual tokens remains a critical bottleneck for real-time robotic deployment. While standard token pruning techniques can alleviate this, these task-agnostic methods struggle to preserve task-critical visual information. To address this challenge, simultaneously preserving both the holistic context and fine-grained details for precise action, we propose Compressor-VLA, a novel hybrid instruction-conditioned token compression framework designed for efficient, task-oriented compression of visual information in VLA models. The proposed Compressor-VLA framework consists of two token compression modules: a Semantic Task Compressor (STC) that distills holistic, task-relevant context, and a Spatial Refinement Compressor (SRC) that preserves fine-grained spatial details. This compression is dynamically modulated by the natural language instruction, allowing for the adaptive condensation of task-relevant visual information. Experimentally, extensive evaluations demonstrate that Compressor-VLA achieves a competitive success rate on the LIBERO benchmark while reducing FLOPs by 59% and the visual token count by over 3x compared to its baseline. The real-robot deployments on a dual-arm robot platform validate the model's sim-to-real transferability and practical applicability. Moreover, qualitative analyses reveal that our instruction guidance dynamically steers the model's perceptual focus toward task-relevant objects, thereby validating the effectiveness of our approach.

cs.RO

An Integrated AI-Enabled System Using One Class Twin Cross Learning (OCT-X) for Early Gastric Cancer Detection

Early detection of gastric cancer, a leading cause of cancer-related mortality worldwide, remains hampered by the limitations of current diagnostic technologies, leading to high rates of misdiagnosis and missed diagnoses. To address these challenges, we propose an integrated system that synergizes advanced hardware and software technologies to balance speed-accuracy. Our study introduces the One Class Twin Cross Learning (OCT-X) algorithm. Leveraging a novel fast double-threshold grid search strategy (FDT-GS) and a patch-based deep fully convolutional network, OCT-X maximizes diagnostic accuracy through real-time data processing and seamless lesion surveillance. The hardware component includes an all-in-one point-of-care testing (POCT) device with high-resolution imaging sensors, real-time data processing, and wireless connectivity, facilitated by the NI CompactDAQ and LabVIEW software. Our integrated system achieved an unprecedented diagnostic accuracy of 99.70%, significantly outperforming existing models by up to 4.47%, and demonstrated a 10% improvement in multirate adaptability. These findings underscore the potential of OCT-X as well as the integrated system in clinical diagnostics, offering a path toward more accurate, efficient, and less invasive early gastric cancer detection. Future research will explore broader applications, further advancing oncological diagnostics. Code is available at https://github.com/liu37972/Multirate-Location-on-OCT-X-Learning.git.

eess.IV

Enhancing Diagnostic Precision in Gastric Bleeding through Automated Lesion Segmentation: A Deep DuS-KFCM Approach

Timely and precise classification and segmentation of gastric bleeding in endoscopic imagery are pivotal for the rapid diagnosis and intervention of gastric complications, which is critical in life-saving medical procedures. Traditional methods grapple with the challenge posed by the indistinguishable intensity values of bleeding tissues adjacent to other gastric structures. Our study seeks to revolutionize this domain by introducing a novel deep learning model, the Dual Spatial Kernelized Constrained Fuzzy C-Means (Deep DuS-KFCM) clustering algorithm. This Hybrid Neuro-Fuzzy system synergizes Neural Networks with Fuzzy Logic to offer a highly precise and efficient identification of bleeding regions. Implementing a two-fold coarse-to-fine strategy for segmentation, this model initially employs the Spatial Kernelized Fuzzy C-Means (SKFCM) algorithm enhanced with spatial intensity profiles and subsequently harnesses the state-of-the-art DeepLabv3+ with ResNet50 architecture to refine the segmentation output. Through extensive experiments across mainstream gastric bleeding and red spots datasets, our Deep DuS-KFCM model demonstrated unprecedented accuracy rates of 87.95%, coupled with a specificity of 96.33%, outperforming contemporary segmentation methods. The findings underscore the model's robustness against noise and its outstanding segmentation capabilities, particularly for identifying subtle bleeding symptoms, thereby presenting a significant leap forward in medical image processing.

eess.IV

Secure Payment System Utilizing MANET for Disaster Areas

Mobile payment system in a disaster area have the potential to provide electronic transactions for people purchasing recovery goods like foodstuffs, clothes, and medicine. Conversely, to enable transactions in a disaster area, current payment systems need communication infrastructures (such as wired networks and cellular networks) which may be ruined during such disasters as large-scale earthquakes and flooding and thus cannot be depended on in a disaster area. In this paper, we introduce a new mobile payment system utilizing infrastructureless MANETs to enable transactions that permit users to shop in disaster areas. Specifically, we introduce an endorsement-based mechanism to provide payment guarantees for a customer-to-merchant transaction and a multilevel endorsement mechanism with a lightweight scheme based on Bloom filter and Merkle tree to reduce communication overheads. Our mobile payment system achieves secure transaction by adopting various schemes such as location-based mutual monitoring scheme and blind signature, while our newly introduce event chain mechanism prevents double spending attacks. As validated by simulations, the proposed mobile payment system is useful in a disaster area, achieving high transaction completion ratio, 65% - 90% for all scenario tested, and is storage-efficient for mobile devices with an overall average of 7MB merchant message size.

cs.DC

Super-resolution Imaging of the Fluorescent Dipole Assembly with Polarized Structured Illumination Microscopy

Fluorescence polarization microscopy images both the intensity and orientation of fluorescent dipoles, which plays a vital role in studying the molecular structure and dynamics of bio-complex. However, it is difficult to resolve the dipole assemblies on the subcellular structure and their dynamics in living cells with super-resolution. Here we report polarized structured illumination microscopy (pSIM), which decouples the entangled spatial and angular structured illumination through interpreting the dipoles in spatio-angular hyperspace. We demonstrate its application on a series of biological filamentous systems such as cytoskeleton networks and lambda-DNA, and report the dynamics of short actin sliding through myosin-coated surface. Further, pSIM reveals "side-by-side" organization of the actin ring structure in the membrane-associated periodic skeleton in hippocampal neurons. It also images the dipole dynamics of green fluorescent proteins labeled to the microtubules in live U2OS cells. pSIM can be applied directly to a large variety of commercial or home-built SIM systems.

physics.optics

Multi-Agent Q-Learning Aided Backpressure Routing Algorithm for Delay Reduction

In queueing networks, it is well known that the throughput-optimal backpressure routing algorithm results in poor delay performance for light and moderate traffic loads. To improve delay performance, state-of-the-art backpressure routing algorithm (called BPmin [1]) exploits queue length information to direct packets to less congested routes to their destinations. However, BPmin algorithm estimates route congestion based on unrealistic assumption that every node in the network knows real-time global queue length information of all other nodes. In this paper, we propose multi-agent Q-learning aided backpressure routing algorithm, where each node estimates route congestion using only local information of neighboring nodes. Our algorithm not only outperforms state-of-the-art BPmin algorithm in delay performance but also retains the following appealing features: distributed implementation, low computation complexity and throughput-optimality. Simulation results show our algorithm reduces average packet delay by 95% for light traffic loads and by 41% for moderate traffic loads when compared to state-of-the-art BPmin algorithm.

cs.NI

Relational Algebra for In-Database Process Mining

The execution logs that are used for process mining in practice are often obtained by querying an operational database and storing the result in a flat file. Consequently, the data processing power of the database system cannot be used anymore for this information, leading to constrained flexibility in the definition of mining patterns and limited execution performance in mining large logs. Enabling process mining directly on a database - instead of via intermediate storage in a flat file - therefore provides additional flexibility and efficiency. To help facilitate this ideal of in-database process mining, this paper formally defines a database operator that extracts the 'directly follows' relation from an operational database. This operator can both be used to do in-database process mining and to flexibly evaluate process mining related queries, such as: "which employee most frequently changes the 'amount' attribute of a case from one task to the next". We define the operator using the well-known relational algebra that forms the formal underpinning of relational databases. We formally prove equivalence properties of the operator that are useful for query optimization and present time-complexity properties of the operator. By doing so this paper formally defines the necessary relational algebraic elements of a 'directly follows' operator, which are required for implementation of such an operator in a DBMS.

cs.DB

Adaptive Traffic Signal Control: Deep Reinforcement Learning Algorithm with Experience Replay and Target Network

Adaptive traffic signal control, which adjusts traffic signal timing according to real-time traffic, has been shown to be an effective method to reduce traffic congestion. Available works on adaptive traffic signal control make responsive traffic signal control decisions based on human-crafted features (e.g. vehicle queue length). However, human-crafted features are abstractions of raw traffic data (e.g., position and speed of vehicles), which ignore some useful traffic information and lead to suboptimal traffic signal controls. In this paper, we propose a deep reinforcement learning algorithm that automatically extracts all useful features (machine-crafted features) from raw real-time traffic data and learns the optimal policy for adaptive traffic signal control. To improve algorithm stability, we adopt experience replay and target network mechanisms. Simulation results show that our algorithm reduces vehicle delay by up to 47% and 86% when compared to another two popular traffic signal control algorithms, longest queue first algorithm and fixed time control algorithm, respectively.

cs.NI

Optimal Scheduling for Incentive WiFi Offloading under Energy Constraint

WiFi offloading, where mobile device users (e.g., smart phone users) transmit packets through WiFi networks rather than cellular networks, is a promising solution to alleviating the heavy traffic burden of cellular networks due to data explosion. However, since WiFi networks are intermittently available, a mobile device user in WiFi offloading usually needs to wait for WiFi connection and thus experiences longer delay of packet transmission. To motivate users to participate in WiFi offloading, cellular network operators give incentives (rewards like coupons, e-coins) to users who wait for WiFi connection and transmit packets through WiFi networks. In this paper, we aim at maximizing users' rewards while meeting constraints on queue stability and energy consumption. However, we face scheduling challenges from random packet arrivals, intermittent WiFi connection and time varying wireless link states. To address these challenges, we first formulate the problem as a stochastic optimization problem. We then propose an optimal scheduling policy, named Optimal scheduling Policy under Energy Constraint (OPEC), which makes online decisions as to when to delay packet transmission to wait for WiFi connection and which wireless link (WiFi link or cellular link) to transmit packets on. OPEC automatically adapts to random packet arrivals and time varying wireless link states, not requiring a priori knowledge of packet arrival and wireless link probabilities. As verified by simulations, OPEC scheduling policy can achieve the maximum rewards while keeping queue stable and meeting energy consumption constraint.

cs.NI

End-to-End Delay Modeling for Mobile Ad Hoc Networks: A Quasi-Birth-and-Death Approach

Understanding the fundamental end-to-end delay performance in mobile ad hoc networks (MANETs) is of great importance for supporting Quality of Service (QoS) guaranteed applications in such networks. While upper bounds and approximations for end-to-end delay in MANETs have been developed in literature, which usually introduce significant errors in delay analysis, the modeling of exact end-to-end delay in MANETs remains a technical challenge. This is partially due to the highly dynamical behaviors of MANETs, but also due to the lack of an efficient theoretical framework to capture such dynamics. This paper demonstrates the potential application of the powerful Quasi-Birth-and-Death (QBD) theory in tackling the challenging issue of exact end-to-end delay modeling in MANETs. We first apply the QBD theory to develop an efficient theoretical framework for capturing the complex dynamics in MANETs. We then show that with the help of this framework, closed form models can be derived for the analysis of exact end-to-end delay and also per node throughput capacity in MANETs. Simulation and numerical results are further provided to illustrate the efficiency of these QBD theory-based models as well as our theoretical findings.

cs.NI

Source Delay in Mobile Ad Hoc Networks

Source delay, the time a packet experiences in its source node, serves as a fundamental quantity for delay performance analysis in networks. However, the source delay performance in highly dynamic mobile ad hoc networks (MANETs) is still largely unknown by now. This paper studies the source delay in MANETs based on a general packet dispatching scheme with dispatch limit $f$ (PD-$f$ for short), where a same packet will be dispatched out up to $f$ times by its source node such that packet dispatching process can be flexibly controlled through a proper setting of $f$. We first apply the Quasi-Birth-and-Death (QBD) theory to develop a theoretical framework to capture the complex packet dispatching process in PD-$f$ MANETs. With the help of the theoretical framework, we then derive the cumulative distribution function as well as mean and variance of the source delay in such networks. Finally, extensive simulation and theoretical results are provided to validate our source delay analysis and illustrate how source delay in MANETs are related to network parameters.

cs.NI