SearcharxivSearch

arXiv subjects

Yizhou Huang

Publications and source records attributed to Yizhou Huang.

At least 19 recordsLinked to original sources

Can We Trust AI in 6G? Verifiable and Auditable AI-Driven Trustworthy Wireless Networks

Mobile network operators are increasingly exploring the use of artificial intelligence (AI) to automate complex network tasks, such as cell selection and mobility management. A fundamental problem arises: there is currently no way to verify that an AI function is making the right decisions or for the right reasons, rather than arriving at correct-looking answers through unreliable shortcuts. In safety-critical and resilience-focused infrastructure, this lack of transparency poses a significant challenge to the widespread adoption of AI technologies in wireless networks. In this paper, we propose a mechanical auditing approach: inspecting a function's internal representations and checking them against machine-verifiable 3GPP specifications. Specifically, we set out a general three-step auditing principle that locates protocol-relevant features, verifies their causal role, and diagnoses how adaptation reshapes their use, grounding it throughout publicly available interpretability and telecommunications research. We present an audit-native network architecture in which a dedicated verification agent continuously checks the reasoning of AI functions in networks, supporting both predeployment certification and runtime auditing. We also discuss how it could be realised, the data and benchmarks, as well as the open challenges that remain before mechanistic auditing can enter telecommunications practice and standardisation.

cs.NI

From Traditional Automation to Embodied Wireless Intelligence: Vision-Language-Action Empowered Physics-Aware Communication Networks

Wireless network automation has progressed from rule-based self-organising networks (SON) to data-driven optimisation, yet existing systems remain fundamentally disembodied. They act on performance indicators without perceiving the physical environment that governs radio propagation. We propose the embodied intelligent empowered base station (eBS), a paradigm that adopts a Vision-Language-Action (VLA) pipeline to transform base stations into autonomous AI agents capable of situated perception, causal physical reasoning, and physics-aware action generation. The eBS employs a two-tier asynchronous architecture: a Semantic Planner powered by a frontier Vision-Language Model (VLM) generates structured action directives on human timescales, whilst a Tactical Controller executes real-time adaptation. Case studies demonstrate that a single VLA pipeline, without task-specific training, can perform zero-shot material reasoning, generalise across viewpoints, and predict dynamic events before signal degradation occurs. These results illustrate a paradigm shift from traditional rule-following network automation to embodied-intelligence-empowered future wireless networks.

cs.NI

Bridging Microscopic Constructions and Continuum Topological Field Theory of Three-Dimensional Non-Abelian Topological Order

Continuum field-theoretical descriptions of topological order are often constructed at long distances without direct reference to microscopic short-distance realizations, guided instead by general principles such as gauge invariance, locality, symmetry, response, and topological invariance. A classic example is provided by Chern--Simons-type topological field theories for two-dimensional anyon systems. Recently, this framework has been extended to three-dimensional topological orders, where particle and loop excitations exhibit highly nontrivial phenomena, including braiding, fusion, and shrinking. Field-theoretical approaches have further led to diagrammatic representations, pentagon and hexagon relations, and \textit{fusion--shrinking consistency} conditions governing these processes. Despite these advances, a long-standing question remains: do such long-distance field-theoretical structures admit faithful microscopic counterparts with tensor-product local Hilbert spaces and short-range interactions? In this work, we answer this question by establishing an explicit correspondence between continuum topological field theory and microscopic lattice constructions of three-dimensional non-Abelian topological order. While Wilson operators encode long-distance topological excitations, we construct microscopic lattice operators that create, fuse, shrink, and braid particles and loops. Using these operators, we compute fusion and shrinking rules, particle--loop and Borromean-Rings braiding phases, and show how non-Abelian shrinking channels can be selectively controlled by the internal degrees of freedom of loop operators. We further show that the lattice shrinking rules satisfy the \textit{fusion--shrinking consistency} relations previously obtained from field theory, establishing these relations as a microscopically verifiable organizing principle for 3D topological order. Remarkably, by...

cond-mat.str-el

Continual Model-Based Reinforcement Learning with Hypernetworks

Effective planning in model-based reinforcement learning (MBRL) and model-predictive control (MPC) relies on the accuracy of the learned dynamics model. In many instances of MBRL and MPC, this model is assumed to be stationary and is periodically re-trained from scratch on state transition experience collected from the beginning of environment interactions. This implies that the time required to train the dynamics model - and the pause required between plan executions - grows linearly with the size of the collected experience. We argue that this is too slow for lifelong robot learning and propose HyperCRL, a method that continually learns the encountered dynamics in a sequence of tasks using task-conditional hypernetworks. Our method has three main attributes: first, it includes dynamics learning sessions that do not revisit training data from previous tasks, so it only needs to store the most recent fixed-size portion of the state transition experience; second, it uses fixed-capacity hypernetworks to represent non-stationary and task-aware dynamics; third, it outperforms existing continual learning alternatives that rely on fixed-capacity networks, and does competitively with baselines that remember an ever increasing coreset of past experience. We show that HyperCRL is effective in continual model-based reinforcement learning in robot locomotion and manipulation scenarios, such as tasks involving pushing and door opening. Our project website with videos is at this link https://rvl.cs.toronto.edu/blog/hypercrl

cs.LG

BioLLMAgent: A Hybrid Framework with Enhanced Structural Interpretability for Simulating Human Decision-Making in Computational Psychiatry

Computational psychiatry faces a fundamental trade-off: traditional reinforcement learning (RL) models offer interpretability but lack behavioral realism, while large language model (LLM) agents generate realistic behaviors but lack structural interpretability. We introduce BioLLMAgent, a novel hybrid framework that combines validated cognitive models with the generative capabilities of LLMs. The framework comprises three core components: (i) an Internal RL Engine for experience-driven value learning; (ii) an External LLM Shell for high-level cognitive strategies and therapeutic interventions; and (iii) a Decision Fusion Mechanism for integrating components via weighted utility. Comprehensive experiments on the Iowa Gambling Task (IGT) across six clinical and healthy datasets demonstrate that BioLLMAgent accurately reproduces human behavioral patterns while maintaining excellent parameter identifiability (correlations $>0.67$). Furthermore, the framework successfully simulates cognitive behavioral therapy (CBT) principles and reveals, through multi-agent dynamics, that community-wide educational interventions may outperform individual treatments. Validated across reward-punishment learning and temporal discounting tasks, BioLLMAgent provides a structurally interpretable "computational sandbox" for testing mechanistic hypotheses and intervention strategies in psychiatric research.

cs.AI

FoSS: Modeling Long Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier State Space Integration

Accurate trajectory prediction is vital for safe autonomous driving, yet existing approaches struggle to balance modeling power and computational efficiency. Attention-based architectures incur quadratic complexity with increasing agents, while recurrent models struggle to capture long-range dependencies and fine-grained local dynamics. Building upon this, we present FoSS, a dual-branch framework that unifies frequency-domain reasoning with linear-time sequence modeling. The frequency-domain branch performs a discrete Fourier transform to decompose trajectories into amplitude components encoding global intent and phase components capturing local variations, followed by a progressive helix reordering module that preserves spectral order; two selective state-space submodules, Coarse2Fine-SSM and SpecEvolve-SSM, refine spectral features with O(N) complexity. In parallel, a time-domain dynamic selective SSM reconstructs self-attention behavior in linear time to retain long-range temporal context. A cross-attention layer fuses temporal and spectral representations, while learnable queries generate multiple candidate trajectories, and a weighted fusion head expresses motion uncertainty. Experiments on Argoverse 1 and Argoverse 2 benchmarks demonstrate that FoSS achieves state-of-the-art accuracy while reducing computation by 22.5% and parameters by over 40%. Comprehensive ablations confirm the necessity of each component.

cs.CV

Non-invertible symmetries and mixed anomalies from conserved current construction in (3+1)D twisted $BF$ topological quantum field theories

We develop a current-based construction of generalized symmetries in $(3+1)$D twisted $BF$ topological quantum field theories (TQFTs), focusing on intrinsically non-invertible higher-form symmetries and their mixed anomalies. Starting from the equations of motion, we extract conserved currents and exponentiate the corresponding charges to obtain topological symmetry operators. This gives a step-by-step procedure for constructing symmetry operators, fusion, and anomaly diagnostics directly from the continuum action. We focus on twisted $BF$ theories with gauge group $G=\prod_i \mathbb{Z}_{N_i}$ and an $a\wedge a\wedge b$ twist, where $a$'s and $b$ are 1-form and 2-form gauge fields, respectively. These theories realize non-Abelian $(3+1)$D TQFTs supporting Borromean-rings braiding and describe three-dimensional non-Abelian topological orders in condensed matter. For $G=(\mathbb{Z}_2)^3$, a microscopic realization is given by the $\mathbb{D}_4$ Kitaev quantum double model. Two distinct classes of conserved currents emerge: Type-I currents generate invertible higher-form symmetries with group-like fusion, while Type-II currents require additional consistency conditions on gauge-field configurations, leading to intrinsically non-invertible symmetries dressed by projectors. We compute the fusion algebra: invertible operators admit inverses, while non-invertible ones exhibit multi-channel fusion governed by projector fusion. We diagnose mixed anomalies by coupling multiple conserved currents to background gauge fields, revealing two outcomes: anomalies canceled by anomaly inflow from a higher-dimensional theory, and intrinsic gauging obstructions encoded in the $(3+1)$D continuum theory. Overall, our results provide a unified and practical approach for constructing and characterizing higher-form symmetries, which can be extended to more general TQFTs and topological orders.

cond-mat.str-el

Agentic AI Empowered Intent-Based Networking for 6G

The transition towards sixth-generation (6G) wireless networks necessitates autonomous orchestration mechanisms capable of translating high-level operational intents into executable network configurations. Existing approaches to Intent-Based Networking (IBN) rely upon either rule-based systems that struggle with linguistic variation or end-to-end neural models that lack interpretability and fail to enforce operational constraints. This paper presents a hierarchical multi-agent framework where Large Language Model (LLM) based agents autonomously decompose natural language intents, consult domain-specific specialists, and synthesise technically feasible network slice configurations through iterative reasoning-action (ReAct) cycles. The proposed architecture employs an orchestrator agent coordinating two specialist agents, i.e., Radio Access Network (RAN) and Core Network agents, via ReAct-style reasoning, grounded in structured network state representations. Experimental evaluation across diverse benchmark scenarios shows that the proposed system outperforms rule-based systems and direct LLM prompting, with architectural principles applicable to Open RAN (O-RAN) deployments. The results also demonstrate that whilst contemporary LLMs possess general telecommunications knowledge, network automation requires careful prompt engineering to encode context-dependent decision thresholds, advancing autonomous orchestration capabilities for next-generation wireless systems.

cs.AI

Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts

Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing multimodal affordance methods face limitations in extracting useful information, mainly due to simple structural designs, basic fusion methods, and large model parameters, making it difficult to meet the performance requirements for practical deployment. To address these issues, this paper proposes the BiT-Align image-depth-text affordance mapping framework. The framework includes a Bypass Prompt Module (BPM) and a Text Feature Guidance (TFG) attention selection mechanism. BPM integrates the auxiliary modality depth image directly as a prompt to the primary modality RGB image, embedding it into the primary modality encoder without introducing additional encoders. This reduces the model's parameter count and effectively improves functional region localization accuracy. The TFG mechanism guides the selection and enhancement of attention heads in the image encoder using textual features, improving the understanding of affordance characteristics. Experimental results demonstrate that the proposed method achieves significant performance improvements on public AGD20K and HICO-IIF datasets. On the AGD20K dataset, compared with the current state-of-the-art method, we achieve a 6.0% improvement in the KLD metric, while reducing model parameters by 88.8%, demonstrating practical application values. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align.

cs.CV

Learning Velocity and Acceleration: Self-Supervised Motion Consistency for Pedestrian Trajectory Prediction

Understanding human motion is crucial for accurate pedestrian trajectory prediction. Conventional methods typically rely on supervised learning, where ground-truth labels are directly optimized against predicted trajectories. This amplifies the limitations caused by long-tailed data distributions, making it difficult for the model to capture abnormal behaviors. In this work, we propose a self-supervised pedestrian trajectory prediction framework that explicitly models position, velocity, and acceleration. We leverage velocity and acceleration information to enhance position prediction through feature injection and a self-supervised motion consistency mechanism. Our model hierarchically injects velocity features into the position stream. Acceleration features are injected into the velocity stream. This enables the model to predict position, velocity, and acceleration jointly. From the predicted position, we compute corresponding pseudo velocity and acceleration, allowing the model to learn from data-generated pseudo labels and thus achieve self-supervised learning. We further design a motion consistency evaluation strategy grounded in physical principles; it selects the most reasonable predicted motion trend by comparing it with historical dynamics and uses this trend to guide and constrain trajectory generation. We conduct experiments on the ETH-UCY and Stanford Drone datasets, demonstrating that our method achieves state-of-the-art performance on both datasets.

cs.CV

Charge Parity Rates in Transmon Qubits with Different Shunting Capacitors

The presence of non-equilibrium quasiparticles in superconducting resonators and qubits operating at millikelvin temperature has been known for decades. One metric for the number of quasiparticles affecting qubits is the rate of single-electron change in charge on the qubit island ($\textit i.e.$ the charge parity rate). Here, we have utilized a Ramsey-like pulse sequence to monitor changes in the parity states of five transmon qubits. The five qubits have shunting capacitors with two different geometries and fabricated from both Al and Ta. The charge parity rate differs by a factor of two for the two transmon designs studied here but does not depend on the material of the shunting capacitor. The underlying mechanism of the source of parity switching is further investigated in one of the qubit devices by increasing the quasiparticle trapping rate using induced vortices in the electrodes of the device. The charge parity rate exhibited a weak dependence on the quasiparticle trapping rate, indicating that the main source of charge parity events is from the production of quasiparticles across the Josephson junction. To estimate this source of quasiparticle production, we simulate and estimate pair-breaking photon absorption rates for our two qubit geometries and find a similar factor of two in the absorption rate for a background blackbody radiation temperature of $T^*\sim$ 350 mK.

quant-ph

Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSM

Motion prediction is crucial for autonomous driving, as it enables accurate forecasting of future vehicle trajectories based on historical inputs. This paper introduces Trajectory Mamba, a novel efficient trajectory prediction framework based on the selective state-space model (SSM). Conventional attention-based models face the challenge of computational costs that grow quadratically with the number of targets, hindering their application in highly dynamic environments. In response, we leverage the SSM to redesign the self-attention mechanism in the encoder-decoder architecture, thereby achieving linear time complexity. To address the potential reduction in prediction accuracy resulting from modifications to the attention mechanism, we propose a joint polyline encoding strategy to better capture the associations between static and dynamic contexts, ultimately enhancing prediction accuracy. Additionally, to balance prediction accuracy and inference speed, we adopted the decoder that differs entirely from the encoder. Through cross-state space attention, all target agents share the scene context, allowing the SSM to interact with the shared scene representation during decoding, thus inferring different trajectories over the next prediction steps. Our model achieves state-of-the-art results in terms of inference speed and parameter efficiency on both the Argoverse 1 and Argoverse 2 datasets. It demonstrates a four-fold reduction in FLOPs compared to existing methods and reduces parameter count by over 40% while surpassing the performance of the vast majority of previous methods. These findings validate the effectiveness of Trajectory Mamba in trajectory prediction tasks.

cs.CV

Diagrammatics, Pentagon Equations, and Hexagon Equations of Topological Orders with Loop- and Membrane-like Excitations

In spacetime dimensions of 4 (i.e., 3+1) and higher, topological orders exhibit spatially extended excitations like loops and membranes, which support diverse topological data characterizing braiding, fusion, and shrinking processes, despite the absence of anyons. Our understanding of these topological data remains less mature compared to 3D, where anyons have been extensively studied and can be fully described through diagrammatic representations. Inspired by recent advancements in field theory descriptions of higher-dimensional topological orders, this paper systematically constructs diagrammatic representations for 4D and 5D topological orders, generalizable to higher dimensions. We introduce elementary diagrams for fusion and shrinking processes, treating them as vectors in fusion and shrinking spaces, respectively, and build complex diagrams by combining these elementary diagrams. Within these vector spaces, we design unitary operations represented by \(F\)-, \(Δ\)-, and \(Δ^2\)-symbols to transform between different bases. We discover \textit{pentagon equations} and \textit{(hierarchical) shrinking-fusion hexagon equations} that impose constraints on the legitimate forms of these unitary operations. We conjecture that all anomaly-free higher-dimensional topological orders must satisfy these conditions and any violations indicate a quantum anomaly. This work opens promising avenues for future research, including the exploration of diagrammatic representations involving braiding and the study of non-invertible symmetries and symmetry topological field theories in higher spacetime dimensions.

hep-th

Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models

Deep learning architectures with powerful reasoning capabilities have driven significant advancements in autonomous driving technology. Large language models (LLMs) applied in this field can describe driving scenes and behaviors with a level of accuracy similar to human perception, particularly in visual tasks. Meanwhile, the rapid development of edge computing, with its advantage of proximity to data sources, has made edge devices increasingly important in autonomous driving. Edge devices process data locally, reducing transmission delays and bandwidth usage, and achieving faster response times. In this work, we propose a driving behavior narration and reasoning framework that applies LLMs to edge devices. The framework consists of multiple roadside units, with LLMs deployed on each unit. These roadside units collect road data and communicate via 5G NSR/NR networks. Our experiments show that LLMs deployed on edge devices can achieve satisfactory response speeds. Additionally, we propose a prompt strategy to enhance the narration and reasoning performance of the system. This strategy integrates multi-modal information, including environmental, agent, and motion data. Experiments conducted on the OpenDV-Youtube dataset demonstrate that our approach significantly improves performance across both tasks.

cs.AI

Fusion rules and shrinking rules of topological orders in five dimensions

As a series of work about 5D (spacetime) topological orders, here we employ the path-integral formalism of 5D topological quantum field theory (TQFT) established in Zhang and Ye, JHEP 04 (2022) 138 to explore non-Abelian fusion rules, hierarchical shrinking rules and quantum dimensions of particle-like, loop-like and membrane-like topological excitations in 5D topological orders. To illustrate, we focus on a prototypical example of twisted $BF$ theories that comprise the twisted topological terms of the $BBA$ type. First, we classify topological excitations by establishing equivalence classes among all gauge-invariant Wilson operators. Then, we compute fusion rules from the path-integral and find that fusion rules may be non-Abelian; that is, the fusion outcome can be a direct sum of distinct excitations. We further compute shrinking rules. Especially, we discover exotic hierarchical structures hidden in shrinking processes of 5D or higher: a membrane is shrunk into particles and loops, and the loops are subsequently shrunk into a direct sum of particles. We obtain the algebraic structure of shrinking coefficients and fusion coefficients. We compute the quantum dimensions of all excitations and find that sphere-like membranes and torus-like membranes differ not only by their shapes but also by their quantum dimensions. We further study the algebraic structure that determines anomaly-free conditions on fusion coefficients and shrinking coefficients. Besides $BBA$, we explore general properties of all twisted terms in $5$D. Together with braiding statistics reported before, the theoretical progress here paves the way toward characterizing and classifying topological orders in higher dimensions where topological excitations consist of both particles and spatially extended objects.

hep-th

Identification and Mitigation of Conducting Package Losses for Quantum Superconducting Devices

Low-loss superconducting rf devices are required when used for quantum computation. Here, we present a series of measurements and simulations showing that conducting losses in the packaging of our superconducting resonator devices affect the maximum achievable internal quality factors (Qi) for a series of thin-film Al quarter-wave resonators with fundamental resonant frequencies varying between 4.9 and 5.8 GHz. By utilizing resonators with different widths and gaps, different volumes of the stored electromagnetic energy were sampled thus affecting Qi. When the backside of the sapphire substrate of the resonator device is adhered to a Cu package with a conducting silver glue, a monotonic decrease in the maximum achievable Qi is found as the electromagnetic sampling volume is increased. This is a result of induced currents in large surface resistance regions and dissipation underneath the substrate. By placing a hole underneath the substrate and using superconducting material for the package, we decrease the ohmic losses and increase the maximum Qi for the larger size resonators.

quant-ph

Stochastic Planning for ASV Navigation Using Satellite Images

Autonomous surface vessels (ASV) represent a promising technology to automate water-quality monitoring of lakes. In this work, we use satellite images as a coarse map and plan sampling routes for the robot. However, inconsistency between the satellite images and the actual lake, as well as environmental disturbances such as wind, aquatic vegetation, and changing water levels can make it difficult for robots to visit places suggested by the prior map. This paper presents a robust route-planning algorithm that minimizes the expected total travel distance given these environmental disturbances, which induce uncertainties in the map. We verify the efficacy of our algorithm in simulations of over a thousand Canadian lakes and demonstrate an application of our algorithm in a 3.7 km-long real-world robot experiment on a lake in Northern Ontario, Canada. Videos are available on our website https://pcctp.github.io/.

cs.RO

Characterization of Asymmetric Gap-Engineered Josephson Junctions and 3D Transmon Qubits

We have fabricated and characterized asymmetric gap-engineered junctions and transmon devices. To create Josephson junctions with asymmetric gaps, Ti was used to proximitize and lower the superconducting gap of the Al counter-electrode. DC IV measurements of these small, proximitized Josephson junctions show a reduced gap and larger excess current for voltage biases below the superconducting gap when compared to standard Al/AlOx/Al junctions. The energy relaxation time constant for an Al/AlOx/Al/Ti 3D transmon was T1 = 1 μs, over two orders of magnitude shorter than the measured T1 = 134 μs of a standard Al/AlOx/Al 3D transmon. Intentionally adding disorder between the Al and Ti layers reduces the proximity effect and subgap current while increasing the relaxation time to T1 = 32 μs.

quant-ph