SearcharxivSearch

arXiv subjects

Chuang Zhang

Publications and source records attributed to Chuang Zhang.

At least 19 recordsLinked to original sources

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.

cs.CV

Monte Carlo Physics-informed Neural Networks for Inverse Multiscale Heat Conduction Problems via the Phonon Boltzmann Transport Equation

Inferring thermal fields and thermophysical properties from limited measurements is a fundamental challenge in micro- and nanoscale heat conduction, where the classical Fourier law breaks down and the phonon Boltzmann transport equation (BTE) is needed to capture non-diffusive transport effects. In this work, we extend Monte Carlo physics-informed neural networks (MC-PINNs), originally developed for forward phonon BTE problems [J. Comput. Phys. 542, 114364, 2025], to inverse multiscale heat conduction problems. Two representative classes of inverse problems are considered: (i) reconstructing the full thermal field from sparse interior temperature measurements when boundary conditions are unknown, and (ii) simultaneously inferring the unknown relaxation time together with the thermal field. Problem-specific MC-PINN architectures and training strategies are designed for each class. The mesh-free Monte Carlo sampling strategy enables a unified treatment across diffusive, transitional, and ballistic transport regimes without requiring a priori knowledge of the relaxation time. The proposed method is evaluated on quasi-one-dimensional, quasi-two-dimensional, and three-dimensional benchmark problems covering a wide range of Knudsen numbers, as well as on a realistic 3D fin field-effect transistor (FinFET) structure. Results demonstrate that MC-PINNs consistently outperform purely data-driven deep neural networks, particularly in the sparse-data regime, and can accurately infer spatially uniform relaxation times. For spatially varying relaxation times, the inferred distributions capture the dominant thermal response, and numerical simulations using the recovered parameters reproduce the macroscopic fields with good accuracy. These findings establish MC-PINNs as an effective and physically consistent framework for inverse thermal analysis at micro- and nanoscales.

physics.comp-ph

Kinetic simulation of magnetic-field-tuned hydrodynamic electron transport in graphene corbino disk

Hydrodynamic electron transport, in which electrical transport in solids resembles fluid hydrodynamics when momentum-conserving electron-electron scattering dominates, has attracted much attention over the past decade. However, its thermal aspects have received considerably less attention. In this paper, electron transport in a graphene Corbino disk is systematically simulated by solving the stationary Boltzmann transport equation with a dual-relaxation-time Callaway model, where momentum-conserving and momentum-relaxing scatterings are explicitly distinguished. By varying the magnetic field intensity and the scattering rates, the electric charge and heat flux responses are compared across the diffusive-to-hydrodynamic crossover under electric-field or temperature-gradient drives. It is shown that magnetic-field-induced deflection of both fluxes is strongly enhanced in the hydrodynamic regime but nearly suppressed in the diffusive regime. Under electric-field driving, a pronounced temperature rise is observed in the hydrodynamic regime due to reduced dissipation, while the diffusive regime remains nearly isothermal. Under temperature-gradient driving, the deflection is reversed relative to the electric-field case. These findings establish that thermal behaviors could provide a sensitive and independent diagnostic of electron hydrodynamics, with the magnetic field being identified as an effective discriminator between collective and dissipative conduction.

cond-mat.mes-hall

Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception

Cooperative 3D perception via Vehicle-to-Everything communication is a promising paradigm for enhancing autonomous driving, offering extended sensing horizons and occlusion resolution. However, the practical deployment of existing methods is hindered at long distances by two critical bottlenecks: the quadratic computational scaling of dense BEV representations and the fragility of feature association mechanisms under significant observation and alignment errors. To overcome these limitations, we introduce Long-SCOPE, a fully sparse framework designed for robust long-distance cooperative 3D perception. Our method features two novel components: a Geometry-guided Query Generation module to accurately detect small, distant objects, and a learnable Context-Aware Association module that robustly matches cooperative queries despite severe positional noise. Experiments on the V2X-Seq and Griffin datasets validate that Long-SCOPE achieves state-of-the-art performance, particularly in challenging 100-150 m long-range settings, while maintaining highly competitive computation and communication costs.

cs.CV

Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning

Large language models (LLMs) demonstrate superior reasoning capabilities compared to small language models (SLMs), but incur substantially higher costs. We propose COllaborative REAsoner (COREA), a system that cascades an SLM with an LLM to achieve a balance between accuracy and cost in complex reasoning tasks. COREA first attempts to answer questions using the SLM, which outputs both an answer and a verbalized confidence score. Questions with confidence below a predefined threshold are deferred to the LLM for more accurate resolution. We introduce a reinforcement learning-based training algorithm that aligns the SLM's confidence through an additional confidence calibration reward. Extensive experiments demonstrate that our method jointly improves the SLM's reasoning ability and confidence calibration across diverse datasets and model backbones. Compared to using the LLM alone, COREA reduces cost by 21.5% and 16.8% on out-of-domain math and non-math datasets, respectively, with only an absolute pass@1 drop within 2%.

cs.CL

Semi-implicit Lax-Wendroff kinetic scheme for electron-phonon coupling

A semi-implicit Lax-Wendroff scheme is developed for electron-phonon coupling process in metals based on the two-temperature kinetic equations. The core of this method is to integrate the evolution information of physical equations into the numerical modeling process, which leads to that the time step or cell size is not limited by the relaxation time and mean free path. Specifically, the finite difference method is used to solve the kinetic model again when reconstructing the interfacial distribution function, through which the particle migration, scattering and electron-phonon coupling processes are coupled together within a single time step. Numerical tests demonstrate that this method could efficiently capture electron-phonon coupling or heat conduction processes from the ballistic to diffusive regimes. It provides a new tool for describing electron-phonon coupling or thermal management in microelectronic devices.

physics.comp-ph

Semi-implicit Lax-Wendroff kinetic scheme for hydrodynamic phonon transport

A semi-implicit Lax-Wendroff kinetic scheme is developed for hydrodynamic phonon transport in solid materials based on the Boltzmann transport equation under the double relaxation time approximation, in which both the normal and resistive scattering processes are accounted. The trapezoidal and midpoint rules are adopted for the temporal integration of the scattering and migration terms under the framework of finite volume method, respectively. Instead of direct numerical interpolation, the kinetic equation is solved again when reconstructing the interfacial flux, in order to realize the coupling of phonon migration and scattering within a numerical time step. Specifically, the finite difference scheme is introduced and the second-order upwind or central schemes are used for the reconstruction of the interfacial distribution function and its spatial gradient. Consequently, the cell size and time step of the present method could be larger than the phonon mean free path and relaxation time in the limit of small Knudsen numbers. Numerical tests demonstrate that the present method can accurately capture multi-scale thermal conduction phenomena within different normal or resistive scattering rates.

physics.comp-ph

Agentic AI for ISAC: Analysis, Framework, and Case Study

Integrated sensing and communication (ISAC) has emerged as a key development direction in the sixth-generation (6G) era, which provides essential support for the collaborative sensing and communication of future intelligent networks. However, as wireless environments become increasingly dynamic and complex, ISAC systems require more intelligent processing and more autonomous operation to maintain efficiency and adaptability. Meanwhile, agentic artificial intelligence (AI) offers a feasible solution to address these challenges by enabling continuous perception-reasoning-action loops in dynamic environments to support intelligent, autonomous, and efficient operation for ISAC systems. As such, we delve into the application value and prospects of agentic AI in ISAC systems in this work. Firstly, we provide a comprehensive review of agentic AI and ISAC systems to demonstrate their key characteristics. Secondly, we show several common optimization approaches for ISAC systems and highlight the significant advantages of generative artificial intelligence (GenAI)-based agentic AI. Thirdly, we propose a novel agentic ISAC framework and prensent a case study to verify its superiority in optimizing ISAC performance. Finally, we clarify future research directions for agentic AI-based ISAC systems.

cs.AI

SparseCoop: Cooperative Perception with Kinematic-Grounded Queries

Cooperative perception is critical for autonomous driving, overcoming the inherent limitations of a single vehicle, such as occlusions and constrained fields-of-view. However, current approaches sharing dense Bird's-Eye-View (BEV) features are constrained by quadratically-scaling communication costs and the lack of flexibility and interpretability for precise alignment across asynchronous or disparate viewpoints. While emerging sparse query-based methods offer an alternative, they often suffer from inadequate geometric representations, suboptimal fusion strategies, and training instability. In this paper, we propose SparseCoop, a fully sparse cooperative perception framework for 3D detection and tracking that completely discards intermediate BEV representations. Our framework features a trio of innovations: a kinematic-grounded instance query that uses an explicit state vector with 3D geometry and velocity for precise spatio-temporal alignment; a coarse-to-fine aggregation module for robust fusion; and a cooperative instance denoising task to accelerate and stabilize training. Experiments on V2X-Seq and Griffin datasets show SparseCoop achieves state-of-the-art performance. Notably, it delivers this with superior computational efficiency, low transmission cost, and strong robustness to communication latency. Code is available at https://github.com/wang-jh18-SVM/SparseCoop.

cs.CV

Extended Multi-Temperature Model for Electron--Phonon Coupling and Ultrafast Thermal Transport in Graphene

Ultrafast thermal transport in low-dimensional materials challenges traditional diffusive models due to reduced scattering, strong electron-phonon coupling, and pronounced non-equilibrium effects. To address these complexities, we extend the macroscopic multi-temperature model by incorporating non-diffusive and non-local phenomena, treating electrons, optical phonons, and acoustic phonons as coupled but thermally distinct subsystems. We benchmark this enhanced framework against the multi-temperature Boltzmann transport equation, enabling detailed resolution of branch-dependent energy relaxation and identifying bottlenecks in thermalization. This approach provides a more accurate and comprehensive description of heat flow in emerging materials, offering novel insights into phonon dynamics and electron-phonon interactions. These theoretical advances pave the way for the improved design and optimization of next-generation nanoelectronic and photothermal devices.

cond-mat.mes-hall

Low-Altitude UAV-Carried Movable Antenna for Joint Wireless Power Transfer and Covert Communications

The proliferation of Internet of Things (IoT) networks has created an urgent need for sustainable energy solutions, particularly for the battery-constrained spatially distributed IoT nodes. While low-altitude uncrewed aerial vehicles (UAVs) employed with wireless power transfer (WPT) capabilities offer a promising solution, the line-of-sight channels that facilitate efficient energy delivery also expose sensitive operational data to adversaries. This paper proposes a novel low-altitude UAV-carried movable antenna-enhanced transmission system joint WPT and covert communications, which simultaneously performs energy supplements to IoT nodes and establishes transmission links with a covert user by leveraging wireless energy signals as a natural cover. Then, we formulate a multi-objective optimization problem that jointly maximizes the total harvested energy of IoT nodes and sum achievable rate of the covert user, while minimizing the propulsion energy consumption of the low-altitude UAV. To address the non-convex and temporally coupled optimization problem, we propose a mixture-of-experts-augmented soft actor-critic (MoE-SAC) algorithm that employs a sparse Top-K gated mixture-of-shallow-experts architecture to represent multimodal policy distributions arising from the conflicting optimization objectives. We also incorporate an action projection module that explicitly enforces per-time-slot power budget constraints and antenna position constraints. Simulation results demonstrate that the proposed approach significantly outperforms some baseline approaches and other state-of-the-art deep reinforcement learning algorithms.

cs.NI

Wireless Laser Power Transfer for Low-altitude Uncrewed Aerial Vehicle-assisted Internet of Things: Paradigms, Challenges, and Solutions

Low-altitude uncrewed aerial vehicles (UAVs) have become integral enablers for the Internet of Things (IoT) by offering enhanced coverage, improved connectivity and access to remote areas. A critical challenge limiting their operational capacity lies in the energy constraints of both aerial platforms and ground-based sensors. This paper explores WLPT as a transformative solution for sustainable energy provisioning in UAV-assisted IoT networks. We first systematically investigate the fundamental principles of WLPT and analysis the comparative advantages. Then, we introduce three operational paradigms for system integration, identify key challenges, and discuss corresponding potential solutions. In case study, we propose a multi-agent reinforcement learning framework to address the coordination and optimization challenges in WLPT-enabled UAV-assisted IoT data collection. Simulation results demonstrate that our framework significantly improves energy sustainability and data freshness. Finally, we discuss some future directions.

cs.NI

Safety Alignment Should Be Made More Than Just A Few Attention Heads

Current safety alignment for large language models(LLMs) continues to present vulnerabilities, given that adversarial prompting can effectively bypass their safety measures.Our investigation shows that these safety mechanisms predominantly depend on a limited subset of attention heads: removing or ablating these heads can severely compromise model safety. To identify and evaluate these safety-critical components, we introduce RDSHA, a targeted ablation method that leverages the model's refusal direction to pinpoint attention heads mostly responsible for safety behaviors. Further analysis shows that existing jailbreak attacks exploit this concentration by selectively bypassing or manipulating these critical attention heads. To address this issue, we propose AHD, a novel training strategy designed to promote the distributed encoding of safety-related behaviors across numerous attention heads. Experimental results demonstrate that AHD successfully distributes safety-related capabilities across more attention heads. Moreover, evaluations under several mainstream jailbreak attacks show that models trained with AHD exhibit considerably stronger safety robustness, while maintaining overall functional utility.

cs.CR

Large AI Model-Enabled Secure Communications in Low-Altitude Wireless Networks: Concepts, Perspectives and Case Study

Low-altitude wireless networks (LAWNs) have the potential to revolutionize communications by supporting a range of applications, including urban parcel delivery, aerial inspections and air taxis. However, compared with traditional wireless networks, LAWNs face unique security challenges due to low-altitude operations, frequent mobility and reliance on unlicensed spectrum, making it more vulnerable to some malicious attacks. In this paper, we investigate some large artificial intelligence model (LAM)-enabled solutions for secure communications in LAWNs. Specifically, we first explore the amplified security risks and important limitations of traditional AI methods in LAWNs. Then, we introduce the basic concepts of LAMs and delve into the role of LAMs in addressing these challenges. To demonstrate the practical benefits of LAMs for secure communications in LAWNs, we propose a novel LAM-based optimization framework that leverages large language models (LLMs) to generate enhanced state features on top of handcrafted representations, and to design intrinsic rewards accordingly, thereby improving reinforcement learning performance for secure communication tasks. Through a typical case study, simulation results validate the effectiveness of the proposed framework. Finally, we outline future directions for integrating LAMs into secure LAWN applications.

cs.NI

Energy Efficient Trajectory Control and Resource Allocation in Multi-UAV-assisted MEC via Deep Reinforcement Learning

Mobile edge computing (MEC) is a promising technique to improve the computational capacity of smart devices (SDs) in Internet of Things (IoT). However, the performance of MEC is restricted due to its fixed location and limited service scope. Hence, we investigate an unmanned aerial vehicle (UAV)-assisted MEC system, where multiple UAVs are dispatched and each UAV can simultaneously provide computing service for multiple SDs. To improve the performance of system, we formulated a UAV-based trajectory control and resource allocation multi-objective optimization problem (TCRAMOP) to simultaneously maximize the offloading number of UAVs and minimize total offloading delay and total energy consumption of UAVs by optimizing the flight paths of UAVs as well as the computing resource allocated to served SDs. Then, consider that the solution of TCRAMOP requires continuous decision-making and the system is dynamic, we propose an enhanced deep reinforcement learning (DRL) algorithm, namely, distributed proximal policy optimization with imitation learning (DPPOIL). This algorithm incorporates the generative adversarial imitation learning technique to improve the policy performance. Simulation results demonstrate the effectiveness of our proposed DPPOIL and prove that the learned strategy of DPPOIL is better compared with other baseline methods.

cs.NI

Generative AI-enhanced Low-Altitude UAV-Mounted Stacked Intelligent Metasurfaces

Wireless communication systems face challenges in meeting the demand for higher data rates and reliable connectivity in complex environments. Stacked intelligent metasurfaces (SIMs) have emerged as a promising technology for advanced wave-domain signal processing, where mobile SIMs can outperform fixed counterparts. In this paper, we propose a novel unmanned aerial vehicle (UAV)-mounted SIM (UAV-SIM) assisted communication system within low-altitude economy (LAE) networks, where UAVs act as both cache-enabled base stations and mobile SIM carriers to enhance uplink transmissions. To maximize network capacity, we formulate a UAV-SIM-based joint optimization problem (USBJOP) that integrates user association, UAV-SIM three-dimensional positioning, and multi-layer SIM phase shift design. Due to the non-convexity and NP-hardness of USBJOP, we decompose it into three subproblems, which are the association between UAV-SIMs and users optimization problem (AUUOP), the UAV location optimization problem (ULOP), and the UAV-SIM phase shifts optimization problem (USPSOP). Then, we solve them through an alternating optimization strategy. Specifically, AUUOP and ULOP are transformed into convex forms solvable via the CVX tool, while USPSOP is addressed by a generative artificial intelligence (GAI)-based hybrid optimization algorithm. Simulation results show that the proposed approach achieves approximately 1.5 times higher network capacity compared with suboptimal schemes, effectively mitigates multi-user interference with increasing SIM layers and meta-atoms, and reduces runtime by 10\% while maintaining solution quality, thereby demonstrating its practicality for real-world deployments.

cs.NI

UGKWP and IUGKP methods for Multi-Scale Phonon Transport with Dispersion and Polarization

This paper presents two novel methods for solving multi-scale phonon transport problems with dispersion and polarization effects: the unified gas-kinetic wave-particle (UGKWP) method and the implicit unified gas-kinetic particle (IUGKP) method. Both approaches are based on solving multiple groups of BGK equations at discrete frequency points. The UGKWP method constructs multiscale macroscopic fluxes at cell interfaces through the integral solution of the unsteady BGK equation and efficiently captures non-equilibrium transport using statistical particles. Its wave-particle adaptive framework ensures computational efficiency across different regimes: in the diffusive limit, it matches the cost of explicit diffusion equation solutions, while in the ballistic limit, it performs comparably to pure particle methods. The IUGKP method, specifically designed for steady-state problems, determines the particle evolution scale based on the physical mean free path. This approach enables rapid convergence at both large and small Knudsen numbers, with the latter facilitated by a newly constructed macroscopic prediction equation. Both methods incorporate an adaptive frequency-space sampling technique that maintains particle counts per cell comparable to single-frequency methods, significantly improving computational efficiency and memory usage. The accuracy and efficiency of both methods are validated through various numerical tests, including large-scale three-dimensional conduction heat transfer simulations. Results demonstrate their effectiveness in handling complex phonon transport phenomena across multiple scales.

physics.comp-ph

Implicit unified gas kinetic particle method for steady-state solution of multiscale phonon transport

This paper presents a highly efficient implicit unified gas-kinetic particle (IUGKP) method for obtaining steady-state solutions of multi-scale phonon transport. The method adapts and reinterprets the integral solution of the BGK equation for time-independent solutions. The distribution function at a given point is determined solely by the surrounding equilibrium states, where the corresponding macroscopic quantities are computed through a weighted sum of equilibrium distribution functions from neighboring spatial positions. From a particle perspective, changes in macroscopic quantities within a cell result from particle transport across cell interfaces. These particles are sampled according to the equilibrium state of their original cells, accounting for their mean free path as the traveling distance. The IUGKP method evolves the solution according to the physical relaxation time scale, achieving high efficiency in large Knudsen number regimes. To accelerate convergence for small Knudsen numbers, an inexact Newton iteration method is implemented, incorporating macroscopic equations for convergence acceleration in the near-diffusive limit. The method also addresses spatial-temporal inconsistency caused by relaxation time variations in physical space through the null-collision concept. Numerical tests demonstrate the method's excellent performance in accelerating multi-scale phonon transport solutions, achieving speedups of one to two orders of magnitude. The IUGKP method proves to be an efficient and accurate computational tool for simulating multiscale non-equilibrium heat transfer, offering significant advantages over traditional methods in both numerical performance and physical applicability.

physics.comp-ph