SearcharxivSearch

arXiv subjects

Zhendong Li

Publications and source records attributed to Zhendong Li.

At least 19 recordsLinked to original sources

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.

cs.CV

Movable Antennas Enabled Wireless Powered Networks: Principles and Technologies

As an emerging framework, movable antenna (MA)-enabled wireless powered networks (WPNs) have attracted growing attention. WPNs integrate wireless communication and energy transfer. MA can dynamically adjust the position of antenna units by introducing additional spatial degrees of freedom, so as to make full use of channel gain, optimize the effect of energy beamforming, and further improve the performance of WPNs. In this article, we first classify the implementations of MA, and review the fundamental principles of WPNs. We then highlight the key advantages of MA-enabled WPNs in enhancing wireless power transfer efficiency, realizing flexible and adaptive beamforming, and improving system robustness and interference resilience. Furthermore, four representative application scenarios and three key enabling technologies are discussed. A case study is also presented to show the improvement of energy harvesting performance brought by MA for WPNs. Finally, we discuss the challenges and future directions of MA-enabled WPNs, aiming to provide reference for future research and practice.

cs.NI

AFDM-Enabled ISAC in Dynamic Environments: Fundamentals, Technologies and Opportunities

Dynamic environments pose fundamental challenges to integrated sensing and communication (ISAC), particularly due to severe Doppler effects, rapidly time-varying channels, and the intricate coupling between delay and Doppler shifts. Affine frequency-division multiplexing (AFDM), with its inherent capability of characterizing and separating delay and Doppler effects, has emerged as a promising waveform for dynamic ISAC. This article provides a comprehensive overview on AFDM-enabled ISAC in dynamic environments, covering its fundamental principles, distinctive advantages, representative application scenarios, and key enabling technologies. We first characterize the key features of ISAC in dynamic environments and introduce the fundamentals of AFDM, followed by an analysis of scenarios where AFDM can provide significant performance benefits. Then, several key enabling technologies for AFDM-based ISAC in dynamic environments are elaborated upon, accompanied by case studies on the critical aspects therein. Finally, open challenges and promising future research directions are discussed, aiming to provide a comprehensive reference for researchers and practitioners while inspiring further innovation in this emerging field.

eess.SP

Adaptive Beam Hopping and Power Control for Dual-Layer Over-the-Air Online Federated Learning in LEO Satellite Networks

This paper investigates over-the-air (OTA) computation enabled online federated learning (FL) in low-Earth orbit (LEO) satellite networks. Specifically, we consider a dual-layer OTA aggregation architecture, where ground devices upload analog model updates to serving satellites via uplink OTA aggregation, and satellites forward the aggregated signals to a data processing center through the second round OTA aggregation. Then, we formulate a long-term data-utilization maximization problem in which devices continuously collect new data and untrained samples gradually lose freshness. The problem is subject to the satellite beam budget, transmit-power limit, and global mean squared error (MSE) constraint that governs end-to-end aggregation distortion. This yields a coupled mixed-integer nonlinear programming (MINLP) problem, involving tightly coupled discrete beam-hopping decisions and continuous power control. Due to the combinatorial action space and nonconvex constraints, the problem is NP-hard and computationally intractable. Furthermore, the time-varying satellite topology and dynamic data generation render it a sequential decision-making problem, necessitating adaptive online scheduling. To address these issues, we cast the problem as a Markov decision process and develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that jointly optimizes adaptive beam hopping and power control, using an MSE-aware reward to balance data utilization and aggregation accuracy. Numerical simulation results verify that the proposed algorithm consistently outperforms other benchmark schemes, achieving superior long-term data utilization and faster FL convergence while satisfying the MSE requirement.

cs.IT

FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production

Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scripting, storyboarding, generation, and editing. It further demands persistent asset management and dynamic task orchestration as intermediate outputs, dependencies, and execution states evolve over time. Existing automated systems typically rely on rigid pipelines that are difficult to adapt to diverse inputs and changing workflows, while general-purpose large language models (LLMs) remain unreliable for long-horizon orchestration and multimodal asset routing. We introduce FRAMEWORKERS, a task-centric and workspace-grounded multi-agent framework for open-ended video production. A central Director formulates video creation as dynamic task management, continuously editing a Task Stack to determine which subtask to execute next and which sub-agent to invoke. An Assistant serves as the execution layer, grounding each selected task in a shared Workspace, retrieving the required assets and context, invoking the assigned sub-agent, and persisting the resulting artifacts. Execution capabilities are exposed through modular sub-agents with registered descriptors, allowing new sub-agents to be integrated without redesigning the orchestration workflow. To improve orchestration reliability, we fine-tune the Director via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for descriptor-conditioned task routing. Experiments show that FRAMEWORKERS outperforms strong LLM planners in routing accuracy, recovers reliably from runtime failures, generalizes to unseen sub-agents without retraining, and achieves higher end-to-end video quality and broader task coverage than fixed pipelines, single-agent systems, and prior multi-agent approaches.

cs.AI

Dual-Layer Over-the-Air Federated Learning in LEO Satellite Networks: Architecture, Key Technologies and Applications

Low Earth orbit (LEO) satellite networks are emerging as a pivotal infrastructure for global edge intelligence. In this context, integrating over-the-air (OTA) computation with adaptive beam hopping (BH) provides an innovative framework that seamlessly merges physical-layer analog aggregation with dynamic resource orchestration. This effectively overcomes the stringent bandwidth and power constraints of space platforms while extending federated learning (FL) to pervasive Internet-of-things (IoT) deployments. In this article, we first outline the fundamental principles of the dual-layer OTA model and introduce the adaptive BH mechanism designed for time-varying topologies. Then, we summarize the distinct advantages of this learning-centric architecture, which include decoupling aggregation latency from device density, optimizing spatio-temporal resource efficiency, and balancing data freshness with channel quality. Several application scenarios are explored to highlight the framework's potential across diverse vertical industries. Furthermore, a specific case is studied to demonstrate the practical efficacy of the proposed scheduling policy. The results reveal substantial performance gains in terms of model convergence speed and data utilization for satellite-based FL systems. Finally, we discuss the implementation challenges and outline future research directions, aiming to provide insights for the evolution of ubiquitous non-terrestrial intelligence.

eess.SP

Antenna Positioning and Beamforming Optimization in MA Enabled Secure ISAC Systems: A Gradient-Based Meta Learning Approach

Integrated sensing and communications (ISAC) significantly improves spectral efficiency but introduces security risks regarding the interception of embedded communication signals. This paper proposes an movable antenna (MA)-enabled secure ISAC system that utilizes the spatial degrees of freedom of MA to mitigate these risks. Then, a problem is formulated to maximize the system secrecy rate by jointly optimizing antenna positioning, transmit beamforming, and artificial noise. However, the principal challenge arises from the non-convexity of the optimization problem and the strong coupling of the optimization variables. Generally, traditional optimization methods for this problem suffer from complex mathematical derivations, while existing deep learning approaches rely heavily on the training data distribution. To address these issues, we introduce a gradient-based meta learning (GML) algorithm, which works without pre-training and demonstrates favorable performance. Specifically, the algorithm establishes a neural network for each optimization variable, where the gradient of the objective function with respect to the variable serves as the input, and the output of the network determines the variable's update step. By handling the constraints and constructing penalty terms, the global loss function is used to guide the optimization process. Extensive numerical simulations confirm that the proposed algorithm achieves satisfactory performance in terms of both communication security and sensing capabilities.

eess.SP

GML-Based Optimization for Movable Antenna Wireless Networks: Challenges and Opportunities

Movable antenna (MA) is proposed as an emerging technology for future wireless networks. By leveraging the additional spatial degrees of freedom, MA can proactively reshape the wireless propagation environment, thereby enhancing network performance.However, fully unlocking the potential of MA networks necessitates the joint optimization of MA antenna positioning and beamforming. For this non-convex and highly coupled problem, existing solutions exhibit significant limitations. Therefore, this paper proposes a gradient-based meta learning (GML) optimization framework. Specifically, we first elaborate on the hardware architecture and channel characteristics of MA, based on which we analyze the primary challenges in optimizing MA wireless networks. Subsequently, we introduce the fundamental logic of the GML framework and compare it with existing methods. Furthermore, we discuss the constraint handling strategies for applying the proposed optimization framework to MA networks. A specific case is studied to show the performance of proposed framework based on numerical simulation. Finally, this paper outlines future research directions for both the GML framework and MA wireless networks.

eess.SP

Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak

Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. We propose SsPCA-MIDAS, which integrates supervised scaled PCA (SsPCA) into the mixed-data sampling framework. We establish consistency and asymptotic normality under weak factors, permitting inference on the prediction target. Simulations show that SsPCA-MIDAS outperforms competing PCA-based and supervised methods, especially when weak factors are prevalent. Applying machine-learning techniques such as boosting to the cleaner factors it extracts yields further gains. An extensive application to U.S. macro-financial forecasting shows that SsPCA-MIDAS selects economically meaningful predictors and improves forecasts of GDP, inflation, unemployment, asset prices, and volatility.

econ.EM

AP Association for RHS-Enabled Cell-Free Uplink MIMO in Industrial Indoor UAV Networks

Indoor industrial UAV uplink networks face serious blockage and shadowing from shelves, metal equipment, and production facilities. UAVs are also often clustered and fly along similar straight inspection routes at fixed heights. These features make traditional small-cell deployment less suitable, especially when high reliability, continuous coverage, and good service for weak UAVs are required. Cell-free networks can improve robustness through distributed access points (APs) and UAV?centric communications. Reconfigurable holographic surface (RHS)-enabled APs provide programmable analog receive beams and generate scalar post-RHS observations, which are jointly processed at the CPU for distributed uplink MIMO detection at relatively low hardware cost. Conventional AP association relies on distance, large-scale fading, or post-combining SINR obtained with user-specific digital combiners. Here, however, all UAVs served by a single-feed RHS AP share one amplitude?constrained receive pattern and one scalar AP output. We therefore derive an SINR-like score from this physical output and, under weak inter-AP disturbance correlation, approximate the CPU-side log-det objective by an additive per-AP surrogate, yielding a low-complexity ranking rule. The results show that the nearest AP is not always the best choice, the AP-UAV height difference may have an optimal value, and larger serving clusters bring diminishing returns. Simulations show that the proposed method improves the minimum UAV data rate, average spectral efficiency, fairness, and energy efficiency compared with benchmark schemes.

cs.NI

Mobile Tracking via Target-Mounted IRS-Assisted ISAC System

This paper proposes a target-mounted intelligent reflecting surface (IRS)-assisted integrated sensing and communication framework for real-time unmanned aerial vehicle (UAV) tracking, addressing challenges such as link blockage and weak radar cross section in the low-altitude economy. By integrating the IRS onto the UAV, the system creates a mobile cooperative target that provides controllable line-of-sight echoes for self-tracking while acting as a mobile relay for ground communication enhancement. We establish a comprehensive three dimensions state evolution model for the maneuvering UAV. Based on this model, an extended Kalman filter is immediately implemented to achieve real time tracking of the moving UAV. To characterize the fundamental theoretical limits of this recursive estimation process, we derive the analytical posterior Cramer Rao bound and a closed form expression for the elliptical tradeoff performance bound to quantify the relationship between sensing precision and communication throughput. To ensure millisecond level responsiveness, we develop a low complexity joint beamforming design. By utilizing the analytical mapping between tracking and communication requirements, the proposed scheme yields closed form solutions for beamforming vectors, effectively bypassing the time consuming numerical iterations of conventional methods. Numerical simulations demonstrate that the proposed framework significantly outperforms traditional fixed-deployment benchmarks across complex maneuvering trajectories, achieving centimeter-level accuracy while substantially reducing transmit power and processing latency.

eess.SP

Transmissive RIS Transceiver-Empowered ISAC Systems: Energy Efficiency Optimization for Perfect and Imperfect CSI

In this paper, a novel transmissive reconfigurable intelligent surface (TRIS) transceiver is employed to enable an integrated sensing and communication (ISAC) system supporting both communication and sensing. Under both perfect and imperfect channel state information (CSI), we study the transmit beamforming design for the TRIS transceiver to maximize the system energy efficiency (EE), subject to per-user minimum-rate guarantees, a minimum beampattern gain toward the sensing target, and per-antenna power constraints. The corresponding EE maximization problems are challenging to solve due to the fractional objective and non-convex constraints. In particular, under imperfect CSI, the resulting semi-infinite constraints further complicate the problem. For the perfect CSI case, we first apply the fractional programming (FP) methodology to obtain more tractable reformulations of the rate functions, and then propose an iterative algorithm based on the majorization-minimization (MM) framework. For the imperfect CSI case, we utilize the S-Procedure to transform the semi-infinite inequality constraints into linear matrix inequalities (LMIs), and further develop an efficient MM-based algorithm with the aid of slack variables. Numerical results demonstrate the convergence and effectiveness of the proposed algorithms and validate the EE gains of the TRIS transceiver-enabled ISAC system.

eess.SP

Constrained Tensor Decomposition-Based Target Sensing for Sparse Non-Uniform Array-Enabled AFDM ISAC Systems

Sparse non-uniform array-enabled affine frequency division multiplexing (AFDM) is a promising candidate for integrated sensing and communication (ISAC), while its performance critically depends on accurate target parameter estimation. In this paper, we propose a constrained tensor decomposition-based sensing framework for delay, Doppler, and angle estimation. Specifically, a manifold-constrained alternating least squares (ALS) algorithm is developed by exploiting the sparse array geometry structure, enabling robust factor matrix extraction and direct angle estimation. From the decomposed factor matrices, we further apply an iterative one dimensional golden section search to refine delay and Doppler shift. Simulation results demonstrate that the proposed algorithm nearly attains Cram\'er-Rao bound (CRB) and significantly outperforms unconstrained ALS and conventional methods, validating its effectiveness for sparse non-uniform array-enabled AFDM ISAC systems.

eess.SP

Tensor-Based Dynamic Channel Estimation for mmWave Movable Antenna MIMO Systems

This paper investigates the dynamic channel estimation algorithm in mmWave movable antenna (MA) multiple-input multiple-output (MIMO) systems. To achieve highly accurate channel estimation, we propose a tensor decomposition-based channel estimation algorithm. First, by leveraging the path response model and utilizing the intrinsic sparsity of mmWave channels, the channel corresponding to MA pairs at the base station and mobile station is transformed into a superposition of channels from sparse paths. Next, the received signal is constructed as a fourth-order tensor to fully capture the high-dimensional structural information of the MA MIMO channel. Then, two tensor decomposition schemes are adopted to extract the factor matrices, and our analysis reveals that the uniqueness of the decomposition can be guaranteed in our model. Subsequently, the propagation loss, frequency offset, angle of arrival/departure, and time delay are obtained based on these factor matrices and the channel matrix can be rebuilt. Additionally, Cram\'er-Rao bound (CRB) is also derived as a performance evaluation standard, proving that the proposed algorithm achieves a higher estimation accuracy and nearly approaches this minimum bound. Moreover, normalized mean square error (NMSE) is selected as the evaluation metrics for estimation accuracy. Finally, simulation results reveal a notable reduction in the estimation error of the proposed algorithm when compared to the baseline algorithms, confirming its estimation advantage.

eess.SP

Knowledge-Centric Agents for Workflow Generation in ComfyUI

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics. To this end, we propose a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels. We first perform knowledge inversion to distill hierarchical representations, ranging from full pseudo-codes and skeletons to high-level strategies, from large collections of real-world workflows. We then conduct knowledge injection through supervised fine-tuning, teaching the model to reason from task descriptions to strategies and from strategies to executable structures. During inference, the model performs reversible reasoning to synthesize executable workflows, augmented by self-refinement for structural coherence. Extensive experiments demonstrate that our method produces workflows with richer node diversity, more coherent structures, and higher execution success rates than existing systems, establishing a new foundation for knowledge-driven, agentic workflow generation.

cs.AI

Collision and coalescence dynamics of bosonic quantum Hall droplets

Recently bosonic quantum Hall droplets have been observed in rapidly rotating two-dimensional Bose-Einstein condensates (BECs), which exhibit robust dynamical stability. Inspired by this, we systematically investigate the collision and coalescence dynamics of these droplets within the Gross-Pitaevskii framework. For two-droplet collisions, we find two distinct collision outcomes, namely merging and separation, that are controlled by the initial relative velocity. The critical velocity exhibits a universal scaling law with the interaction and the particle number as $v_c \propto (gN)^{1/4}$, which can be interpreted from a simplified analytical model, revealing the essential role of the collision time. It differs fundamentally from the mechanism governing the conventional Lee-Huang-Yang stabilized quantum droplets. Furthermore, while the collision can change the shape of the droplet significantly, the center of mass trajectory remains nearly unaffected, owing to the conservation of angular momentum. For overlapping stationary droplets, vortex arrays can emerge through Kelvin-Helmholtz instability driven by phase-induced shear flow. Although two droplets may merge into a larger one, extended states cannot be constructed from multiple overlapping droplets. Instead, the system dynamically reorganizes into new isolated droplets, revealing the localized property in the bulk region. Our results reveal the unique nonequilibrium dynamics of quantum Hall droplets and suggest new pathways for manipulating strongly correlated rotating quantum fluids.

cond-mat.quant-gas

Clifford disentanglers for entanglement reduction in molecular electronic structure simulations

Entanglement is a key bottleneck limiting the efficiency of tensor-network and quantum simulations of molecular electronic structures. Here, we systematically assess and extend Clifford disentanglers as a structure-preserving approach to entanglement reduction: they can modify the entanglement structure of qubit wavefunctions while retaining the Pauli-string form of qubit Hamiltonians. To enable a practical search over Clifford transformations, we classify Clifford operators by their action on the Schmidt spectrum across a bipartition, reducing the two- and four-qubit search spaces to 20 and 91392 representatives, respectively. Embedded in an iterative Clifford-augmented matrix product state framework, these transformations reduce the energy errors at fixed bond dimension for the molecular test cases studied and mitigate the dependence on orbital orderings and fermion-to-qubit mappings. We further show that Clifford disentanglers can also benefit quantum simulations such as the shallow-circuit variational quantum eigensolver calculations. Together, these results establish Clifford disentanglers as a useful structure-preserving entanglement-engineering tool for tensor-network and quantum simulations of molecular electronic structure, while also clarifying their correlation dependence and motivating future developments.

quant-ph

Agents' Last Exam

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.

cs.AI