SearcharxivSearch

arXiv subjects

Ziqian Zhang

Publications and source records attributed to Ziqian Zhang.

17 recordsLinked to original sources

Sharp Regularity Thresholds for Viscous Katz--Pavlović Dyadic Models

We study the unforced viscous Katz--Pavlović dyadic model with non-negative smooth initial data for two shell settings. For super-exponential scales $N_n=N_0^{b^n}$, with $N_0>1$ and $1 (\log2)/(2\logΛ)$. In particular, for $Λ>2^{3/2}$, the result reaches $α=1/3$ and, together with the known classical blow-up theorem, identifies the sharp threshold in classical shell ratios.

math.AP

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.

cs.CV

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail to differentiate question difficulty, limiting their ability to effectively distinguish models' capabilities. To address this limitation, we propose RankLLM, a novel framework designed to quantify both question difficulty and model competency. RankLLM introduces difficulty as the primary criterion for differentiation, enabling a more fine-grained evaluation of LLM capabilities. RankLLM's core mechanism facilitates bidirectional score propagation between models and questions. The core intuition of RankLLM is that a model earns a competency score when it correctly answers a question, while a question's difficulty score increases when it challenges a model. Using this framework, we evaluate 30 models on 35,550 questions across multiple domains. RankLLM achieves 90% agreement with human judgments and consistently outperforms strong baselines such as IRT. It also exhibits strong stability, fast convergence, and high computational efficiency, making it a practical solution for large-scale, difficulty-aware LLM evaluation.

cs.CL

Brillouin-Enhanced Photonic Stepped-Frequency Radar

Photonic stepped-frequency (SF) radar offers high range resolution and only requires low-speed driving electronics, but existing architectures face challenges in achieving low phase noise and uniform frequency steps simultaneously. Here, we demonstrate a photonic SF radar system that exploits dual Brillouin lasers in a shared fiber cavity to simultaneously suppress phase noise and ensure uniform frequency stepping. Phase noise is reduced through Brillouin optomechanical suppression and common-mode noise rejection upon photomixing. Frequency-step uniformity is enforced via lasing at a series of uniformly spaced cavity resonances. The system generates an X-band SF waveform spanning 1.31 GHz, achieving >23 dB of phase-noise improvement at a 100 kHz offset relative to a low-cost driving voltage-controlled oscillator. The demonstrated system reduces the dependence of the output waveform quality on noise in the driving electronics, offering a path towards high-performance radar sensing.

physics.optics

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware heterogeneity across bimanual robotic platforms. To bridge this gap, we introduce RoboCOIN, a large-scale multi-embodiment bimanual manipulation dataset comprising over 180,000 demonstrations collected from 15 distinct robotic platforms. Spanning 16 diverse environments-including residential, commercial, and industrial settings-the dataset features 421 bimanual tasks systematically categorized by 39 bimanual collaboration actions and 432 objects. A key innovation of our work is the hierarchical capability pyramid, which provides granular annotations ranging from trajectory-level concepts to segment-level subtasks and frame-level kinematics. Furthermore, we present CoRobot, an efficient data processing pipeline powered by the Robot Trajectory Markup Language (RTML), designed to facilitate quality assessment, automated annotation, and unified multi-embodiment and data management. Extensive experiments demonstrate the effectiveness of RoboCOIN in enhancing the performance of various bimanual manipulation models across a wide spectrum of robotic embodiments. The entire dataset and codebase are fully open-sourced, providing a valuable resource for advancing research in bimanual and multi-embodiment manipulation.

cs.RO

Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics

Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech production, which also encompasses speaker traits, including emotional states. To bridge this gap, we conducted a series of experiments to investigate the potential of the phonation excitation information and articulatory kinematics for SER. Due to the scarcity of training data for this purpose, we introduce a portrayed emotional dataset, STEM-E2VA, which includes audio and physiological data such as electroglottography (EGG) and electromagnetic articulography (EMA). EGG and EMA provide information of phonation excitation and articulatory kinematics, respectively. Additionally, we performed emotion recognition using estimated physiological data derived through inversion methods from speech, instead of collected EGG and EMA, to explore the feasibility of applying such physiological information in real-world SER. Experimental results confirm the effectiveness of incorporating physiological information about speech production for SER and demonstrate its potential for practical use in real-world scenarios.

cs.SD

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrarily Shaped Modular Aerial Robot Systems

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior work on MARS self-reconfiguration has solely focused on maximizing controllability margins to tolerate a single rotor or unit fault for rectangular-shaped MARS. We propose TransforMARS, a general fault-tolerant reconfiguration framework that transforms arbitrarily shaped MARS under multiple rotor and unit faults while ensuring continuous in-air stability. Specifically, we develop algorithms to first identify and construct minimum controllable assemblies containing faulty units. We then plan feasible disassembly-assembly sequences to transport MARS units or subassemblies to form target configuration. Our approach enables more flexible and practical feasible reconfiguration. We validate TransforMARS in challenging arbitrarily shaped MARS configurations, demonstrating substantial improvements over prior works in both the capacity of handling diverse configurations and the number of faults tolerated. The videos and source code of this work are available at the anonymous repository: https://anonymous.4open.science/r/TransforMARS-1030/

cs.RO

Benchmark Study of Transient Stability during Power-Hardware-in-the-Loop and Fault-Ride-Through capabilities of PV inverters

The deployment of PV inverters is rapidly expanding across Europe, where these devices must increasingly comply with stringent grid requirements.This study presents a benchmark analysis of four PV inverter manufacturers, focusing on their Fault Ride Through capabilities under varying grid strengths, voltage dips, and fault durations, parameters critical for grid operators during fault conditions.The findings highlight the influence of different inverter controls on key metrics such as total harmonic distortion of current and voltage signals, as well as system stability following grid faults.Additionally, the study evaluates transient stability using two distinct testing approaches.The first approach employs the current standard method, which is testing with an ideal voltage source. The second utilizes a Power Hardware in the Loop methodology with a benchmark CIGRE grid model.The results reveal that while testing with an ideal voltage source is cost-effective and convenient in the short term, it lacks the ability to capture the dynamic interactions and feedback loops of physical grid components.This limitation can obscure critical real world factors, potentially leading to unexpected inverter behavior and operational challenges in grids with high PV penetration.This study underscores the importance of re-evaluating conventional testing methods and incorporating Power Hardware in the Loop structures to achieve test results that more closely align with real-world conditions.

eess.SY

Charge-Discharge Coupling Strategy for Dispatching Problems with Electric Tractors at Airports

Airports worldwide are actively promoting the transition of ground service vehicles from traditional fuel-powered vehicles to electric vehicles. The key to the successful implementation of this transition lies in the development of efficient electric vehicle dispatching models that comprehensively consider the charge-discharge processes of electric vehicles. However, due to the nonlinear characteristics of charge-discharge processes, finding precise solutions poses a significant challenge. Previous researchers have often used traditional energy consumption models and constant charging rates to simplify calculations, but this has resulted in inaccurate estimates of the remaining battery charge level. Furthermore, the lack of diverse pacing and charging strategies for airport ground service vehicles necessitates more adaptable solutions to enhance operational efficiency. To address these challenges, this paper uses airport electric tractors as a case study, develops an accurate model that takes into account the start-stop process and a piecewise linear charging function, designs an improved genetic algorithm that incorporates a greedy algorithm and an adaptive strategy, and develops charge-discharge coupling strategies for different configuration scenarios at Nanjing Lukou Airport to meet current and future needs. The research results indicate that compared to traditional genetic algorithms, the proposed improved genetic algorithm significantly enhances solution accuracy and convergence speed. Additionally, with the increase in flight scale, airports can appropriately enhance their charging strategies; airports with dispersed aircraft stands should devise higher pacing strategies compared to those with dense aircraft stands.

math.OC

GDN: A Stacking Network Used for Skin Cancer Diagnosis

Skin cancer, the primary type of cancer that can be identified by visual recognition, requires an automatic identification system that can accurately classify different types of lesions. This paper presents GoogLe-Dense Network (GDN), which is an image-classification model to identify two types of skin cancer, Basal Cell Carcinoma, and Melanoma. GDN uses stacking of different networks to enhance the model performance. Specifically, GDN consists of two sequential levels in its structure. The first level performs basic classification tasks accomplished by GoogLeNet and DenseNet, which are trained in parallel to enhance efficiency. To avoid low accuracy and long training time, the second level takes the output of the GoogLeNet and DenseNet as the input for a logistic regression model. We compare our method with four baseline networks including ResNet, VGGNet, DenseNet, and GoogLeNet on the dataset, in which GoogLeNet and DenseNet significantly outperform ResNet and VGGNet. In the second level, different stacking methods such as perceptron, logistic regression, SVM, decision trees and K-neighbor are studied in which Logistic Regression shows the best prediction result among all. The results prove that GDN, compared to a single network structure, has higher accuracy in optimizing skin cancer detection.

cs.CV

A Survey of Progress on Cooperative Multi-agent Reinforcement Learning in Open Environment

Multi-agent Reinforcement Learning (MARL) has gained wide attention in recent years and has made progress in various fields. Specifically, cooperative MARL focuses on training a team of agents to cooperatively achieve tasks that are difficult for a single agent to handle. It has shown great potential in applications such as path planning, autonomous driving, active voltage control, and dynamic algorithm configuration. One of the research focuses in the field of cooperative MARL is how to improve the coordination efficiency of the system, while research work has mainly been conducted in simple, static, and closed environment settings. To promote the application of artificial intelligence in real-world, some research has begun to explore multi-agent coordination in open environments. These works have made progress in exploring and researching the environments where important factors might change. However, the mainstream work still lacks a comprehensive review of the research direction. In this paper, starting from the concept of reinforcement learning, we subsequently introduce multi-agent systems (MAS), cooperative MARL, typical methods, and test environments. Then, we summarize the research work of cooperative MARL from closed to open environments, extract multiple research directions, and introduce typical works. Finally, we summarize the strengths and weaknesses of the current research, and look forward to the future development direction and research problems in cooperative MARL in open environments.

cs.MA

Learning to Coordinate with Anyone

In open multi-agent environments, the agents may encounter unexpected teammates. Classical multi-agent learning approaches train agents that can only coordinate with seen teammates. Recent studies attempted to generate diverse teammates to enhance the generalizable coordination ability, but were restricted by pre-defined teammates. In this work, our aim is to train agents with strong coordination ability by generating teammates that fully cover the teammate policy space, so that agents can coordinate with any teammates. Since the teammate policy space is too huge to be enumerated, we find only dissimilar teammates that are incompatible with controllable agents, which highly reduces the number of teammates that need to be trained with. However, it is hard to determine the number of such incompatible teammates beforehand. We therefore introduce a continual multi-agent learning process, in which the agent learns to coordinate with different teammates until no more incompatible teammates can be found. The above idea is implemented in the proposed Macop (Multi-agent compatible policy learning) algorithm. We conduct experiments in 8 scenarios from 4 environments that have distinct coordination patterns. Experiments show that Macop generates training teammates with much lower compatibility than previous methods. As a result, in all scenarios Macop achieves the best overall coordination ability while never significantly worse than the baselines, showing strong generalization ability.

cs.MA

Implementing a new fully stepwise decomposition-based sampling technique for the hybrid water level forecasting model in real-world application

Various time variant non-stationary signals need to be pre-processed properly in hydrological time series forecasting in real world, for example, predictions of water level. Decomposition method is a good candidate and widely used in such a pre-processing problem. However, decomposition methods with an inappropriate sampling technique may introduce future data which is not available in practical applications, and result in incorrect decomposition-based forecasting models. In this work, a novel Fully Stepwise Decomposition-Based (FSDB) sampling technique is well designed for the decomposition-based forecasting model, strictly avoiding introducing future information. This sampling technique with decomposition methods, such as Variational Mode Decomposition (VMD) and Singular spectrum analysis (SSA), is applied to predict water level time series in three different stations of Guoyang and Chaohu basins in China. Results of VMD-based hybrid model using FSDB sampling technique show that Nash-Sutcliffe Efficiency (NSE) coefficient is increased by 6.4%, 28.8% and 7.0% in three stations respectively, compared with those obtained from the currently most advanced sampling technique. In the meantime, for series of SSA-based experiments, NSE is increased by 3.2%, 3.1% and 1.1% respectively. We conclude that the newly developed FSDB sampling technique can be used to enhance the performance of decomposition-based hybrid model in water level time series forecasting in real world.

cs.LG

Fast Teammate Adaptation in the Presence of Sudden Policy Change

In cooperative multi-agent reinforcement learning (MARL), where an agent coordinates with teammate(s) for a shared goal, it may sustain non-stationary caused by the policy change of teammates. Prior works mainly concentrate on the policy change during the training phase or teammates altering cross episodes, ignoring the fact that teammates may suffer from policy change suddenly within an episode, which might lead to miscoordination and poor performance as a result. We formulate the problem as an open Dec-POMDP, where we control some agents to coordinate with uncontrolled teammates, whose policies could be changed within one episode. Then we develop a new framework, fast teammates adaptation (Fastap), to address the problem. Concretely, we first train versatile teammates' policies and assign them to different clusters via the Chinese Restaurant Process (CRP). Then, we train the controlled agent(s) to coordinate with the sampled uncontrolled teammates by capturing their identifications as context for fast adaptation. Finally, each agent applies its local information to anticipate the teammates' context for decision-making accordingly. This process proceeds alternately, leading to a robust policy that can adapt to any teammates during the decentralized execution phase. We show in multiple multi-agent benchmarks that Fastap can achieve superior performance than multiple baselines in stationary and non-stationary scenarios.

cs.MA

Multi-agent Continual Coordination via Progressive Task Contextualization

Cooperative Multi-agent Reinforcement Learning (MARL) has attracted significant attention and played the potential for many real-world applications. Previous arts mainly focus on facilitating the coordination ability from different aspects (e.g., non-stationarity, credit assignment) in single-task or multi-task scenarios, ignoring the stream of tasks that appear in a continual manner. This ignorance makes the continual coordination an unexplored territory, neither in problem formulation nor efficient algorithms designed. Towards tackling the mentioned issue, this paper proposes an approach Multi-Agent Continual Coordination via Progressive Task Contextualization, dubbed MACPro. The key point lies in obtaining a factorized policy, using shared feature extraction layers but separated independent task heads, each specializing in a specific class of tasks. The task heads can be progressively expanded based on the learned task contextualization. Moreover, to cater to the popular CTDE paradigm in MARL, each agent learns to predict and adopt the most relevant policy head based on local information in a decentralized manner. We show in multiple multi-agent benchmarks that existing continual learning methods fail, while MACPro is able to achieve close-to-optimal performance. More results also disclose the effectiveness of MACPro from multiple aspects like high generalization ability.

cs.MA

Photonic Radar for Contactless Vital Sign Detection

Vital sign detection is used across ubiquitous scenarios in medical and health settings. Contact and wearable sensors have been widely deployed. However, they are unsuitable for patients with burn wounds or infants with insufficient attaching areas. Contactless detection can be achieved using camera imaging, but it is susceptible to ambient light conditions and creates privacy concerns. Here, we report the first demonstration of a photonic radar for non-contact vital signal detection to overcome these challenges. This photonic radar can achieve millimeter range resolution based on synthesized radar signals with a bandwidth of up to 30 GHz. The high resolution of the radar system enables accurate respiratory detection from breathing simulators and a cane toad as a human proxy. Moreover, we demonstrated that the optical signals generated from the proposed system can enable vital sign detection based on light detection and ranging (LiDAR). This demonstration reveals the potential of a sensor-fusion architecture that can combine the complementary features of radar and LiDAR for improved sensing accuracy and system resilience. The work provides a novel technical basis for contactless, high-resolution, and high-privacy vital sign detection to meet the increasing demands in future medical and healthcare applications.

physics.app-ph

Photonic Generation of Radar Signals with 30 GHz Bandwidth and Ultra-High Time-Frequency Linearity

Photonic generation of radio-frequency signals has shown significant advantages over the electronic counterparts, allowing the high precision generation of radio-frequency carriers up to the terahertz-wave region with flexible bandwidth for radar applications. Great progress has been made in photonics-based radio-frequency waveform generation. However, the approaches that rely on sophisticated benchtop digital microwave components, such as synthesizers and digital-to-analog converters have limited achievable bandwidth and thus resolution for radar detections. Methods based on voltage-controlled analog oscillators exhibit high time-frequency non-linearity, causing degraded sensing precision. Here, we demonstrate, for the first time, a photonic stepped-frequency (SF) waveform generation scheme enabled by MHz electronics with a tunable bandwidth exceeding 30 GHz and intrinsic time-frequency linearity. The ultra-wideband radio-frequency signal generation is enabled by using a polarization-stabilized optical cavity to suppress intra-cavity polarization-dependent instability; meanwhile, the signal's high-linearity is achieved via consecutive MHz acousto-optic frequency-shifting modulation without the necessity of using electro-optic modulators that have bias-drifting issues. We systematically evaluate the system's signal quality and imaging performance in comparison with conventional photonic radar schemes that use high-speed digital electronics, confirming its feasibility and excellent performance for high-resolution radar applications.

physics.app-ph