SearcharxivSearch

arXiv subjects

Yi Peng

Publications and source records attributed to Yi Peng.

At least 19 recordsLinked to original sources

Stability of admissible solutions for coexisting phase transitions for one-dimensional compressible van der Waals fluids

In this paper, we investigate the dynamic stability of certain steady-state solutions to the periodic boundary value problem for compressible isentropic Navier-Stokes system under the van der Waals equation of state in one space dimension. These steady-state solutions correspond to the admissible solutions describing two-phase coexisting phase transitions, where the integral average of the specific volume belongs to the Maxwell region. We first construct a semi-discrete staggered grid difference scheme to prove the local existence of solutions to the periodic problem, without imposing the standard stability hypothesis \(p_v<0\). Then, by virtue of rigorous piecewise a priori estimates, we demonstrate that the periodic boundary value problem for van der Waals fluids possesses a global solution existing for all time, and this solution converges uniformly to the admissible steady state as time tends to infinity. This result firmly establishes the nonlinear stability of the admissible phase-transition solutions under general small initial disturbances.

math.AP

Nonlinear Bias-Compensated Adaptive Filter and Its Application for Time-Series Prediction

Most existing nonlinear adaptive filtering algorithms only account for output noise, neglecting the fact that input noise is also prevalent in practice. Although the recently proposed bias-compensated kernel least mean square (BCKLMS) algorithm addresses input noise in the nonlinear errors-in-variables (EIV) model, it still suffers from two major limitations. First, the use of a fixed-size dictionary restricts network growth but also prevents it from fully capturing the characteristics of the input signal. Second, as an least mean square (LMS) based algorithm, it exhibits poor robustness in the presence of non-Gaussian noise in the output signal. To overcome these issues, this paper proposes the random Fourier bias-compensated filter under general adaptive function (RFFBCGA) algorithm. Within the random Fourier feature based bias-compensated (RFFBC) framework, the proposed algorithm not only maintains a fixed network structure and effectively mitigates input noise interference through the BC term, but also achieves improved characterization of the input signal. Moreover, by leveraging the flexible form of the general adaptive (GA) function, the algorithm's robustness across various noise scenarios is further enhanced. Extensive simulations, including real-world time series prediction tasks, demonstrate the superiority of the proposed method.

cs.LG

Online Censoring-Based Widely Linear Total Least lncosh Method for Improved Power System Frequency Estimation

Recently, under the presumption of a noise-free input, the augmented complex least lncosh (ACLlncosh) method was introduced for a power system frequency estimate and showed robust performance when impulsive noise polluted the output signal. However, in practical terms, noise often contaminates input signals, which drastically reduces the efficiency of the ACLlncosh method. To enhance robustness against noisy input-output while maintaining resilience to impulsive noise in the output signal, this paper proposed an online censoring-based widely linear total least lncosh (OC-WL-TLlnC) method. This method improves performance under both balanced and unbalanced settings by filtering out less valuable data via online censoring, hence reducing the computing burden. Furthermore, a variable parameter approach is incorporated to accelerate convergence and improve steady-state accuracy, thereby ensuring adaptability to dynamic power system conditions. The proposed methods significantly enhance frequency estimate performance by addressing the constraints of current techniques and offering a computationally efficient, noise-resilient solution for real-time power system monitoring.

eess.SP

Decentralized Variational Bayesian UKF with Maximum Generalized Student's t-kernel Correntropy for Wide-Area Power System state estimation

A Conventional centralized state estimators exhibit limited robustness in large-scale grids and face practical deployment hurdles. To overcome these challenges, this paper proposes a decentralized maximum generalized Student's t-kernel correntropy Variational Bayesian unscented Kalman filter (D-MGST-VBUKF). The algorithm optimizes the estimation performance at three levels for the regionalized state estimation needs: first, to address non-Gaussian measurement noise in practical systems, we propose the cost function using MGST, retaining Student's t robustness while improving adaptability to complex noise by expanding the degree-of-freedom parameter; secondly, the VB inference framework is constructed to model the unknown noise distribution online, and the joint optimization of the noise statistical characteristics and state estimation is realized by constructing the conjugate prior distribution; finally, the regional state fusion mechanism is established based on the topological correlation characteristics of the power grid, and the global consistency correction of the local estimation results is realized by constructing the state coordination equation of the boundary nodes. Simulation experiments in IEEE 14-bus and IEEE 39-bus system show that the method has stronger robustness compared with the traditional algorithm under non-Gaussian noise environment and unknown noise environment.

eess.SP

Outlier-Robust unscented Kalman filter based on generalized correntropy induced

Conventional Kalman filtering (KF) approaches exhibit significant limitations in addressing nonlinear state estimation problems contaminated by non-Gaussian noise disturbances. To overcome these challenges, this work proposes a robust iterative square root unscented Kalman Filter based on the generalized correntropy induced (SR-GCI-IUKF). While sharing the maximum correntropy criterion's (MCC) ability to characterize higher-order noise statistics, the proposed GCI framework exhibits intrinsic kernel bandwidth insensitivit a critical advantage enabling robust adaptation to diverse complex noise environments through its generalized kernel structure. For nonlinear state estimation challenges, the algorithm constructs a nonlinear error generalization model that dynamically corrects measurement-induced errors during the state update phase, thereby significantly enhancing estimation accuracy in strongly nonlinear regimes. Furthermore, the square-root decomposition implementation ensures numerical robustness by preserving covariance matrix positive definiteness throughout recursive operations. Theoretical stability guarantees are established through rigorous error dynamics analysis, demonstrating bounded estimation variance under non-Gaussian disturbances. Finally, experiments are carried out in nonlinear systems, land vehicle navigation systems as well as power system FASE to compare other robust algorithms, and it is determined that the proposed algorithm has stronger robustness.

eess.SP

A Fast Robust Adaptive filter using Improved Data-Reuse Method

Adaptive filter in complex scenarios demands algorithms that integrate fast convergence, low complexity, and robust performance under diverse noise conditions. To address this challenge, we propose a online censoring robust total generalized adaptive filter using improved data-reused method (RTGA-IDROC) algorithm. The proposed RTGA variant possesses the advantages of both the total least squares (TLS) strategy and the robust generalized adaptive (RGA) function. This algorithm not only effectively handles input noise under the errors-in-variables (EIV) model but also achieves excellent performance across diverse noise environments. Furthermore, to meet the high demand for convergence speed in practical applications, an improved data reuse (IDR) method is introduced, enabling faster convergence in the early stages of iteration without compromising steady-state performance. The increased computational complexity brought by the IDR method is mitigated using the online censoring (OC) strategy. We also modify the OC threshold for real-valued algorithms, as the original threshold was defined for the complex domain. Beyond these algorithmic enhancements, a local stability analysis for the proposed algorithm is provided, and the theoretical steady-state mean-square deviation (MSD) is derived. Finally, simulation experiments in system identification and acoustic echo cancellation (AEC) scenarios validate the superior performance of the proposed algorithm.

eess.SP

Outstanding TC Enhancement in 5d-3d Y2NiIrO6 by Compression

Understanding and predicting the properties of 5d compounds critically depend on the identification of the superexchange interactions from which their magnetism emerges. The study of pressure effects on double perovskite Y2NiIrO6 (YNIO) provides deep insight toward this goal. At ambient pressure, YNIO is a ferrimagnetic insulator with the Ir4+-5d Jeff = 1/2 Mott-insulating state. Under physical pressure up to 17 GPa, the compound exhibits concurrent compression on Ni/Ir-O bond lengths and Ni-O-Ir bond angles, leading to an increase of the Curie temperature from 192 to 243 K. On the contrary, external pressure increases distanced Ir-Ir interaction and in turn induces magnetic frustration in Sr2IrO4/Sr3Ir2O7 due to the extended 5d orbitals. In YNIO, the rock-salt ordered Ni-Ir naturally blocks extended superexchange beyond the nearest neighbor, and in turn suppresses such magnetic frustration. Moreover, the orthogonal Ni eg-Ir t2g pathway in YNIO is robust under lattice distortion, while the superexchange is weakened by bond bending in La2NiMnO6 with a similar half-filed eg-t2g configuration. Our findings establish a framework for elucidating the mechanism of 5d-3d superexchange and guide bond-engineered magnetism in iridate-related systems.

cond-mat.str-el

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDream-O1-Image, a natively unified generative foundation model via pixel-space Diffusion Transformer, that pioneers a paradigm shift from modular architectures to an end-to-end in-context visual generation engine. By mapping raw image pixels, text tokens, and task-specific conditions into a single shared token space, HiDream-O1-Image achieves a structural unification of multimodal inputs within an Unified Transformer (UiT) architecture. This native encoding paradigm eliminates the need for separate VAEs or disjoint pre-trained text encoders, allowing the model to treat diverse generation and editing tasks as a consistent in-context reasoning process. Extensive experiments show that HiDream-O1-Image excels across various generation tasks, including text-to-image generation, instruction-based editing, and subject-driven personalization. Notably, with only 8B parameters, HiDream-O1-Image (8B) achieves performance parity with or even surpasses established state-of-the-art models with significantly larger parameters (e.g., 27B Qwen-Image). Crucially, to validate the immense scalability of this paradigm, we successfully scale the architecture up to over 200B parameters. Experimental results demonstrate that this massive-scale version HiDream-O1-Image-Pro (200B+) unlocks unprecedented generative capabilities and superior performance, establishing new state-of-the-art benchmarks. Ultimately, HiDream-O1-Image highlights the immense potential of natively unified architectures and charts a highly scalable path toward next-generation multimodal AI.

cs.CV

A Regularized Framework and Admissible Solutions for Liquid-Vapor Phase Transitions in Steady Compressible Flows

We investigate the well-posedness of the periodic boundary value problem for the steady compressible isentropic Navier-Stokes system under the van der Waals equation of state. The main difficulty arises from the non-monotonicity of the pressure, which induces liquid-vapor phase transitions and consequently leads to both physical instabilities and mathematical non-uniqueness of solutions. It is shown that the occurrence of a phase transition is determined by whether the integral average of the specific volume lies inside the gas-liquid coexistence region defined by the Maxwell construction. By introducing an artificial viscosity, we construct an approximate system. When the integral average of the specific volume falls within the Maxwell region, the approximate solution converges, as the artificial viscosity tends to zero, to the equilibrium states given by Maxwell's construction, with the diffuse interface sharpening into a discontinuity. Conversely, if the integral average of the specific volume lies outside this region, the limiting solution remains outside as well, meaning that no phase transition occurs. These results demonstrate that the non-monotonicity of the pressure, combined with the condition that the integral average of the specific volume belongs to the Maxwell region, can act as a nucleation mechanism for phase transitions in the isentropic gas-liquid problem. Furthermore, the proposed approximation not only offers a regularized framework for describing phase transitions but also provides, from a rigorous mathematical viewpoint, a definition of admissible solutions related to phase transitions. The detailed proof relies on the artificial viscosity method, the calculus of variations, the anti-derivative technique, phase-plane analysis, and the level-set method.

math.AP

Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling

The recent surge in popularity of Nano-Banana and Seedream 4.0 underscores the community's strong interest in multi-image composition tasks. Compared to single-image editing, multi-image composition presents significantly greater challenges in terms of consistency and quality, yet existing models have not disclosed specific methodological details for achieving high-quality fusion. Through statistical analysis, we identify Human-Object Interaction (HOI) as the most sought-after category by the community. We therefore systematically analyze and implement a state-of-the-art solution for multi-image composition with a primary focus on HOI-centric tasks. We present Skywork UniPic 3.0, a unified multimodal framework that integrates single-image editing and multi-image composition. Our model supports an arbitrary (1~6) number and resolution of input images, as well as arbitrary output resolutions (within a total pixel budget of 1024x1024). To address the challenges of multi-image composition, we design a comprehensive data collection, filtering, and synthesis pipeline, achieving strong performance with only 700K high-quality training samples. Furthermore, we introduce a novel training paradigm that formulates multi-image composition as a sequence-modeling problem, transforming conditional generation into unified sequence synthesis. To accelerate inference, we integrate trajectory mapping and distribution matching into the post-training stage, enabling the model to produce high-fidelity samples in just 8 steps and achieve a 12.5x speedup over standard synthesis sampling. Skywork UniPic 3.0 achieves state-of-the-art performance on single-image editing benchmark and surpasses both Nano-Banana and Seedream 4.0 on multi-image composition benchmark, thereby validating the effectiveness of our data pipeline and training paradigm. Code, models and dataset are publicly available.

cs.CV

P-norm based Fractional-Order Robust Subband Adaptive Filtering Algorithm for Impulsive Noise and Noisy Input

Building upon the mean p-power error (MPE) criterion, the normalized subband p-norm (NSPN) algorithm demonstrates superior robustness in $\alpha$-stable noise environments ($1 < \alpha \leq 2$) through effective utilization of low-order moment hidden in robust loss functions. Nevertheless, its performance degrades significantly when processing noise input or additive noise characterized by $\alpha$-stable processes ($0 < \alpha \leq 1$). To overcome these limitations, we propose a novel fractional-order NSPN (FoNSPN) algorithm that incorporates the fractional-order stochastic gradient descent (FoSGD) method into the MPE framework. Additionally, this paper also analyzes the convergence range of its step-size, the theoretical domain of values for the fractional-order $\beta$, and establishes the theoretical steady-state mean square deviation (MSD) model. Simulations conducted in diverse impulsive noise environments confirm the superiority of the proposed FoNSPN algorithm against existing state-of-the-art algorithms.

eess.SP

A Data Annotation Requirements Representation and Specification (DARS)

With the rise of AI-enabled cyber-physical systems, data annotation has become a critical yet often overlooked process in the development of these intelligent information systems. Existing work in requirements engineering (RE) has explored how requirements for AI systems and their data can be represented. However, related interviews with industry professionals show that data annotations and their related requirements introduce distinct challenges, indicating a need for annotation-specific requirement representations. We propose the Data Annotation Requirements Representation and Specification (DARS), including an Annotation Negotiation Card to align stakeholders on objectives and constraints, and a Scenario-Based Annotation Specification to express atomic and verifiable data annotation requirements. We evaluate DARS with an automotive perception case related to an ongoing project, and a mapping against 18 real-world data annotation error types. The results suggest that DARS mitigates root causes of completeness, accuracy, and consistency annotation errors. By integrating DARS into RE, this work improves the reliability of safety-critical systems using data annotations and demonstrates how engineering frameworks must evolve for data-dependent components of today's intelligent information systems.

cs.SE

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real tool-execution traces. To address these limitations, we present Skywork-R1V4, a 30B (A3B) parameter multimodal agentic model that unifies multimodal planning, active image manipulation ("thinking with images"), deep multimodal search, and, most critically, interleaved reasoning that dynamically alternates between visual operations and external knowledge retrieval. Trained solely via supervised fine-tuning on fewer than 30,000 high-quality, planning-execution-consistent trajectories and validated through stepwise consistency filtering, Skywork-R1V4 achieves state-of-the-art results across perception and multimodal search benchmarks: it scores 66.1 on MMSearch and 67.2 on FVQA, surpassing Gemini 2.5 Flash on all 11 metrics. Skywork-R1V4 exhibits emergent long-horizon reasoning at inference time, successfully orchestrating more than 10 tool calls to solve complex, multi-step tasks. Our results demonstrate that sophisticated agentic multimodal intelligence can be achieved through carefully curated supervised learning alone, without any reliance on reinforcement learning.

cs.CV

From Machine Learning Documentation to Requirements: Bridging Processes with Requirements Languages

In software engineering processes for machine learning (ML)-enabled systems, integrating and verifying ML components is a major challenge. A prerequisite is the specification of ML component requirements, including models and data, an area where traditional requirements engineering (RE) processes face new obstacles. An underexplored source of RE-relevant information in this context is ML documentation such as ModelCards and DataSheets. However, it is uncertain to what extent RE-relevant information can be extracted from these documents. This study first investigates the amount and nature of RE-relevant information in 20 publicly available ModelCards and DataSheets. We show that these documents contain a significant amount of potentially RE-relevant information. Next, we evaluate how effectively three established RE representations (EARS, Rupp's template, and Volere) can structure this knowledge into requirements. Our results demonstrate that there is a pathway to transform ML-specific knowledge into structured requirements, incorporating ML documentation in software engineering processes for ML systems.

cs.SE

PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments

Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in static, fully observable settings, limiting their effectiveness in real-world environments where information is often incomplete due to occlusion or limited field of view. Humans, in contrast, actively explore and interact with their environment-moving, examining, and manipulating objects-to gather information through a closed-loop process integrating perception, reasoning, and action. Inspired by this human capability, we introduce the Active Visual Reasoning (AVR) task, extending visual reasoning to partially observable, interactive environments. AVR necessitates agents to: (1) actively acquire information via sequential physical actions, (2) integrate observations across multiple steps for coherent reasoning, and (3) dynamically adjust decisions based on evolving visual feedback. To rigorously evaluate AVR, we introduce CLEVR-AVR, a simulation benchmark featuring multi-round interactive environments designed to assess both reasoning correctness and information-gathering efficiency. We present AVR-152k, a large-scale dataset that offers rich Chain-of-Thought (CoT) annotations detailing iterative reasoning for uncertainty identification, action-conditioned information gain prediction, and information-maximizing action selection, crucial for training agents in a higher-order Markov Decision Process. Building on this, we develop PhysVLM-AVR, an MLLM achieving state-of-the-art performance on CLEVR-AVR, embodied reasoning (OpenEQA, RoboVQA), and passive visual reasoning (GeoMath, Geometry30K). Our analysis also reveals that current embodied MLLMs, despite detecting information incompleteness, struggle to actively acquire and integrate new information through interaction, highlighting a fundamental gap in active reasoning capabilities.

cs.CV

Confinement reduces surface accumulation of swimming bacteria

Many swimming bacteria naturally inhabit confined environments, yet how confinement influences their swimming behaviors remains unclear. Here, we combine experiments, continuum modeling and particle-based simulations to investigate near-surface bacterial swimming in dilute suspensions under varying confinement. Confinement reduces near-surface accumulation and facilitates bacterial escape. These effects are quantitatively captured by models incorporating the force quadrupole, a higher-order hydrodynamic singularity, that generates a rotational flow reorienting bacteria away from surfaces. Under strong confinement, bacterial trajectories straighten due to the balancing torques exerted by opposing surfaces. These findings highlight the role of hydrodynamic quadrupole interactions in near-surface bacterial motility, with implications for microbial ecology, infection control, and industrial applications.

physics.bio-ph

Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model

Recent advances in multimodal models have demonstrated impressive capabilities in unified image generation and editing. However, many prominent open-source models prioritize scaling model parameters over optimizing training strategies, limiting their efficiency and performance. In this work, we present UniPic2-SD3.5M-Kontext, a 2B-parameter DiT model based on SD3.5-Medium, which achieves state-of-the-art image generation and editing while extending seamlessly into a unified multimodal framework. Our approach begins with architectural modifications to SD3.5-Medium and large-scale pre-training on high-quality data, enabling joint text-to-image generation and editing capabilities. To enhance instruction following and editing consistency, we propose a novel Progressive Dual-Task Reinforcement strategy (PDTR), which effectively strengthens both tasks in a staged manner. We empirically validate that the reinforcement phases for different tasks are mutually beneficial and do not induce negative interference. After pre-training and reinforcement strategies, UniPic2-SD3.5M-Kontext demonstrates stronger image generation and editing capabilities than models with significantly larger generation parameters-including BAGEL (7B) and Flux-Kontext (12B). Furthermore, following the MetaQuery, we connect the UniPic2-SD3.5M-Kontext and Qwen2.5-VL-7B via a connector and perform joint training to launch a unified multimodal model UniPic2-Metaquery. UniPic2-Metaquery integrates understanding, generation, and editing, achieving top-tier performance across diverse tasks with a simple and scalable training paradigm. This consistently validates the effectiveness and generalizability of our proposed training paradigm, which we formalize as Skywork UniPic 2.0.

cs.CV

Non-uniqueness of weak solutions to the 3D Hall-MHD equations on the plane

We prove the non-uniqueness of weak solutions with non-trivial magnetic fields to the 3D Hall-MHD equations on the plane in the space $C^0_t L_x^2$ through the convex integration scheme and by constructing new errors and new intermittent flows. In particular, based on the construction of 3D intermittent flows, we obtain the $2\frac{1}{2}$D Mikado flows through a projection onto the plane. Moreover, we prove that the constructed weak solution do not conserve the magnetic helicity and find that weak solutions of the ideal Hall-MHD equations in $C^{\bar{\beta}}_{t,x}$ ($\bar{\beta}>0$) are the strong vanishing viscosity and resistive limit of weak solutions to the Hall-MHD equations.

math.AP