SearcharxivSearch

arXiv subjects

Jun Zheng

Publications and source records attributed to Jun Zheng.

At least 19 recordsLinked to original sources

MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downstream utility, frequently introducing outdated, misleading, or conflicting evidence into the active context. We present MeClear, a task conditioned memory clearance framework that identifies memories featuring negative downstream utility through cooperative attribution and selectively suppresses them from agent execution. MeClear combines Leave One Out screening with sampled cooperative Shapley attribution to distribute utility across interacting evidence, effectively resolving redundant conflict masking where single removal evaluations fail. Utilizing attribution rankings, MeClear executes a query scoped minimal clearance strategy over a nested filtration, verifying task recovery on the cleared context without permanently altering the persistent memory bank. Comprehensive experimental evaluations across ten long dialogue memory pools demonstrate that MeClear achieves a target recall of 85.9% and an overall task recovery rate of 82.3%, representing a 25.5 percentage point improvement over Leave One Out (LOO) baselines.

cs.AI

Input-to-State Stabilization of a Coupled ODE-PDE System with Time-Varying Coefficients via Composite Boundary Control

This paper proposes a novel composite boundary feedback control law that ensures input-to-state stability (ISS) for a coupled ODE-parabolic PDE system with time-varying coefficients in both subsystems. In controller design, we circumvent the need to directly solve coupled time-varying parabolic-hyperbolic kernel equations by employing an analytic pre-defined gain function and a time-varying Volterra kernel function to design the control law explicitly. In stability analysis, to address the simultaneous challenges of Dirichlet boundary disturbances and time-varying coefficients, we employ the square root of a time-varying positive definite matrix and a superlinear function to construct a nonquadratic Lyapunov function for the ODE and a generalized Lyapunov functional for the PDE, respectively, in the target system, thereby establishing the ISS in the $L^2$-norm of the closed-loop system. Numerical simulations are presented to illustrate the effectiveness of the proposed control scheme.

math.OC

Exploring the Performance Frontier of Compact Unified Image Generation Models

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.

cs.CV

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

cs.CV

CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and introduces multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings, serving as an effective proxy for public human evaluations.

cs.CV

Polynomial iISS for a class of Timoshenko equations with external disturbances and infinite history memory

This paper investigates the robustness of Timoshenko equations subject to external disturbances and infinite history memory within the integral input-to-state stability (iISS) framework. While polynomial iISS (PiISS) is a powerful tool for characterizing the influence of these two factors, the existing definitions of PiISS rely on graph norms of the state, which are inconsistent with the classical iISS notion for infinite-dimensional systems. In this paper, we provide a rigorous PiISS definition for a class of Timoshenko equations solely via the state norm, and derive sufficient conditions for PiISS under the Tolksdorf condition on the relaxation function, as well as exponential iISS (EiISS) under a stronger condition. The well-posedness is established by using the semigroup and elliptic equation theories. The PiISS and EiISS are assessed by constructing appropriate Lyapunov functionals and employing various a priori estimates of solutions, thereby fully characterizing the impact of disturbances and history memory on system stability.

math.OC

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This limitation overlooks a crucial aspect of real-world apparel presentation: active human-garment interaction. To bridge this gap, we introduce and formalize a new challenging task: Interactive Video Virtual Try-On (Interactive VVT), where subjects in the video actively engage with their clothing. This task introduces unique challenges beyond simple texture preservation, including: (1) resolving the semantic ambiguity of interactions from standard pose information, and (2) learning complex garment deformations from video where interactive moments are sparse and brief. To address these challenges, we propose iTryOn, a novel framework built upon a large-scale video diffusion Transformer. iTryOn pioneers a multi-level interaction injection mechanism to guide the generation of complex dynamics. At the spatial level, we introduce a garment-agnostic 3D hand prior to provide fine-grained guidance for precise hand-garment contact, effectively resolving spatial ambiguity. At the semantic level, iTryOn leverages global captions for overall context and time-stamped action captions for localized interactions, synchronized via our novel Action-aware Rotational Position Embedding (A-RoPE). Extensive experiments demonstrate that iTryOn not only achieves state-of-the-art performance on traditional VVT benchmarks but also establishes a commanding lead in the new interactive setting, marking a significant step towards more dynamic and controllable virtual try-on experiences.

cs.CV

Robust synchronization for multi-agent systems governed by PDEs with observable and unobservable disturbances

This paper investigates robust synchronization for multi-agent systems (MASs) governed by parabolic partial differential equations in the presence of both observable and unobservable disturbances. Using only boundary output measurements, a disturbance observer is designed to estimate observable Dirichlet boundary disturbances while ensuring robustness of the observer error system with unobservable disturbances occurring in the domain. Using only the reference signal and local output information, distributed synchronization controllers are then constructed to enable all agents to track the reference trajectory. In particular, exponential tracking is achieved in the absence of unobservable disturbances, while robustness is preserved when additional unobservable disturbances occur during controller implementation. We further analyze the impact of unobservable Dirichlet-Robin boundary disturbances on synchronization performance by proving the boundedness of solutions to the synchronization error system. Moreover, to characterize the influence of all disturbances, input-to-state stability (ISS) is established for the closed-loop system. For the involved systems, the generalized Lyapunov method and the recursion technique are extensively employed in the stability analysis, and the lifting technique and semigroup theory are used to prove the well-posedness. Simulation results validate the proposed control scheme, demonstrating effective disturbance estimation and rejection, robust synchronization, and the ISS properties under various scenarios.

eess.SY

Unified Lyapunov Method for ISS of PDEs: A Tutorial on Constructing Generalized Lyapunov Functionals for Parabolic and Hyperbolic Equations

This tutorial provides an overview of the generalized Lyapunov method (GLM) for analyzing input-to-state stability (ISS) of partial differential equations (PDEs). We begin by revisiting the classical Lyapunov method and the standard ISS-Lyapunov theorem, highlighting their limitations when applied to systems with complex boundary disturbances. In contrast, the GLM, based on the concept of generalized Lyapunov functionals (GLFs) that explicitly depend on the external input, offers greater flexibility and efficiency, particularly for PDEs with Dirichlet-type disturbances. The main objective of this tutorial is to demonstrate how to systematically construct GLFs to establish ISS estimates in $L^q$ spaces with any $q\in[2,\infty]$ for different PDEs. Specifically, we consider three representative classes of PDEs: (i) an $N$-dimensional nonlinear parabolic equation with mixed nonlinear boundary disturbances, (ii) a first order nonlinear hyperbolic equation with boundary disturbances, and (iii) a second order linear hyperbolic equation, i.e., a wave equation, with boundary damping and disturbances. For each case, we provide step-by-step constructions of appropriate GLFs and derive explicit ISS estimates, illustrating the general applicability of the GLM. Finally, we discuss open challenges and future directions, including the systematic construction of GLFs for broader classes of PDEs and their applications in controller design.

math.OC

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Recent advances in image generation and editing have opened new opportunities for virtual try-on. However, existing methods still struggle to meet complex real-world demands. We present Tstars-Tryon 1.0, a commercial-scale virtual try-on system that is robust, realistic, versatile, and highly efficient. First, our system maintains a high success rate across challenging cases like extreme poses, severe illumination variations, motion blur, and other in-the-wild conditions. Second, it delivers highly photorealistic results with fine-grained details, faithfully preserving garment texture, material properties, and structural characteristics, while largely avoiding common AI-generated artifacts. Third, beyond apparel try-on, our model supports flexible multi-image composition (up to 6 reference images) across 8 fashion categories, with coordinated control over person identity and background. Fourth, to overcome the latency bottlenecks of commercial deployment, our system is heavily optimized for inference speed, delivering near real-time generation for a seamless user experience. These capabilities are enabled by an integrated system design spanning end-to-end model architecture, a scalable data engine, robust infrastructure, and a multi-stage training paradigm. Extensive evaluation and large-scale product deployment demonstrate that Tstars-Tryon1.0 achieves leading overall performance. To support future research, we also release a comprehensive benchmark. The model has been deployed at an industrial scale on the Taobao App, serving millions of users with tens of millions of requests.

cs.CV

A Spectral-based ISS small-gain theorem for boundary control systems with infinite couplings

We study the input-to-state stability (ISS) of boundary control systems allowing for infinitely many boundary couplings. Using semigroup perturbation theory and the theory of positive linear operators on Banach lattices, we derive a spectral small-gain condition ensuring exponential ISS. We further investigate linear Boltzmann-type equations on an infinite network of intersecting circles, incorporating delays, scattering, and disturbances acting at the junction. For this class of systems, we prove that a spectral small-gain condition on the transmission operator matrix guarantees exponential ISS with respect to disturbances propagating through the network. Moreover, we derive explicit ISS estimates for {certain} classes of dynamical processes. Finally, we demonstrate the practical applicability of our results by considering two important classes of time-delayed transmission conditions.

math.OC

Mode-Resolved Multiband Ballistic Transport and Conductance Thresholds in Bilayer Graphene Junctions

We study ballistic transport in bilayer graphene junctions and show how electrostatic gating, interlayer bias, and homogeneous strain provide complementary control over electron transmission. In the absence of strain, transport is governed by symmetry constraints that suppress transmission at specific incidence angles despite the availability of states. An interlayer bias lifts this suppression through mode mixing and opens a tunable transport gap. Within a full four-band description, we identify a distinct conductance threshold that marks the onset of propagation of the upper band inside the barrier. This produces a clear change in the slope of the conductance and serves as an experimentally accessible transport fingerprint of the multiband structure and interlayer coupling. Homogeneous in-plane strain acts as a geometric control mechanism. By reshaping the band structure in momentum space, it redistributes the angular transmission window and suppresses conductance without introducing disorder. Importantly, strain preserves the underlying symmetry-based decoupling responsible for transmission suppression while shifting its condition away from normal incidence. These results provide a unified framework for interpreting angle-resolved transport in bilayer graphene and establish multiband ballistic transport as a practical probe of band-structure geometry.

cond-mat.mes-hall

Hausdorff measure of the free boundary for the $p$-obstacle problem with subcritical exponents

This paper investigates a class of $p$-obstacle problems with subcritical exponents having the form \begin{align} \mathrm{div}\left( a(x)|\nabla u|^{p-2}\nabla u\right) =m_1\chi_{\{u>0\}}-m_2u^{\lambda-1}\chi_{\{u>0\}} \ \text{in}\ \Omega,\notag \end{align} where $\Omega $ is a smooth bounded domain in $ \mathbb{R}^N (N \geq 2)$, $m_1,m_2$ are positive constants, the coefficient function $a \in C^2(\Omega)$ has a positive lower bound, and $2 \leq p < \lambda <p^*:= \frac{Np}{N-p}$ when $p<N$ and $N \geq 3$, or $2\leq p < \lambda <+\infty$ when $ N = 2$. By using the mountain-pass lemma, combined with the penalty method, we first establish the existence of non-negative weak solutions. Then, using the De Giorgi-Nash iteration, we prove the $L^\infty$ bound and local $C^{1,\alpha}$ continuity for the solutions. In addition, we prove local porosity of the free boundary based on the optimal growth and non-degeneracy of solutions near the free boundary. Furthermore, by means of Lebesgue measure estimates for gradient level sets, we show that at least one solution corresponds the free boundary having locally finite $(N-1)$-dimensional Hausdorff measure.

math.AP

Well-posedness of boundary control systems and application to ISS for coupled heat equations with boundary disturbances and delays

This paper studies the existence of solutions and, in particular, the well-posedness of a class of boundary control systems. Our main result provides explicit and verifiable conditions on the system data that guarantee continuous dependence of solutions on the initial data and $L^p$-inputs. The proof relies on a new boundedness estimate for the input/output maps of linear time-invariant infinite-dimensional systems with unbounded control and observation operators. The developed technique is applied to derive specific conditions for the exponential input-to-state stability of boundary-coupled heat equations with boundary disturbances and time-delays.

math.OC

Look Forward to Walk Backward: Efficient Terrain Memory for Backward Locomotion with Forward Vision

Legged robots with egocentric forward-facing depth cameras can couple exteroception and proprioception to achieve robust forward agility on complex terrain. When these robots walk backward, the forward-only field of view provides no preview. Purely proprioceptive controllers can remain stable on moderate ground when moving backward but cannot fully exploit the robot's capabilities on complex terrain and must collide with obstacles. We present Look Forward to Walk Backward (LF2WB), an efficient terrain-memory locomotion framework that uses forward egocentric depth and proprioception to write a compact associative memory during forward motion and to retrieve it for collision-free backward locomotion without rearward vision. The memory backbone employs a delta-rule selective update that softly removes then writes the memory state along the active subspace. Training uses hardware-efficient parallel computation, and deployment runs recurrent, constant-time per-step inference with a constant-size state, making the approach suitable for onboard processors on low-cost robots. Experiments in both simulations and real-world scenarios demonstrate the effectiveness of our method, improving backward agility across complex terrains under limited sensing.

cs.RO

Input-to-state stabilization of an ODE cascaded with a parabolic equation involving Dirichlet-Robin boundary disturbances

This paper focuses on the input-to-state stabilization problem for an ordinary differential equation (ODE) cascaded by parabolic partial differential equation (PDE) in the presence of Dirichlet-Robin boundary disturbances, as well as in-domain disturbances. For the cascaded system with a Dirichlet pointwise interconnection, the ODE takes the value of a Robin boundary condition at the ODE-PDE interface as its direct input, and the PDE is driven by a Dirichlet boundary input at the opposite end. We first employ the backstepping method to design a boundary controller and to decouple the cascaded system. This decoupling facilitates independent stability analysis of the PDE and ODE systems sequentially. Then, to address the challenges posed by Dirichlet boundary disturbances to the application of the classical Lyapunov method, we utilize the generalized Lyapunov method to establish the ISS in the max-norm for the cascaded system involving Dirichlet boundary disturbances and two other types of disturbances. The obtained result indicates that even in the presence of different types of disturbances, ISS analysis can still be conducted within the framework of Lyapunov stability theory. For the well-posedness of the target system, it is conducted by using the technique of lifting and the semigroup method. Finally, numerical simulations are conducted to illustrate the effectiveness of the proposed control scheme and ISS properties for a cascaded system with different disturbances.

math.OC

CI4A: Semantic Component Interfaces for Agents Empowering Web Automation

While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused on enhancing model grounding capabilities through techniques such as Reinforcement Learning. However, rather than compelling agents to adapt to human-centric interfaces, we propose constructing interaction interfaces specifically optimized for agents. This paper introduces Component Interface for Agent (CI4A), a semantic encapsulation mechanism that abstracts the complex interaction logic of UI components into a set of unified tool primitives accessible to agents. We implemented CI4A within Ant Design, an industrial-grade front-end framework, covering 23 categories of commonly used UI components. Furthermore, we developed a hybrid agent featuring an action space that dynamically updates according to the page state, enabling flexible invocation of available CI4A tools. Leveraging the CI4A-integrated Ant Design, we refactored and upgraded the WebArena benchmark to evaluate existing SoTA methods. Experimental results demonstrate that the CI4A-based agent significantly outperforms existing approaches, achieving a new SoTA task success rate of 86.3%, alongside substantial improvements in execution efficiency.

cs.AI

Mode-selective cloaking and phase-matching cavity resonances in bilayer graphene transport

We study ballistic electron transport through electrostatic barriers in AB-stacked bilayer graphene within a full four-band framework. A mode-resolved analysis reveals how propagating and evanescent channels couple across electrostatic interfaces and how channel selectivity governs transport at normal incidence. We show that perfect transmission can occur at discrete energies due to phase matching of a single internal mode within an individual barrier, without activating the decoupled channels. This effect is interpreted as a phase-matching cavity, namely, an effective cavity formed by internal phase coherence inside the barrier, which yields perfect transmission at discrete energies without true bound states and without opening additional transport channels. For single- and double-barrier geometries, we derive compact analytical expressions for the transmission and identify the corresponding resonance conditions. Extending the analysis to multibarrier structures using a transfer-matrix approach, we demonstrate how perfect resonances driven by internal phase matching coexist with Fabry-Perot-type resonances arising from interbarrier interference. Our results provide a unified, channel-resolved description of tunneling suppression and resonance-assisted transport in bilayer graphene barrier systems.

cond-mat.mes-hall