Searcharxiv⌕ Search

arXiv subjects

Cheng Chen

Publications and source records attributed to Cheng Chen.

At least 55 records · Page 3Linked to original sources

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/

cs.CV↗

Symmetric Rank-$k$ Methods

This paper proposes a novel class of block quasi-Newton methods for convex optimization which we call symmetric rank-$k$ (SR-$k$) methods. Each iteration of SR-$k$ incorporates the curvature information with~$k$ Hessian-vector products achieved from the greedy or random strategy. We prove that SR-$k$ methods have the local superlinear convergence rate of $\mathcal{O}\big((1-k/d)^{t(t-1)/2}\big)$ for minimizing smooth and strongly convex function, where $d$ is the problem dimension and $t$ is the iteration counter. This is the first explicit superlinear convergence rate for block quasi-Newton methods, and it successfully explains why block quasi-Newton methods converge faster than ordinary quasi-Newton methods in practice. We also leverage the idea of SR-$k$ methods to study the block BFGS and block DFP methods, showing their superior convergence rates.

math.OC↗

Role of $K^*_0(700)$ exchange in the $p \bar{p} \to Λ\barΛ$ reaction

Based on the effective Lagrangian approach, we investigate the $p \bar{p} \to Λ\barΛ$ reaction. Within this framework, we provide a dynamical explanation by analyzing its total and differential cross sections, as well as the polarization of the produced $Λ$ hyperon. Incorporating the $t$-channel exchange of the scalar $K^*_0(700)$ and pseudoscalar $K$ mesons, complemented by an $s$-channel contribution from the vector excited state, we can reproduce the current experimental data fairly well in a wide energy region. Compared to the conventional $K$ and $K^*(892)$ mesons exchange, the $K^*_0(700)$ meson exchange plays a more essential role in simultaneously capturing the observed features of the total and differential cross sections. The introduction of the vector $s$-channel resonance improves the description of the spin observables. This work gives a perspective to inspect the role of $K^*_0(700)$ and serves as a test to search for the resonances in the reaction $p \bar{p} \to Λ\barΛ$ at threshold.

hep-ph↗

Blind Gradient-Ascent Phase Alignment for Multi-Aperture Coherent Digital Combining Under Aperture-Dependent Phase Disturbance

Multi-aperture reception can provide spatial diversity in free-space optical (FSO) communication by collecting signal replicas at separate apertures. When the branches are accurately phase-aligned, their received optical fields can also be added constructively to obtain coherent-combining gain. In this paper, we propose blind gradient-ascent phase alignment (BGAPA), which iteratively adjusts one phase correction per aperture by directly maximizing the combined output power. Closed-form analytical gradients provide a deterministic update that requires no symbol decisions, unlike the stochastic perturbation-based estimate of SPGD or the decision-directed feedback of DD-LMS. To isolate phase-tracking capability, the numerical model includes independent aperture-dependent phase disturbance but excludes amplitude scintillation and polarization-dependent distortion. Under this controlled phase-only setting, BGAPA obtains an SNR improvement closer to the ideal 6.02~dB coherent-combining gain than block-wise cross-correlation, SPGD, DD-LMS, and CMA/RDE-based equalization when the aperture count is increased by a factor of four. In particular, increasing the aperture count from 64 to 256 yields an SNR improvement of about 5.7~dB. In a separate amplitude-tolerance test with $N=16$ and $f_{\max}=1$~MHz, the first observed BGAPA trial above the HD-FEC threshold of $3.8\times10^{-3}$ occurs at an actual phase RMS of approximately 278~rad, whereas DD-LMS becomes unreliable at substantially smaller phase excursions. The reported step size is optimized separately at each operating point. BGAPA is fully blind and updates its phase parameters directly from the received aperture fields without training symbols, pilots, or decision-directed feedback.

physics.optics↗

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0. OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%. These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.

cs.AI↗

Many-Body Physics with Rydberg Atoms: Quantum Simulation and Non-equilibrium Dynamics

Rydberg atoms, characterized by their strong and long-range dipole-dipole interactions, provide a versatile platform for exploring intriguing collective and many-body effects. Recently, the experimental realization of these effects in dense ensembles and reconfigurable atomic arrays has attracted significant interest, particularly for applications in quantum simulations and non-equilibrium physics. This review focuses on such recent development, discussing the theoretical foundations of the interactions between Rydberg atoms and the ensuing many-body physics, while providing a critical survey of experimental techniques for their precise manipulation and observation. We further discuss recent breakthroughs in leveraging Rydberg collective effects to probe novel many-body phases and non-equilibrium dynamics of these systems. By synthesizing theoretical insights with experimental milestones, we provide a comprehensive perspective on this rapidly evolving field and its transformative potential for future quantum technologies.

quant-ph↗

Heterogeneous-Gradient Phase--Polarization Alignment and Maximal-Ratio Weight Allocation for Multi-Aperture Coherent FSO Reception

Multi-aperture coherent reception can improve freespace optical (FSO) links by converting spatial diversity into coherent combining gain. In turbulent links, the aperture branches are simultaneously affected by relative phase errors, polarization mismatch, and unequal signal-to-noise ratios (SNRs). Existing methods treat phase/polarization alignment and branch-weight allocation as separate operations, or absorb all impairments into a high-dimensional MIMO equalizer that obscures the physical meaning of each aperture's contribution. This paper proposes a structured blind combining method based on heterogeneous gradient sources: phase and per-aperture polarization parameters are updated by closed-form analytical gradients that maximize the combined output power, while aperture weights and an optional global polarization angle are updated by gradients derived from the constellation-radius error. An exponential parameterization pn = eqn/N ensures positivity without clipping. The internal variable qn is adapted by radius-error gradients, thereby allocating maximal-ratio-combining-like weights according to the quality of the already aligned branches.

physics.optics↗

Dirac Spin Liquid Candidate in a Rydberg Quantum Simulator

We experimentally investigate a frustrated spin-exchange antiferromagnet in a quantum simulator, composed of N = 114 dipolar Rydberg atoms arranged into a kagome array. Motivated by a recent theoretical proposal of a gapless U(1) Dirac spin liquid ground state, we use local addressing to adiabatically prepare low-energy states. We measure the local polarization and spin-spin correlations over this adiabatic protocol, and observe our system move from a staggered product state, through an intermediate magnetic crystal, and finally into a disordered, correlated liquid. We estimate the entropy density of this atomic liquid to be similar to that of frustrated magnetic insulators at liquid nitrogen temperatures. We compare the correlations in our liquid to those of a simple, parameter-free ansatz for the Dirac spin liquid, and find good agreement in the sign structure and spatial decay. Finally, we probe the static susceptibility of our system to a local field perturbation and to a geometrical distortion. Our results establish Rydberg atom arrays as a promising platform for the preparation and microscopic characterization of quantum spin liquid candidates.

cond-mat.quant-gas↗

Archetypal Microbiome Profiles as Indicators of Nitrous Oxide Emission States in Activated Sludge

Nitrous oxide (N2O) emissions from water resource recovery facilities (WRRFs) fluctuate over time and can arise from multiple microbial pathways, making source attribution and full-scale prediction difficult. The difficulty is compounded by the high dimensionality of activated sludge microbiomes, whose complex and dynamic community structure can obscure relationships with N2O emission patterns. This study evaluated whether interpretable, low-dimensional representations of activated sludge microbiomes can be correlated with N2O emission states. Temporal 16S rRNA gene amplicon profiles and N2O emission metrics were collected from two full-scale WRRFs in Switzerland. Genus-level relative-abundance profiles were summarized using archetypal analysis (AA), which represents each sample as a convex combination of a small number of interpretable community profiles. In both WRRFs, three archetypes captured most explainable variation in community composition (63%--73%) and defined a simplex state space in which samples clustered near vertices and edges, indicating that community compositions were organized around distinct archetypal states and their mixtures. Without using emission labels while training, the archetypal state space aligned strongly with binary N2O emission states: high-emission observations in both plants concentrated around a specific archetype, and temporal trajectories showed consistent high weights of this archetype during high-emission periods. Functional summaries suggested site-specific but pathway-relevant interpretations of the high-N2O archetype. Temperature further structured the archetypal state space, indicating seasonal forcing of microbiome configurations associated with elevated N2O. Overall, AA provides an interpretable framework to track microbiome regime shifts and may support operational tracking of high-N2O emission states in full-scale WRRFs.

q-bio.QM↗

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion

Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the large language model (LLM). This static architecture fundamentally limits the ability of LLMs to achieve comprehensive alignment with hierarchical visual knowledge, compromising their capacity to accurately integrate local details with global semantics into coherent reasoning. To resolve this, we introduce Cross-Layer Injection (CLI), a novel and lightweight framework that forges a dynamic many-to-many bridge between the two modalities. CLI consists of two synergistic, parameter-efficient components: an Adaptive Multi-Projection (AMP) module that harmonizes features from diverse vision layers, and an Adaptive Gating Fusion (AGF) mechanism that empowers the LLM to selectively inject the most relevant visual information based on its real-time decoding context. We validate the effectiveness and versatility of CLI by integrating it into LLaVA-OneVision and LLaVA-1.5. Extensive experiments on 18 diverse benchmarks demonstrate significant performance improvements, establishing CLI as a scalable paradigm that unlocks deeper multimodal understanding by granting LLMs on-demand access to the full visual hierarchy.

cs.CV↗

High-Precision Calibration Workflow Achieves Above $99.9\%$ CZ Gate Fidelity on a Scalable Superconducting Processor

High-fidelity universal two-qubit gates are critical for building fault-tolerant quantum computers. In scalable superconducting processors, shortened coherence times introduce more incoherent errors in gate operations. With a constrained error budget, there is reduced tolerance for coherent errors stemming from parameter deviations. In this work, we develop a closed-loop workflow to enhance the CZ gate calibration precision. Utilizing the echoed leakage error amplification (ELEA) and the repurposed context-aware fidelity estimation (CAFE) circuits, we suppress the population leakage to non-computational states, and, for the first time, demonstrate a CZ gate fidelity exceeding $99.9\%$ on an 84-qubit processor, with coherent error suppressed to $0.007\%$. Meanwhile, we obtain a median fidelity of $99.25\%$ among 72 CZ gates, demonstrating that the workflow can be generalized to the calibration of parallel CZ gates. Finally, we realize automated calibration and observe enhanced stability of the CZ gate throughout 9-hour comparative monitoring experiments. Our results, realized on a completely domestic platform, establish an efficient and automated route to quantum computation with superconducting quantum systems.

quant-ph↗

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two scenarios: in-domain, which requires retaining the reference subject features as much as possible, and cross-domain, which preserves the intrinsic features of the subject while allowing subject-irrelevant properties to vary flexibly according to the text prompt. Existing methods primarily focus on maximizing subject fidelity in in-domain scenarios, which limits their editability and adaptability in cross-domain scenarios, such as novel styles, semantic combinations, or domain attributes. In this study, we propose that an ideal S2V method should flexibly shuttle between different domains, achieving strong performance in both in-domain and cross-domain scenarios. To this end, we propose DomainShuttle, which could achieve high fidelity and generative flexibility for open domain video personalization. Specifically, we introduce Domain-MoT, which decouples videos and reference features and introduces the domain-aware AdaLN for domain-specific modeling of reference images. We then introduce the Video-Reference DualRoPE scheme, which places reference image tokens and video tokens in separate RoPE spaces to enable precise subject-level spatial modeling, and Cross-Pair Consistent Loss, which aims to extract intrinsic subject features unaffected by irrelevant features. Extensive experiments demonstrate that DomainShuttle achieves significant performance improvements over existing methods, exhibiting high subject fidelity and generative flexibility across diverse open domain application scenarios.

cs.CV↗

Maintaining Leiden Communities in Large Dynamic Graphs

Community detection is a foundational capability in large-scale industrial graph analytics, powering applications such as fraud-ring discovery, recommendation systems, and hierarchical indexing for retrieval-augmented generation. Among modularity-based methods, the Leiden algorithm has been widely adopted in production because it delivers high-quality communities with connectivity guarantees. However, real-world graphs evolve continuously, and timely community updates are needed to keep downstream features and retrieval indices fresh. Meanwhile, existing dynamic Leiden approaches recompute the communities whenever their vertices and edges change, thereby almost degrading to near-full recomputation under frequent updates. To alleviate the efficiency issue, we study the efficient maintenance of Leiden communities in large dynamic graphs and present a novel algorithm, called Hierarchical Incremental Tree Leiden (HIT-Leiden). We first provide a boundedness analysis showing that prior incremental Leiden methods can incur essentially unbounded work even for small updates. Guided by this analysis, we propose HIT-Leiden which effectively reduces the range of affected vertices by maintaining connected components and hierarchical community structures. Extensive experiments on large real-world dynamic graphs demonstrate that HIT-Leiden achieves community quality comparable to the state-of-the-art competitors while delivering speedups of up to five orders of magnitude over existing solutions. The production deployment results show that HIT-Leiden meets stringent latency requirements under high-rate updates at scale.

cs.SI↗

SAMTok: Representing Any Mask with Two Words

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-level encoders, specialized segmentation decoders, and incompatible training objectives. To address these challenges, we present SAMTok, a discrete mask tokenizer that converts any region mask into two special tokens and reconstructs the mask using these tokens with high fidelity. By treating masks as new language tokens, SAMTok enables base MLLMs (such as the QwenVL series) to learn pixel-wise capabilities through standard next-token prediction and simple reinforcement learning, without architectural modifications and specialized loss design. SAMTok builds on SAM2 and is trained on 209M diverse masks using a mask encoder and residual vector quantizer to produce discrete, compact, and information-rich tokens. With 5M SAMTok-formatted mask understanding and generation data samples, QwenVL-SAMTok attains state-of-the-art or comparable results on region captioning, region VQA, grounded conversation, referring segmentation, scene graph parsing, and multi-round interactive segmentation. We further introduce a textual answer-matching reward that enables efficient reinforcement learning for mask generation, delivering substantial improvements on GRES and GCG benchmarks. Our results demonstrate a scalable and straightforward paradigm for equipping MLLMs with strong pixel-wise capabilities. Our code and models are available.

cs.CV↗

Adapting Vision-Language Models from Iconic to Inclusive for Multi-Label Recognition Without Labels

Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision-language models (VLMs) enable zero-shot recognition without labeled data. However, due to their intrinsic design, these models often prioritize the most iconic object and omit other contextual positives. This intrinsic bias conflicts with the nature of multi-label learning, thereby limiting their applicability. In this work, we propose an unsupervised framework that adapts VLMs from iconic recognition toward inclusive understanding, enabling label-free multi-label image recognition. Our approach consists of two key stages, ``cutting'' and ``sewing'': In the cutting stage, we present the multi-sampling response estimator to prevent the model from concentrating only on one single object. In the second sewing stage, the multi-object blend adaptation is introduced to adjust the labels to better conform to the multi-label distribution while preserving the intrinsic characteristics of the original model within only one epoch. Extensive experiments show that our framework significantly outperforms existing unsupervised approaches on four public datasets, even surpassing several representative weakly supervised baselines. These results demonstrate the potential of adapting pre-trained VLMs for more comprehensive visual understanding without manual annotations. Our code is publicly available at https://github.com/iCVTEAM/TailorCLIP.

cs.CV↗

Final-state rescattering mechanism of the $Δ(1232)^{++}$ production in $Λ^+_c \to K^- π^+ p$ decay

We investigate the production of the $Δ(1232)^{++}$ resonance in the charmed baryon weak decay $Λ^+_c \to K^- π^+ p$, focusing on the $π^+ p$ final-state rescattering mechanism. The direct $W^+$ exchange diagram is expected to be suppressed, hence we adopt the $W^+$ internal emission process $Λ^+_c \to p \bar K^{*0}(892)$ followed by the subsequent decay $\bar{K}^{*0} \to K^- π^+$ as the dominant source of the final state particles. The $Δ(1232)^{++}$ resonance is then generated via $π^+ p$ rescattering within a triangle loop mechanism. Our calculations incorporate both the tree-level $\bar K^{*0}(892)$ and the dynamically generated $\bar{K}^*_0(700)$ state arising from the $S$-wave $K π$ final state interaction. We find that our theoretical results can reproduce the bump and peak structures in the $K^- π^+$ invariant mass distributions for the $\bar{K}^*_0(700)$ and $\bar{K}^{*0}(892)$, respectively. Meanwhile, the peak for the $Δ(1232)^{++}$ in the $π^+ p$ invariant mass distributions is also well described. The $Δ(1232)^{++}$ signal naturally emerges from rescattering effects, and adopting the pole parameters of $Δ(1232)$ resonance yields an improved description of the experimental data. In addition, we obtain a branching fraction ratio $\mathcal{B}[Λ_c^+ \to Δ(1232)^{++} K^-] / \mathcal{B}[Λ_c^+ \to p \bar{K}^{*0}(892)] \approx 0.5$, which is lower than the experimentally measured value. This discrepancy suggests that interference effects are likely significant in this decay process. Future high-precision measurements will further verify the proposed rescattering mechanism.

hep-ph↗

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness

This paper focuses on the decentralized stochastic optimization problem $f(\mathbf{x})=\frac{1}{m}\sum_{i=1}^m f_i(\mathbf{x})$ over a connected network of $n$ agents, where each local function has the form of $f_i(\mathbf{x}) = {\mathbb E}\left[F(\mathbf{x};{\boldsymbol ξ}_i)\right]$ which satisfies the $(L_0,L_1)$-smooth condition but possibly nonconvex and each random variable ${\boldsymbol ξ}_i$ follows distribution ${\mathcal D}_i$. We propose a novel algorithm called decentralized normalized stochastic gradient descent (DNSGD), which can achieve an $ε$-stationary point at each local agent. We present a new framework for analyzing decentralized first-order methods in the $(L_0,L_1)$-smooth setting, based on the Lyapunov function related to the product of the gradient norm and the consensus error. We show that the proposed algorithm attains the upper bounds on the sample complexity of ${\mathcal O}(m^{-1}(L_fσ^2Δ_fε^{-4} + σ^2ε^{-2} + L_f^{-2}L_1^3σ^2Δ_fε^{-1} + L_f^{-2}L_1^2σ^2))$ per agent and the communication complexity of $\tilde{\mathcal O}((L_fε^{-2} + L_1ε^{-1})γ^{-1/2}Δ_f)$, where $L_f=L_0 +L_1ζ$, $σ^2$ is the variance of the stochastic gradient, $Δ_f$ is the initial optimal function value gap, $γ$ is the spectral gap of the network, and $ζ$ is the degree of the gradient dissimilarity. In the special case of $L_1=0$, the above results (nearly) match the lower bounds of decentralized stochastic nonconvex optimization under the standard smoothness. We also conduct numerical experiments to show the empirical superiority of our method.

math.OC↗