SearcharxivSearch

arXiv subjects

Chao Wang

Publications and source records attributed to Chao Wang.

At least 19 recordsLinked to original sources

Particle Dynamics of Flow Matching and Classifier-Free Guidance from a Stagewise Geometry Perspective

Flow matching, together with classifier-free guidance (CFG), is widely used in generative modeling, yet much of the theoretical understanding remains distribution-wise. Since practical sampling follows individual trajectories, distribution-level guarantees alone do not fully capture how trajectories interact with the data geometry or how guidance reshapes it. To overcome this limitation, we establish a unified stagewise geometric theory of attraction and absorption for both continuous dynamics and explicit Euler discretization. Specifically, with $t\in[0,1]$ running from noise to data, we show that unconditional flow trajectories are successively attracted toward a neighborhood of the global mean, the data convex hull, and a neighborhood of a possibly nonconvex local cluster. Across these stages, the corresponding distance satisfies a common contraction estimate, yielding an ${O}(1-t)$ decay of the distance in the final stage. For CFG, the same structure persists with an extrapolated mean, an inflated conditional convex hull, and, near the target cluster, the restored local geometry of conditional flow matching. We further show that a general time schedule $a(t)$ replaces the $O(1-t)$ decay by $O(1-a(t))$. Together, these results provide a unified particle-level geometric account of flow matching and CFG across continuous and discrete sampling.

cs.LG

CR-VLA-Force: Learning Control-aware Compliance VLA Model for Robust Contact-rich Robotic Manipulation

Integrating visuomotor policies or Vision-Language-Action (VLA) models with force/torque (F/T) perception has demonstrated significant progress in imitation learning for robotic manipulation. However, existing force-aware VLA models frequently exhibit limited capability in precise force tracking and rapid successive adjustments. This deficiency stems from the limitations of action-chunk execution strategies and the substantial latency between perception and real-time control. Such limitations can lead to task failures and safety risks, particularly when the execution of an action chunk exerts excessive interaction forces without timely adjustment. To overcome this challenge, we propose the Control-aware Compliance VLA (CC-VLA) framework for reactive control. The CC-VLA model employs a multimodal mixture-of-experts (MoE) to encode force signal sequences and vision-language fused feature. Furthermore, it utilizes a multi-stage training strategy to ensure robust perception within the visual-semantic space and effective force perception under sparse sampling conditions. Additionally, a VLA-guided adaptive compliance controller is designed to facilitate precise position tracking during contact-free motion and optimal force-position tracking for contact-rich tasks. To facilitate high-precision F/T data acquisition, we also implement an adversaria shared teleoperation strategy for contact-rich demonstrations that bolsters system safety and interactivity. Extensive real-world experiments demonstrate that CC-VLA significantly improves success rates in challenging force-perception tasks and enhances force-control precision, while providing multi-level safety and robustness under the tested partial-OOD pose-shift settings.

cs.RO

The $H^*H^*V$ charge couplings from light-cone sum rules

We present an updated and rigorous determination of the strong charge couplings $g_{H^*H^*V}$ (with $H \in \{D, B\}$ and $V \in \{\rho, \omega, K^*, \phi\}$) using the framework of light-cone sum rules (LCSR). The theoretical precision is significantly enhanced by establishing the leading-power hard-collinear factorization formula with next-to-leading-order (NLO) $\alpha_s$ corrections, alongside the systematic inclusion of power-suppressed contributions up to the next-to-next-to-leading power (NNLP). Our numerical analysis demonstrates a subtle cancellation of the factorization-scale dependence at NLO and reveals highly stable Borel plateaus, yielding robust predictions for the couplings. By parameterizing the $\mathcal{O}(1/m_{H^*})$ power corrections, we extract the universal static coupling $\beta = 0.73 \pm 0.13$ and demonstrate that these charge couplings are remarkably insensitive to heavy-quark mass breaking effects. Finally, our investigation into SU(3) flavor symmetry breaking shows that its minute physical effects are currently overshadowed by uncertainties in the non-perturbative vector meson distribution amplitudes.

hep-ph

ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation

Volumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts. To address these challenges, we propose \ourname, a Gaussian splatting based framework for volumetric video reconstruction. Our key insight is that explicitly tracking long term complex motion with individual Gaussian primitives is inherently unstable. Instead, we organize Gaussians around time conditioned anchors that localize their spatial and temporal support, thereby reducing long range motion complexity. We further introduce a temporal windowing strategy to activate only anchors relevant to the queried time, which improves scalability and temporal coherence. In addition, to ensure spatial and temporal stability, we design a compact set of multi level anchor features that encode global features, local spatial features, and local temporal features, jointly constraining Gaussian generation. Extensive experiments demonstrate that \ourname \ consistently outperforms prior methods on long sequence volumetric videos with complex motions. Project page: https://github.com/WuJH2001/ATGS.

cs.CV

HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees

Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes. Existing systems primarily target full-attention models and lack dense, differentiable hybrid-attention execution compatible with activation recomputation. We present HARTS (Hybrid-Attention RL over Tree Structures). HARTS jointly plans microbatches, data-parallel (DP) replica assignments, and microbatch-slot schedules using non-replay compact-token work after prefix compression. For chunkwise linear attention, a linear-time algorithm coordinates chunk-boundary state recovery and replay and produces the minimum number of sequential linear-attention calls under our packed execution model. HARTS preserves the chunkwise state partitioning of trajectory-wise training: it does not repeat projections, MLP/MoE computation, or final outputs, and performs only bounded state replay for numerical alignment. Per round, HARTS batches all branches into one packed call, propagates gradients through differentiable state handoffs, supports activation recomputation, and restores per-token log-probabilities. For deterministic, no-token-drop top-$k$ MoE routing, semantic multiplicities restore MoE-objective token weights and load statistics. Existing RL objectives retain their interface. To our knowledge, HARTS is the first system to demonstrate arbitrary-rollout-tree prefix-sharing speedups on a real hybrid-attention model. On an Agentic RL workload generated from SWE-bench tasks, HARTS achieves $4.81$--$4.87\times$ forward/backward/gradient speedup with activation recomputation across multiple parallel configurations. Its numerical differences are comparable to baseline self-rerun variation, and its reward trend is similar to the baseline over the first 120 steps of $\tau^3$-Bench training.

cs.LG

Adaptive workforce exploration in complex productivity landscapes

Specialization and task allocation enhance efficiency and innovation across diverse systems, from biological organisms to socioeconomic institutions. The evolution of task distribution and its influence on organizational productivity encapsulate the dynamics between task dependencies and adaptive strategies. We explore the organizational division of labor, inspired by the NK model of rugged landscapes, which is widely applied in evolutionary biology, and incorporate interdependencies among the attributes of technical experts within an organization. Our model considers two types of employees characterized by their task allocation strategies: specialists, who are permanently assigned to a single task, and generalists, who stochastically select a task at each time step. We investigate how the ruggedness of the productivity landscape, shaped by task interdependency, affects the organization's capacity to optimize labor division and meet market demands. Using group selection algorithms, we reveal the emergence of nonlinear adaptive dynamics, providing insights into how companies can adapt their strategies to meet market demands and foster innovation.

cond-mat.stat-mech

CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition

Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. However, most existing time series forecasting (TSF) approaches remain inherently passive. Even when incorporating operational decisions as auxiliary covariates, they typically optimize for correlation-based extrapolation under historical policies. This design suffers from autoregressive inertia and conflates endogenous market evolution with decision-induced transitions, leading to policy-insensitive rollouts and unreliable counterfactual analysis. To bridge this gap, we propose CEDAR (Controlled and Event-Driven Demand forecasting via Action-aware Residual decomposition), a two-stage framework for robust decision-conditioned simulation. In Stage I, an Action-Interleaved Transformer learns controllable action-conditioned state transitions for rollout under planned interventions. In Stage II, a Residual Correction Module leverages external event signals and LLM-assisted text representations to align noisy event descriptions with product context and correct event-driven deviations. Our study is enabled by a large-scale real-world dataset from Alibaba 1688, comprising approximately 32 million product trajectories with paired state-action sequences and aligned event signals. Extensive offline experiments and online controlled experiments in production demonstrate that CEDAR consistently improves simulation accuracy over strong TSF baselines and delivers practical gains for real-world budget planning.

cs.LG

PURA: Provably Unbiased and Robust Multi-Bit Watermarking for AI-Generated Text Attribution

Fine-grained attribution of AI-generated text is becoming increasingly important for accountability and auditing, yet existing multi-bit watermarking methods still struggle to simultaneously preserve the base generation distribution, support high-capacity payloads, and remain recoverable after editing. We present PURA, a provably unbiased and robust multi-bit watermarking method for text attribution. Instead of perturbing token probabilities directly, PURA embeds payloads in the latent sampling space via keyed inverse transform sampling, and recovers them by treating observed tokens as soft interval evidence and aggregating such evidence across the sequence. This design preserves the base generation distribution exactly while substantially improving recovery stability under post-editing and channel perturbations. Building on this recovery paradigm, we further develop a unified robustness analysis and show that, under bounded attack strength, the per-bit error probability decays exponentially with sequence length. Extensive experiments show that PURA substantially outperforms existing unbiased baselines in the high-payload regime. For example, when embedding 36 bits in 200 tokens, PURA achieves a 91.7\% message match rate, more than three times that of the strongest unbiased baseline, while preserving text quality and remaining statistically close to unwatermarked text, and incurring only millisecond-level verification overhead. Our code is available at

cs.CR

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.

cs.SI

Local Well-Posedness for Compressible Capillary-Gravity Water Waves with Acute Contact Angles

Our purpose is to investigate the local well-posedness of the compressible Euler equations in a two-dimensional bounded corner domain with acute contact angles. This configuration describes a free surface intersecting the fixed bottom at two points, where the fluid is subject to a gravitational field and the interface between the fluid and air is influenced by capillary forces. When the contact angles are less than $\pi/2$, we establish a local existence theory for the solution, with dissipation effects occurring at the contact points. The main analytical challenge arises from contact point singularities, which renders previous methods for dealing with compressible free boundary problems inadequate. To overcome this, we first establish the geometric structure for the compressible Euler equations, an approach originally introduced by Shatah and Zeng \cite{Shatah2008} for incompressible fluids. Additionally, we provide a singularity analysis for the wave equations in the corner domain, which ensures the validity of calculations near the corner. Finally, based on the geometric structure and singularity analysis, we obtain a priori energy estimates. Using these estimates, we also prove the local well-posedness of the system in a geometric formulation. To our knowledge, this is the first result addressing compressible Euler equations with a free boundary that involves contact points.

math.AP

Adaptive Unequal Error Protection for Semantic Split Learning over Wireless Channels

We propose a task-aware semantic split learning (SL) framework for wireless edge-cloud inference, in which the reliability of transmitted latent representations is dynamically adapted to their relevance for the downstream task. An autoencoder (AE)-based physical (PHY) layer enables end-to-end learning of the communication interface, while unequal error protection (UEP) is realized via mutual information (MI)-driven prioritization of latent components during training. The gradient of the estimated MI with respect to each latent component serves as a sensitivity-based proxy for task relevance, providing a fully learning-driven prioritization that adapts to both the data distribution and the downstream task. We further show that this prioritization translates into measurable physical-layer effects: MI-guided UEP assigns significantly higher transmit power to the most task-critical latent components compared to the equal error protection (EEP) baseline. Experiments on real-world IoT sensing data demonstrate consistent gains over equal and fixed-UEP baselines across SNR regimes. Additional analysis confirms ranking stability, estimator robustness and generalization across datasets and task types, indicating broad applicability of the proposed framework.

cs.IT

Long-Wave Spectral Instability of Shear Layers for the Compressible Euler Equations

We study the long-wave spectral instability of the two-dimensional compressible Euler equations around smooth monotone shear layers. We construct decaying half-line solutions of the compressible Rayleigh equation through a long-wave expansion and derive a second-order expansion of the matching Wronskian. For every fixed Mach number $m>0$, we prove the existence of unstable modes for sufficiently small wavenumbers. For $m<\sqrt2$, this holds for a class of profiles, with $c_i$ tending to a positive constant as $\alpha\to0$. At $m=\sqrt{2}$, the profile $U_s(Y)=\tanh Y$ admits an unstable mode with $c\to0$ and $c_i$ of order $\alpha^{1/3}$. For $m>\sqrt2$, the same profile remains unstable, with $c\to c_*(m)\in(0,1)$ and $c_i>0$ of order $\alpha$. The corresponding temporal growth rates are of order $\alpha$, $\alpha^{4/3}$ and $\alpha^2$, respectively, showing a change in the long-wave instability scaling at $m=\sqrt2$. In the zero-thickness limit, the supercritical unstable eigenvalue approaches the real axis, consistently with the stability results for supersonic compressible vortex sheets in \cite{CS1,CS2}.

math.AP

PPOM: Marginalizing Patch-Grid Phase for CLIP-Based Generalizable Vision-Language Prompt Tuning

Prompt tuning adapts CLIP-based vision-language models with few trainable parameters, yet its predictions remain sensitive to the spatial sampling imposed by a frozen vision transformer. In particular, non-overlapping patch tokenization makes predictions depend on the alignment (phase) between image and the patch lattice. To reduce prediction sensitivity to patch-grid alignment, we introduce Patch-Phase Orbit Marginalization (PPOM), a training-free inference operator that treats phase shift as a nuisance variable. Given a patch stride, PPOM evaluates the identity view and reflection-padded translations, pairs opposite shifts into horizontal, vertical, and diagonal antithetic families, and assigns equal mass to these families and the identity prediction to avoid view-count bias during phase integration. In summary, PPOM provides a deterministic interface between prompt adaptation and patch-grid sensitivity. Across multiple prompt-learning hosts, PPOM improves host performance without re-training.

cs.CV

The $D^*D^*\pi$ and $B^*B^*\pi$ couplings from light-cone sum rules

We revisit the calculation of the strong couplings $D^{*}D^{*}\pi$ and $B^{*}B^{*}\pi$ from the light-cone sum rules (LCSR) using the pion light-cone distribution amplitudes. The accuracy of the underlying correlation function is upgraded by establishing the hard-collinear factorization formula at the leading power up to the next-to-leading order in $\alpha_s$. Furthermore, the next-to-leading power contributions are systematically incorporated at the leading order by evaluating the two-particle and three-particle higher-twist pion distribution amplitudes up to twist-4 accuracy. By matching the QCD-level spectral representations with the hadronic dispersion relations, we present a solid numerical analysis that accounts for the finite heavy quark masses and carefully evaluates the systematic uncertainty originating from the two-dimensional quark-hadron duality ansatz. We predict $g_{D^{*}D^{*}\pi} = 5.71_{-0.53}^{+0.67} \text{ GeV}^{-1}$ and $g_{B^{*}B^{*}\pi} = 4.90_{-0.44}^{+0.57} \text{ GeV}^{-1}$. Finally, by parameterizing the $1/m_H$ (with $H=D,\,B$) power corrections to extract the universal static coupling $\hat{g} = 0.30 \pm 0.04$, we compare our results with previous theoretical determinations and experimental data, highlighting the significance of heavy quark spin symmetry breaking effects.

hep-ph

Secure Cooperative THz ISAC via Mamba Empowered Graph Neural Network Precoding

The terahertz (THz) band offers abundant spectrum resources for high-throughput communication and ultra high-precision localization. This paper investigates secure communication in cooperative THz orthogonal frequency-division multiplexing (OFDM) bistatic integrated sensing and communications (ISAC) systems, where multiple base stations (BSs) equipped with extremely large-scale antenna arrays (ELAAs) collaboratively serve downlink users while concurrently locating multiple targets. Malicious targets are assumed to act as potential eavesdroppers attempting to intercept confidential information intended for legitimate users. To mitigate these threats, we formulate a joint optimization problem for analog beamforming, digital precoding, true-time delayers (TTDs), and sensing signal covariance matrix design. The objective is to maximize the minimum secrecy rate subject to Cramer-Rao bound (CRB) constraints that ensure localization accuracy. This problem is highly challenging due to the non-convex CRB constraint, strongly coupled variables, high computational complexity from ELAA, and near-field channel modeling. To address these challenges, we propose a novel data-driven framework that integrates graph neural networks (GNNs) with the Mamba architecture. Our proposed framework first encodes the interactions among users, targets, and BSs into a heterogeneous graph and then employs message passing to optimize vertex features. The Mamba blocks further enhance this process through their selection mechanism and state space modeling capabilities, enabling dynamic and context-aware optimization of beamforming, TTD configurations, and sensing parameters. Numerical simulations validate that the proposed method outperforms both conventional and learning-based baselines, while offering high computational efficiency and strong generalization across different network conditions.

cs.IT

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge

Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata, so index traffic and nonzero extraction become critical SpMM bottlenecks. We introduce the Payload-to-Metadata Ratio (PMR) and show that improving PMR raises effective compute intensity in decoding. We present UnionSparse, an index-efficient framework that combines Index-Efficient Bitmap Encoding (IE-BME) with a SpMM kernel using Low-Bit Shared-Memory Parallel Decoding (LSPD). IE-BME amortizes metadata and aligns sparse traversal with fragment assembly, while LSPD improves small-batch execution. Under W4A4 quantization and 30%--70% sparsity, UnionSparse outperforms FlashLLM and SpInfer by 2.30x and 1.43x, and CUTLASS and cuBLAS Tensor Core by 1.56x and 3.46x, respectively. These results establish payload-extraction efficiency as a first-order concern for low-bit sparse inference on edge GPUs. Source code is available at: https://github.com/Victor-Alen/UnionSparse.

cs.DC

Intelligent Wiretap Code Design: Exploiting Wireless Endogenous Security via Information Theory and Deep Learning Integration

Recent advancements in wireless endogenous security have explored leveraging the inherent randomness of wireless channels to enhance communication security, providing an effective alternative to traditional encryption methods. This paper proposes a wiretap coding scheme within the semantic communication framework, which leverages discrete semantic representations compatible with conventional digital modulation to jointly enhance communication security and reliability. We investigate two eavesdropping scenarios: (i) the eavesdropper employs a maximum a posteriori (MAP) decoder, and (ii) the eavesdropper has access to a decoder identical to that of the legitimate receiver. In the first scenario, we exploit mutual information as a metric to guide the design of an optimized coding strategy, minimizing information leakage while enhancing communication reliability. In the second scenario, considering the limitations of the eavesdropper's decoding capability, we employ generalized mutual information (GMI) to characterize recoverability under the prescribed decoding rule and guide reliability-aware code optimization.

cs.IT

CSI Reconstruction in Fluid Antenna Systems Without Spatial Covariance Priors

Fluid antenna systems (FASs) exploit many candidate ports for spatial diversity, but hardware constraints allow channel observations at only a few active ports. Whether full-port CSI can be recovered without pre-acquired channel statistics remains open. Under the Clarke isotropic scattering model, we show that the channel lies in a low-dimensional spatial modal subspace determined by the scattering environment rather than the total port count. Consequently, recovery becomes feasible when the number of observed ports reaches the modal dimension (i.e., $M\geq r$), even when $M\ll N$. We further establish a sharp feasibility threshold: reliable recovery is impossible below this dimension regardless of SNR, whereas accuracy improves with additional observations above it. By decomposing the recovery error into modal truncation, estimation, and learning components, we derive explicit tradeoffs among RF chains, pilot overhead, transmit power, and training data. These results enable scalable prior-free full-port CSI recovery with few active ports.

cs.IT