SearcharxivSearch

arXiv subjects

Zhang Chen

Publications and source records attributed to Zhang Chen.

At least 19 recordsLinked to original sources

Uniform Large Deviations of Mckean-Vlasov Stochastic Fractional $(\alpha,p)$-Laplacian Equations Driven by Superlinear Noise on $\mathbb{R}^d$

The global-in-time well-posedness and uniform large deviation principles (LDPs) are investigated for a wide class of Mckean-Vlasov stochastic non-local fractional $(\alpha,p)$-Laplacian equations with $\alpha \in (0,1)$ and $p>2$ driven by superlinear multiplicative noise defined on the whole space $\mathbb{R}^d$, where the non-local nonlinear fractional $(\alpha,p)$-Laplace operator is defined by a singular, symmetrical and translation invariant kernel function, the distribution-dependent drift terms have arbitrary polynomial growth and the distribution-dependent diffusion terms have superlinear growth. The global-in-time well-posedness is established under these conditions by using the monotone method and a domain expansion argument. Under additional conditions on the growth of diffusion terms, we establish the Freidlin-Wentzell and Dembo-Zeitouni uniform LDPs by using the generalized weak convergence method developed by Salins (Probab. Surv., 16:99-142, 2019). The idea of uniform tail-ends estimates, the pseudo monotone technique and the Arzel\`{a}-Ascoli theorem are combined to prove the weak-to-strong continuity of solution operators of the controlled equations in order to overcome many difficulties caused by the noncompactness of Sobolev embeddings on $\mathbb{R}^d$ and the nonlinearity of the fractional $(\alpha,p)$-Laplace operator. The superlinearly growing diffusion term is carefully controlled by using the dissipative drift terms and several algebraic inequalities.

math.PR

SPARE-GS: Structural Parsimony and Resource Efficiency for 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis in real-time; however its training efficiency and representation compactness are hindered by excessive primitive proliferation. To address this challenge, we formulate the structural evolution of 3DGS as a global budget-constrained optimization problem and derive an optimality condition, which requires the marginal utility of structural resources to be balanced across spatial regions under a finite primitive budget. Based on this formulation, we propose SPARE-GS, a general plug-and-play framework that dynamically aligns the distribution of 3D Gaussian primitives with regional representational demand. SPARE-GS estimates capacity-normalized regional demand, assigns adaptive target quotas, and uses regional budget deviations to coordinate densification, pruning and adaptive termination toward a more balanced structural allocation. Extensive experiments across standard, accelerated, and structure-enhanced 3DGS pipelines demonstrate that SPARE-GS reduces the Gaussian count and training time by an average of 30.38% and 23.81%, respectively, while improving the average PSNR. Moreover, the resulting compact representations reduce downstream processing time and improve the rate-distortion performance of diverse compression and pruning methods, demonstrating the broad applicability of global structural budget regulation.

cs.CV

Existence and vanishing noise limit of measure attractors for McKean-Vlasov $p$-Laplacian lattice systems with delay driven by L\'evy noise

This paper is concerned with the existence and the limiting behavior of measure attractors of distribution laws of the solution segment process for the McKean-Vlasov stochastic $p$-Laplace lattice system with time delay driven by L\'evy noise. The nonlinear drift and diffusion terms are allowed to have superlinear growth. Due to time delay, the Skorohod metric space is employed to describe the trajectories of the solutions with jumps. We first prove the existence and uniqueness of c\`adl\`ag solutions for the lattice system, and then define a non-autonomous cocycle acting on the Borel probability measures in the Skorohod space. This cocycle is continuous in bounded subsets of the space of probability measures only when time is sufficiently large. We then prove the existence of pullback absorbing sets and the asymptotic compactness of the cocycle as well as the existence and uniqueness of pullback measure attractors. We finally investigate the limiting behavior of measure attractors of the lattice system as the noise intensity approaches zero, and establish the optimal convergence rate of singleton measure attractors in the Wasserstein distance of order $\theta$.

math.PR

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.

cs.RO

REFINE: Super-efficient 3D Gaussian Splatting Pruning via Rendering-Free Primitive Importance

Existing pruning methods for 3D Gaussian splatting (3DGS) suffer from either severe quality degradation or prohibitive computational overhead. In this paper, we propose REFINE, a highly accelerated 3DGS pruning framework centered on a novel rendering-free primitive importance metric. Our approach leverages an analytically approximated, rendering-aware Hessian field to quantify the expected perceptual error induced by the removal of individual primitives. By modeling the joint modulation of visibility, projection geometry and the content adaptive hyperparameter, we entirely bypass costly forward rendering passes and derive an anisotropic perceptual weight field that serves as a high-fidelity proxy for primitive importance. Extensive experiments across multiple benchmark datasets demonstrate that REFINE maintains highly competitive rendering quality while achieving a $3,000\times$ reduction in pruning-related computational complexity, translating to a practical $\sim 20\times$ speedup in device latency compared to state-of-the-art pruning methods.

cs.CV

Historical Developments in Probability Measures for Asset Pricing: From State Prices to Modern Pricing Kernels

This review summarizes the historical development of probability measures in asset pricing, from early mathematical finance and state price theory to risk-neutral valuation, martingale measures, forward measures, stochastic discount factors, incomplete-market measure selection, benchmark pricing, robust and nonlinear pricing, and modern data-driven probability transformations. The central theme is that asset pricing is not merely an exercise in estimating physical probabilities. Instead, pricing theory constructs, transforms, or selects probability measures so that market prices can be represented as expectations after discounting, numeraire normalization, marginal utility weighting, entropy penalization, calibration, or information conditioning. The paper emphasizes landmark contributions including Bachelier's probabilistic model of speculation, Arrow-Debreu state-contingent claims, Black-Scholes-Merton option pricing, Harrison-Kreps and Harrison-Pliska's martingale formalization, Delbaen and Schachermayer's fundamental theorem, Breeden-Litzenberger implied state price densities, change of numeraire methods, Hansen-Jagannathan stochastic discount factor restrictions, Cochrane's SDF synthesis, and recent empirical and machine learning work on learned pricing kernels. Text-, attention-, and sentiment-based probability transformations are treated as recent information-adjusted forecasting extensions that complement, rather than replace, martingale, numeraire, SDF, and incomplete-market frameworks. The paper also collects key formulas for state prices, stochastic discount factors, Radon-Nikodym densities, Girsanov changes of measure, risk-neutral valuation, forward measures, implied densities, coherent risk measures, benchmark pricing, learned SDFs, and information-adjusted forecasting.

q-fin.MF

Autoregressive Appearance Prediction for 3D Gaussian Avatars

A photorealistic and immersive human avatar experience demands capturing fine, person-specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large, high-quality datasets, which often introduce ambiguities and spurious correlations when very similar poses correspond to different appearances. Models that fit these details during training can overfit and produce unstable, abrupt appearance changes for novel poses. We propose a 3D Gaussian Splatting avatar model with a spatial MLP backbone that is conditioned on both pose and an appearance latent. The latent is learned during training by an encoder, yielding a compact representation that improves reconstruction quality and helps disambiguate pose-driven renderings. At driving time, our predictor autoregressively infers the latent, producing temporally smooth appearance evolution and improved stability. Overall, our method delivers a robust and practical path to high-fidelity, stable avatar driving.

cs.CV

SAM Molecular Stacking with Heterogeneous Orientationfor High-Performance Perovskite Photovoltaics

This study demonstrates that thermal-evaporated SAM (eSAM) films, particularly in a thick configuration, spontaneously adopt a heterogeneous molecular orientation, forming a vertical-to-horizontal gradient in molecular packing. This unique architecture establishes a graded energy barrier, which is shown to facilitate more efficient hole transport compared with the single energy barrier presented by conventional thin SAMs. In conclusion, while solution-processed SAMs present formidable scalability challenges, the thermal evaporation of SAMs offers a viable pathway toward industrial-scale fabrication. The strategy of employing thick eSAM films with gradient molecular packing not only circumvents the uniformity issues of solution methods but also introduces a superior structure for charge transport, positioning it as a promising enabler for the commercialization of high-efficiency perovskite photovoltaics. The inability to achieve uniform hole transport with solution-processed self-assembled monolayers (SAMs) constitutes a fundamental bottleneck for scaling perovskite photovoltaics. Herein, we demonstrate that thermal-evaporated SAMs (eSAMs) overcome this limitation by enabling precise thickness control. Crucially, a thickened eSAM spontaneously forms a vertical-to-horizontal gradient in molecular orientation, which creates a descending energy barrier that directionally facilitates hole transport. This tailored interface also ensures excellent surface coverage and directs the growth of high-quality perovskite films. Consequently, the resultant photovoltaic devices set new benchmarks, delivering impressive power conversion efficiencies (PCEs) of 21.46% (small-area, 0.108 cm2) and 19.38% (large-area module, 15.52 cm2) for fully vacuum-evaporated devices, while also setting an impressive PCE of 23.67% for eSAM-based devices with solution-processed perovskites.

cond-mat.mtrl-sci

Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals

We introduce a multimodal industrial fault analysis dataset collected from a single-speed chain conveyor (SSCC) system, targeting system-level fault detection in production lines. The dataset consists of multimodal signals, including three audio and four vibration channels. It covers normal operation and four representative fault types under multiple speeds, loads, and both clean and realistic factory-noise conditions reproduced on-site. It is explicitly designed to support channel-wise analysis and multimodal fusion research. We establish standardized evaluation protocols for unsupervised fault detection with normal-only training and supervised fault classification with balanced dataset splits across different operating conditions and fault types. A unified channel-wise kNN baseline is provided to enable fair comparison of representation quality without task-specific training. The dataset offers a practical and extensible benchmark for robust multimodal industrial fault analysis.

cs.SD

SARAH: Spatially Aware Real-time Agentic Humans

As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech-aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current methods lack this spatial awareness. We close this gap with the first real-time, fully causal method for spatially-aware conversational motion, deployable on a streaming VR headset. Given a user's position and dyadic audio, our approach produces full-body motion that aligns gestures with speech while orienting the agent according to the user. Our architecture combines a causal transformer-based VAE with interleaved latent tokens for streaming inference and a flow matching model conditioned on user trajectory and audio. To support varying gaze preferences, we introduce a gaze scoring mechanism with classifier-free guidance to decouple learning from control: the model captures natural spatial alignment from data, while users can adjust eye contact intensity at inference time. On the Embody 3D dataset, our method achieves state-of-the-art motion quality at over 300 FPS -- 3x faster than non-causal baselines -- while capturing the subtle spatial dynamics of natural conversation. We validate our approach on a live VR system, bringing spatially-aware conversational agents to real-time deployment. Please see https://evonneng.github.io/sarah/ for details.

cs.CV

Objective Quality Assessment of Point Clouds Using Multi-scale Implicit Structural Similarity

The unstructured and irregular nature of points poses a significant challenge for accurate point cloud quality assessment (PCQA), particularly in establishing accurate perceptual feature correspondence. To tackle this, we propose the Multi-scale Implicit Structural Similarity Measurement (MS-ISSM). Unlike traditional point-to-point matching, MS-ISSM utilizes radial basis function (RBF) to represent local features continuously, transforming distortion measurement into a comparison of implicit function coefficients. This approach effectively circumvents matching errors inherent in irregular data. Additionally, we propose a ResGrouped-MLP quality assessment network, which robustly maps multi-scale feature differences to perceptual scores. The network architecture departs from traditional flat multi-layer perceptron (MLP) by adopting a grouped encoding strategy integrated with residual blocks and channel-wise attention mechanisms. This hierarchical design allows the model to preserve the distinct physical semantics of luma, chroma, and geometry while adaptively focusing on the most salient distortion features across High, Medium, and Low scales. Experimental results on multiple benchmarks demonstrate that MS-ISSM outperforms state-of-the-art metrics in both reliability and generalization. The source code is available at: https://github.com/ZhangChen2022/MS-ISSM.

cs.CV

DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

We present DualMat, a novel dual-path diffusion framework for estimating Physically Based Rendering (PBR) materials from single images under complex lighting conditions. Our approach operates in two distinct latent spaces: an albedo-optimized path leveraging pretrained visual knowledge through RGB latent space, and a material-specialized path operating in a compact latent space designed for precise metallic and roughness estimation. To ensure coherent predictions between the albedo-optimized and material-specialized paths, we introduce feature distillation during training. We employ rectified flow to enhance efficiency by reducing inference steps while maintaining quality. Our framework extends to high-resolution and multi-view inputs through patch-based estimation and cross-view attention, enabling seamless integration into image-to-3D pipelines. DualMat achieves state-of-the-art performance on both Objaverse and real-world data, significantly outperforming existing methods with up to 28% improvement in albedo estimation and 39% reduction in metallic-roughness prediction errors.

cs.CV

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence to the physical world, fueling the flourishing of vision-language-action (VLA) models. Despite seemingly diverse approaches, we observe that current VLA models can be unified under a single framework: vision and language inputs are processed by a series of VLA modules, producing a chain of \textit{action tokens} that progressively encode more grounded and actionable information, ultimately generating executable actions. We further determine that the primary design choice distinguishing VLA models lies in how action tokens are formulated, which can be categorized into language description, code, affordance, trajectory, goal state, latent representation, raw action, and reasoning. However, there remains a lack of comprehensive understanding regarding action tokens, significantly impeding effective VLA development and obscuring future directions. Therefore, this survey aims to categorize and interpret existing VLA research through the lens of action tokenization, distill the strengths and limitations of each token type, and identify areas for improvement. Through this systematic review and analysis, we offer a synthesized outlook on the broader evolution of VLA models, highlight underexplored yet promising directions, and contribute guidance for future research, hoping to bring the field closer to general-purpose intelligence.

cs.RO

Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can both comprehend and generate dyadic behavioral dynamics. To this end, we introduce the Seamless Interaction Dataset, a large-scale collection of over 4,000 hours of face-to-face interaction footage from over 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand dyadic embodied dynamics, unlocking breakthroughs in virtual agents, telepresence experiences, and multimodal content analysis tools. We also develop a suite of models that utilize the dataset to generate dyadic motion gestures and facial expressions aligned with human speech. These models can take as input both the speech and visual behavior of their interlocutors. We present a variant with speech from an LLM model and integrations with 2D and 3D rendering methods, bringing us closer to interactive virtual agents. Additionally, we describe controllable variants of our motion models that can adapt emotional responses and expressivity levels, as well as generating more semantically-relevant gestures. Finally, we discuss methods for assessing the quality of these dyadic motion models, which are demonstrating the potential for more intuitive and responsive human-AI interactions.

cs.CV

Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges

Large language models (LLMs) exhibit probabilistic output characteristics, yet conventional evaluation frameworks rely on deterministic scalar metrics. This study introduces a Bayesian approach for LLM capability assessment that integrates prior knowledge through probabilistic inference, addressing limitations under limited-sample regimes. By treating model capabilities as latent variables and leveraging a curated query set to induce discriminative responses, we formalize model ranking as a Bayesian hypothesis testing problem over mutually exclusive capability intervals. Experimental evaluations with GPT-series models demonstrate that the proposed method achieves superior discrimination compared to conventional evaluation methods. Results indicate that even with reduced sample sizes, the approach maintains statistical robustness while providing actionable insights, such as probabilistic statements about a model's likelihood of surpassing specific baselines. This work advances LLM evaluation methodologies by bridging Bayesian inference with practical constraints in real-world deployment scenarios.

cs.CL

Steady-State Drifting Equilibrium Analysis of Single-Track Two-Wheeled Robots for Controller Design

Drifting is an advanced driving technique where the wheeled robot's tire-ground interaction breaks the common non-holonomic pure rolling constraint. This allows high-maneuverability tasks like quick cornering, and steady-state drifting control enhances motion stability under lateral slip conditions. While drifting has been successfully achieved in four-wheeled robot systems, its application to single-track two-wheeled (STTW) robots, such as unmanned motorcycles or bicycles, has not been thoroughly studied. To bridge this gap, this paper extends the drifting equilibrium theory to STTW robots and reveals the mechanism behind the steady-state drifting maneuver. Notably, the counter-steering drifting technique used by skilled motorcyclists is explained through this theory. In addition, an analytical algorithm based on intrinsic geometry and kinematics relationships is proposed, reducing the computation time by four orders of magnitude while maintaining less than 6% error compared to numerical methods. Based on equilibrium analysis, a model predictive controller (MPC) is designed to achieve steady-state drifting and equilibrium points transition, with its effectiveness and robustness validated through simulations.

cs.RO

PanoDreamer: Consistent Text to 360-Degree Scene Generation

Automatically generating a complete 3D scene from a text description, a reference image, or both has significant applications in fields like virtual reality and gaming. However, current methods often generate low-quality textures and inconsistent 3D structures. This is especially true when extrapolating significantly beyond the field of view of the reference image. To address these challenges, we propose PanoDreamer, a novel framework for consistent, 3D scene generation with flexible text and image control. Our approach employs a large language model and a warp-refine pipeline, first generating an initial set of images and then compositing them into a 360-degree panorama. This panorama is then lifted into 3D to form an initial point cloud. We then use several approaches to generate additional images, from different viewpoints, that are consistent with the initial point cloud and expand/refine the initial point cloud. Given the resulting set of images, we utilize 3D Gaussian Splatting to create the final 3D scene, which can then be rendered from different viewpoints. Experiments demonstrate the effectiveness of PanoDreamer in generating high-quality, geometrically consistent 3D scenes.

cs.CV

RBFIM: Perceptual Quality Assessment for Compressed Point Clouds Using Radial Basis Function Interpolation

One of the main challenges in point cloud compression (PCC) is how to evaluate the perceived distortion so that the codec can be optimized for perceptual quality. Current standard practices in PCC highlight a primary issue: while single-feature metrics are widely used to assess compression distortion, the classic method of searching point-to-point nearest neighbors frequently fails to adequately build precise correspondences between point clouds, resulting in an ineffective capture of human perceptual features. To overcome the related limitations, we propose a novel assessment method called RBFIM, utilizing radial basis function (RBF) interpolation to convert discrete point features into a continuous feature function for the distorted point cloud. By substituting the geometry coordinates of the original point cloud into the feature function, we obtain the bijective sets of point features. This enables an establishment of precise corresponding features between distorted and original point clouds and significantly improves the accuracy of quality assessments. Moreover, this method avoids the complexity caused by bidirectional searches. Extensive experiments on multiple subjective quality datasets of compressed point clouds demonstrate that our RBFIM excels in addressing human perception tasks, thereby providing robust support for PCC optimization efforts.

cs.CV