SearcharxivSearch

arXiv subjects

Hongwen Zhang

Publications and source records attributed to Hongwen Zhang.

At least 19 recordsLinked to original sources

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.

cs.CV

Giant and Continuous Ionic Current Oscillation Induced by Dynamic Surface Charge Regulation in Cylindrical Mesopores

Nanofluidic ionic oscillators based on the dynamic regulation of surface charges hold great promise for neuromorphic computing, biosensing, and ionic circuits. Here, by dynamically adjusting the local charge inversion on pore walls, we present a simple and effective strategy to achieve periodic current oscillations by harnessing the transient adsorption and desorption of Ca2+ ions in cylindrical mesopores under concentration gradients. Based on the combined precision current measurements and multiphysics simulations, we demonstrate that local overadsorption of Ca2+ ions may induce asymmetric bipolar charge distributions along the pore axis, which periodically reverses the direction of electroosmotic flow and modulates the local ion concentration inside the pore, generating highly regular current oscillations. Notably, both the oscillation frequency and the open-state probability of the pore vary nearly linearly with the applied voltage. Moreover, under dynamic voltage scanning, the system exhibits typical memristive hysteresis, and the switching between the open and closed states is highly reproducible. This work not only reveals the dynamic, heterogeneous surface charge regulation by divalent ions, but also provides a simple, material agnostic method for constructing ionic oscillators and memristors based on dynamic adsorption/desorption of multivalent ions.

physics.chem-ph

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action (V-A) paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action (V-3D-A) paradigm, leveraging intermediate 3D representations. However, these methods often struggle with action inaccuracies due to the complexity and dynamic nature of manipulation scenes. In this paper, we adopt a V-4D-A framework that enables direct action reasoning from motion-aware 4D representations via a Gaussian Action Field (GAF). GAF extends 3D Gaussian Splatting (3DGS) by incorporating learnable motion attributes, allowing 4D modeling of dynamic scenes and manipulation actions. To learn time-varying scene geometry and action-aware robot motion, GAF provides three interrelated outputs: reconstruction of the current scene, prediction of future frames, and estimation of init action via Gaussian motion. Furthermore, we employ an action-vision-aligned denoising framework, conditioned on a unified representation that combines the init action and the Gaussian perception, both generated by the GAF, to further obtain more precise actions. Extensive experiments demonstrate significant improvements, with GAF achieving +11.5385 dB PSNR, +0.3864 SSIM and -0.5574 LPIPS improvements in reconstruction quality, while boosting the average +7.3% success rate in robotic manipulation tasks over state-of-the-art methods.

cs.RO

Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation

Recent trends in sparse-view 3D reconstruction have taken two different paths: feed-forward reconstruction that predicts pixel-aligned point maps without a complete geometry, and generative 3D reconstruction that generates complete geometry but often with poor input-alignment. We present Mix3R, a novel generative 3D reconstruction method which mixes feed-forward reconstruction and 3D generation into a single framework in an aligned manner. Mix3R generates a 3D shape in two stages: a sparse voxel generation stage and a textured geometry generation stage. Unlike pure generative methods, our first-stage generation jointly produces a coarse 3D structure (sparse voxels), per-view point maps and camera parameters aligned to that 3D structure. This is made possible by introducing a Mixture-of-Transformers architecture that inserts global self-attentions to a feed-forward reconstruction model and a 3D generative model, both pretrained on large-scale data. This design effectively retains the pretrained priors but enables better 2D-3D alignment. Based on the initial aligned generations of sparse 3D voxels and point maps, we compute an overlap-based attention bias that is directly added to another pretrained textured geometry generation model, enabling it to correctly place input textures onto generated shapes in a training-free manner. Our design brings mutual benefits to both feed-forward reconstruction and 3D generation: The feed-forward branch learns to ground its predictions to a generative 3D prior, and conversely, the 3D generation branch is conditioned on geometrically informative features from the feed-forward branch. As a result, our method produces 3D shapes with better input alignment compared with pure 3D generative methods, together with camera pose estimations more accurate than previous feed-forward reconstruction methods. Our project page is at https://jsnln.github.io/mix3r/

cs.CV

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexities in multi-agent coordination, and limited generalization across objects. In this paper, we present SynAgent, a unified framework that enables scalable and physically plausible cooperative manipulation by leveraging Solo-to-Cooperative Agent Synergy to transfer skills from single-agent human-object interaction to multi-agent human-object-human scenarios. To maintain semantic integrity during motion transfer, we introduce an interaction-preserving retargeting method based on an Interact Mesh constructed via Delaunay tetrahedralization, which faithfully maintains spatial relationships among humans and objects. Building upon this refined data, we propose a single-agent pretraining and adaptation paradigm that distills synergistic collaborative behaviors from abundant single-human data through decentralized training and multi-agent PPO. Finally, we develop a trajectory-conditioned generative policy using a conditional VAE, trained via multi-teacher distillation from motion imitation priors to achieve stable and controllable object-level trajectory execution. Extensive experiments demonstrate that SynAgent significantly outperforms existing baselines in both cooperative imitation and trajectory-conditioned control, while generalizing across diverse object geometries. Codes and data will be available after publication. Project Page: https://yw0208.github.io/synagent/

cs.CV

OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer

In this paper, we introduce OmniHands, a universal approach to recovering interactive hand meshes and their relative movement from monocular or multi-view inputs. Our approach addresses two major limitations of previous methods: lacking a unified solution for handling various hand image inputs and neglecting the positional relationship of two hands within images. To overcome these challenges, we develop a universal architecture with novel tokenization and contextual feature fusion strategies, capable of adapting to a variety of tasks. Specifically, we propose a Relation-aware Two-Hand Tokenization (RAT) method to embed positional relation information into the hand tokens. In this way, our network can handle both single-hand and two-hand inputs and explicitly leverage relative hand positions, facilitating the reconstruction of intricate hand interactions in real-world scenarios. As such tokenization indicates the relative relationship of two hands, it also supports more effective feature fusion. To this end, we further develop a 4D Interaction Reasoning (FIR) module to fuse hand tokens in 4D with attention and decode them into 3D hand meshes and relative temporal movements. The efficacy of our approach is validated on several benchmark datasets. The results on in-the-wild videos and real-world scenarios demonstrate the superior performances of our approach for interactive hand reconstruction. More video results can be found on the project page: https://OmniHand.github.io.

cs.CV

SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis

Hand-Object Interaction (HOI) generation plays a critical role in advancing applications across animation and robotics. Current video-based methods are predominantly single-view, which impedes comprehensive 3D geometry perception and often results in geometric distortions or unrealistic motion patterns. While 3D HOI approaches can generate dynamically plausible motions, their dependence on high-quality 3D data captured in controlled laboratory settings severely limits their generalization to real-world scenarios. To overcome these limitations, we introduce SyncMV4D, the first model that jointly generates synchronized multi-view HOI videos and 4D motions by unifying visual prior, motion dynamics, and multi-view geometry. Our framework features two core innovations: (1) a Multi-view Joint Diffusion (MJD) model that co-generates HOI videos and intermediate motions, and (2) a Diffusion Points Aligner (DPA) that refines the coarse intermediate motion into globally aligned 4D metric point tracks. To tightly couple 2D appearance with 4D dynamics, we establish a closed-loop, mutually enhancing cycle. During the diffusion denoising process, the generated video conditions the refinement of the 4D motion, while the aligned 4D point tracks are reprojected to guide next-step joint generation. Experimentally, our method demonstrates superior performance to state-of-the-art alternatives in visual realism, motion plausibility, and multi-view consistency.

cs.CV

SharpTimeGS: Sharp and Stable Dynamic Gaussian Splatting via Lifespan Modulation

Novel view synthesis of dynamic scenes is fundamental to achieving photorealistic 4D reconstruction and immersive visual experiences. Recent progress in Gaussian-based representations has significantly improved real-time rendering quality, yet existing methods still struggle to maintain a balance between long-term static and short-term dynamic regions in both representation and optimization. To address this, we present SharpTimeGS, a lifespan-aware 4D Gaussian framework that achieves temporally adaptive modeling of both static and dynamic regions under a unified representation. Specifically, we introduce a learnable lifespan parameter that reformulates temporal visibility from a Gaussian-shaped decay into a flat-top profile, allowing primitives to remain consistently active over their intended duration and avoiding redundant densification. In addition, the learned lifespan modulates each primitives' motion, reducing drift in long-lived static points while retaining unrestricted motion for short-lived dynamic ones. This effectively decouples motion magnitude from temporal duration, improving long-term stability without compromising dynamic fidelity. Moreover, we design a lifespan-velocity-aware densification strategy that mitigates optimization imbalance between static and dynamic regions by allocating more capacity to regions with pronounced motion while keeping static areas compact and stable. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art performance while supporting real-time rendering up to 4K resolution at 100 FPS on one RTX 4090.

cs.CV

CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template Decomposition

Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing details or hinder end-to-end learning. In this paper, we revisit point-based solutions and propose to decompose explicit garment-related templates and then add pose-dependent wrinkles to them. In this way, the clothing deformations are disentangled such that the pose-dependent wrinkles can be better learned and applied to unseen poses. Additionally, to tackle the seam artifact issues in recent state-of-the-art point-based methods, we propose to learn point features on a body surface, which establishes a continuous and compact feature space to capture the fine-grained and pose-dependent clothing geometry. To facilitate the research in this field, we also introduce a high-quality scan dataset of humans in real-world clothing. Our approach is validated on two existing datasets and our newly introduced dataset, showing better clothing deformation results in unseen poses. The project page with code and dataset can be found at https://zhanghongwen.cn/closet.

cs.CV

PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images

We present PyMAF-X, a regression-based approach to recovering parametric full-body models from monocular images. This task is very challenging since minor parametric deviation may lead to noticeable misalignment between the estimated mesh and the input image. Moreover, when integrating part-specific estimations into the full-body model, existing solutions tend to either degrade the alignment or produce unnatural wrist poses. To address these issues, we propose a Pyramidal Mesh Alignment Feedback (PyMAF) loop in our regression network for well-aligned human mesh recovery and extend it as PyMAF-X for the recovery of expressive full-body models. The core idea of PyMAF is to leverage a feature pyramid and rectify the predicted parameters explicitly based on the mesh-image alignment status. Specifically, given the currently predicted parameters, mesh-aligned evidence will be extracted from finer-resolution features accordingly and fed back for parameter rectification. To enhance the alignment perception, an auxiliary dense supervision is employed to provide mesh-image correspondence guidance while spatial alignment attention is introduced to enable the awareness of the global contexts for our network. When extending PyMAF for full-body mesh recovery, an adaptive integration strategy is proposed in PyMAF-X to produce natural wrist poses while maintaining the well-aligned performance of the part-specific estimations. The efficacy of our approach is validated on several benchmark datasets for body, hand, face, and full-body mesh recovery, where PyMAF and PyMAF-X effectively improve the mesh-image alignment and achieve new state-of-the-art results. The project page with code and video results can be found at https://zhanghongwen.cn/pymaf-x.

cs.CV

Modulation of Electroosmotic Flow through Short Nanopores by Charged Exterior Surfaces

Electroosmotic flow (EOF) through nanoporous membranes has broad applications in micro- and nanofluidic systems, particularly in biomedical diagnostics and chemical analysis. The use of short nanopores enables high fluid flux, and the presence of exterior surface charges can further enhance ion flux through short nanopores. Here, systematic simulations are conducted to explore the modulation of EOF by exterior surface charges. Our results indicate that charged exterior surfaces can provide an additional effective pathway for fluid flow, significantly increasing both the EOF velocity and output pressure. By analyzing the dependence of EOF velocity on the area of the charged exterior surface, we derive the effective width (Lcs_eff) of the charged ring region extending beyond the pore boundary. This parameter is quantitatively examined under various nanopore configurations and applied conditions. Lcs_eff is found to be proportional to the pore diameter, surface charge density, and applied voltage, while inversely proportional to pore length and salt concentration. These findings provide valuable insights into the modulation of EOF by exterior surface charges and offer theoretical guidance for optimizing the structural and functional properties of nanoporous membranes in practical applications of EOF.

physics.chem-ph

Modulation of Ionic Current Rectification in Short Unipolar Nanopores

With controlled ionic current rectification (ICR) achieved through a strategically designed non-uniform surface charge distribution, short unipolar nanopores exhibit promising applications in nanofluidic sensors, ionic circuits, and ion amplifiers. By systematically investigating how the charged length on inner pore walls modulates ion transport, we found that both the maximum ICR degree and the corresponding charged-length proportion were influenced by nanopore parameters and simulation conditions. For 100 nm-long unipolar nanopores, the highest ICR degree is obtained at a charged-length proportion of ~0.3, due to the corresponding most significant ion enrichment and depletion inside the nanopore under opposite biases. This charged-length proportion of ~0.3 consistently appears as a characteristic value across most considered cases. For short unipolar nanopores, the presence of exterior surface charges significantly enhances the ICR degree by facilitating ion transport through nanopores. The effective widths of charged regions beyond nanopore borders on outer surfaces exhibit direct proportionality to the pore diameter, surface charge density, and applied voltage, and inverse proportionality to the pore length and salt concentration. Our research may provide useful guidance for the design of unipolar nanopores and porous membranes incorporating such charge configurations.

physics.chem-ph

Characteristics of mono-, di-, and trivalent cations in electric double layers: a molecular dynamic investigation

Ionic behaviors, including ion distributions and hydration characteristics at solid-liquid interfaces, are important research interests in many important applications, such as electric double-layer capacitors and water lubrication. Here, we systematically investigated the concentration distributions, hydration numbers, and screening properties of Li+, Na+, K+, Ca2+, Mg2+, and La3+ ions inside electric double layers (EDLs) at various charge densities. For the surface charge density weaker than -0.16 C m-2, monovalent cations mainly accumulate in the outer Helmholtz plane (OHP). As the charge density magnitude increases, monovalent cations start to dehydrate and migrate to the inner Helmholtz plane (IHP), following the order of K+, Na+, and Li+. This size-dependent behavior arises from the lower hydration energy of larger ions. While for the di- and trivalent ions, no obvious IHP appears. Based on ion distributions, the screening effect of counterions on surface charges is evaluated by analyzing the net charge distributions. As the charge density changes from 0 to -0.32 C m-2, due to the stronger accumulation of cations in EDLs, the location of the neutral plane changes from ~1.2 to ~0.4 nm. When the charge density reaches a threshold, excessive accumulation of cations can induce charge inversion. The threshold value and the maximum reversed charge are found to correlate with the ion size, cation valence, and concentration.

physics.chem-ph

HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception

Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction approaches struggle to coordinate fine-grained manipulations with long-range navigation. To address these limitations, we propose HOSIG, a novel framework for synthesizing full-body interactions through hierarchical scene perception. Our method decouples the task into three key components: 1) a scene-aware grasp pose generator that ensures collision-free whole-body postures with precise hand-object contact by integrating local geometry constraints, 2) a heuristic navigation algorithm that autonomously plans obstacle-avoiding paths in complex indoor environments via compressed 2D floor maps and dual-component spatial reasoning, and 3) a scene-guided motion diffusion model that generates trajectory-controlled, full-body motions with finger-level accuracy by incorporating spatial anchors and dual-space classifier-free guidance. Extensive experiments on the TRUMANS dataset demonstrate superior performance over state-of-the-art methods. Notably, our framework supports unlimited motion length through autoregressive generation and requires minimal manual intervention. This work bridges the critical gap between scene-aware navigation and dexterous object manipulation, advancing the frontier of embodied interaction synthesis. Codes will be available after publication. Project page: http://yw0208.github.io/hosig

cs.CV

Modulation of Memristive Characteristics by Dynamic Nanoprecipitation inside Conical Nanopores

Nanofluidic memristors have demonstrated great potential for neuromorphic system applications with the advantages of low energy consumption and excellent biocompatibility. Here, an effective way is developed to regulate the memristive behavior of conical nanopores by leveraging the reversible formation and dissolution of nanoprecipitates induced by ion enrichment and depletion in nanopores under opposite voltages. Through the interplay between precipitation dynamics at the pore tip and the ion enrichment/depletion inside the nanopore, conical nanopores exhibit pronounced current hysteresis loops in the presence of CaHPO4, a slightly soluble inorganic salt. The memristive characteristics are found to be strongly dependent on the concentration of CaHPO4, besides the applied voltage amplitude and scan rate. Under the stimulation of pulse voltages, ionic current demonstrates stable learning and forgetting processes with robust switching stability and effective reset capability, which is similar to the short-term plasticity characteristics of biological synapses. Our research may provide a straightforward and tunable approach for the design of nanofluidic memristors.

physics.chem-ph

Ion transport through differently charged nanoporous membranes: from a single nanopore to multi-nanopores

Nanoporous membranes, leveraging their high-throughput characteristics, have been widely applied in fields such as molecular separation and energy conversion. Due to interpore interactions, besides the applied voltage and solution environment, the ion transport properties in porous membranes are influenced by the pore number and spacing. Here, to understand and control the transport properties of nanopore arrays, we systematically investigate the ion transport characteristics through membranes with different charge properties, pore numbers, and interpore distances. Using numerical simulations, we analyzed local ionic concentrations and electric potential in nanopore arrays containing nanopores with uniformly charged walls as well as unipolar diodes i.e., pores containing a junction between a charged zone and a neutral zone, and showed significant ion concentration polarization (ICP) for all studied cases. As the number of pores increased and the interpore spacing decreased, the enhanced interpore interactions through ICP led to a greater deviation of the total ionic current from the linear superposition of single-pore currents. Conversely, in bipolar nanopores whose walls contain a junction between positively and negatively charged zones ICP becomes negligible, and interpore interactions are substantially reduced. Furthermore, for membranes with various charge properties, the total current through nanopore arrays presents different quantitative dependence on the pore number under varying pore spacings. Our findings clarify the mechanism of interpore interactions in modulating ion transport through porous membranes, providing critical insights for designing nanofluidic devices based on nanopore arrays, such as nanopore-array sensors.

cond-mat.soft

Ionic current rectification under concentration gradients and its application in evaluating surface charge properties of micropores

Ionic current rectification (ICR) induced by electroosmotic flow (EOF) under concentration gradients can find many applications in micro/nanofluidic sensing and ionic circuits. Here, we focused on the cases with micropores of moderate length-diameter ratios, through experimental research and systematical simulations, the EOF-induced ICR was found to exhibit voltage-dependent ratios. In the considered cases with a weak EOF or strong ionic diffusion, a large deviation appears between the ion concentration inside the micropore and the bulk value, which fails the prediction by solution conductivity gradients. Based on our simulation results, effective equations were developed for the theoretical description of ion concentration distributions along the micropore axis under coupled concentration gradient and electric field. With the predicted ion distributions inside micropores, the ICR ratio can be conveniently calculated with the derived electrical resistance of the microfluidic system, which applies to micropores of 200 to 1000 nm in diameter. Because the surface charge density is the only unknown input parameter, our developed equations can be used to evaluate the surface charge density of micropores with the measured EOF-induced ICR ratio under concentration gradients.

physics.chem-ph

DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective

We present a novel framework to reconstruct human avatars from monocular videos. Recent approaches have struggled either to capture the fine-grained dynamic details from the input or to generate plausible details at novel viewpoints, which mainly stem from the limited representational capacity of the avatar model and insufficient observational data. To overcome these challenges, we propose to leverage the advanced video generative model, Human4DiT, to generate the human motions from alternative perspective as an additional supervision signal. This approach not only enriches the details in previously unseen regions but also effectively regularizes the avatar representation to mitigate artifacts. Furthermore, we introduce two complementary strategies to enhance video generation: To ensure consistent reproduction of human motion, we inject the physical identity into the model through video fine-tuning. For higher-resolution outputs with finer details, a patch-based denoising algorithm is employed. Experimental results demonstrate that our method outperforms recent state-of-the-art approaches and validate the effectiveness of our proposed strategies.

cs.CV