SearcharxivSearch

arXiv subjects

Xiaoxuan Ma

Publications and source records attributed to Xiaoxuan Ma.

At least 19 recordsLinked to original sources

SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image

Part-aware 3D asset generation enables applications such as editing, articulation, simulation, and fabrication, yet existing methods can generate visually complete individual parts without ensuring that they form a valid physical assembly. Consequently, generated neighboring parts may interpenetrate, lack valid connections, or collapse under gravity. We propose a physics-guided framework for improving single-image part-aware 3D generation with physically compatible geometry and stable connections. Our method resolves inter-part penetration, recovers a contact graph between neighboring parts, and introduces parameterized connectors at their contact surfaces. Using feedback from physical simulation, we refine connector placement, orientation, and dimensions to improve assembly stability while preserving the generated geometry. We further introduce a physics-based evaluation protocol that complements conventional geometric metrics by directly testing assembly validity and stability under gravity. Experiments comparing against multiple part-aware 3D generators show substantial improvements in physical realizability and stability while maintaining geometric quality. We additionally validate the resulting parts through 3D printing and real-world assembly.

cs.GR

Observation of Magnetic-Anisotropy Crossover and High-Temperature Skyrmions in the Dirac Magnet Fe3Ge with a Distorted Kagome Lattice

Topological materials that simultaneously host robust high-temperature skyrmions and nontrivial electronic band structures have attracted tremendous interest owing to their distinctive advantages for both fundamental research and prospective technological applications. Here, we report the observation of robust skyrmions in the Dirac kagome magnet Fe3Ge, which exhibits a high Curie temperature of ~ 650 K. At room temperature, Fe3Ge shows a large intrinsic anomalous Hall conductivity of ~ 380 Ω-1cm-1, originating from its nontrivial electronic band topology. Systematic magnetization measurements reveal a spin reorientation transition at ~ 375 K, indicating a crossover from easy-plane to easy-axis magnetic anisotropy. Below the spin reorientation temperature, a large topological Hall effect is observed, arising from microscopic noncoplanar spin structures. Lorentz transmission electron microscopy shows that mesoscopic skyrmions are stabilized in the easy-axis magnetic anisotropy regime and persist over an exceptionally wide temperature window of 375-650 K, far exceeding that of most previously reported skyrmion-hosting materials. These results establish Fe3Ge as a promising platform for exploring diverse topological properties, with strong potential for advancing future high-temperature spintronic applications, ranging from next-generation information storage to logic computing devices.

cond-mat.mtrl-sci

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing single-image reconstruction methods fall short in capturing the physical structure of a scene. As a result, they often produce geometrically plausible but physically inconsistent results, including object floating and penetration, which lead to unstable behavior in physics simulations. Image-conditioned scene generation methods improve physical plausibility but often rely on strong scene priors, yielding plausible yet inaccurate object arrangements that fail to match the input image. We propose REST3D, a single-image reconstruction framework that can reconstruct physically stable 3D scenes by integrating physical scene understanding with physics-constrained refinement. We first introduce an agentic physical scene understanding technique that constructs a scene-tree representation capturing object physical states and inter-object relationships from a gravity-support perspective, providing a structural prior for reconstruction. Leveraging this structure, we initialize the scene using image-to-3D models, followed by scene-tree-guided alignment and physics-constrained optimization to resolve physical violations while preserving visual consistency with the input image. Experiments show that our method significantly reduces physical errors and improves simulation stability on both synthetic and real-world datasets while maintaining strong reconstruction quality. We further demonstrate the reconstructed scenes in VR-based human-object interaction, showing their potential for immersive applications.

cs.CV

InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding

Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene understanding remain largely category-oriented. In contrast, instance-aware 3DGS methods typically rely on per-scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed-forward 3DGS framework for generalizable 3D reconstruction and instance-aware scene understanding from pose-free multi-view images. In a single forward pass, InstanceSplat constructs an instance-aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language-aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross-view-consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance-centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language-aligned semantics strengthen the discrimination of confusing same-category instances, and instance regions aggregate semantic evidence into coherent object-level predictions. Experiments on novel-view synthesis, instance segmentation, and open-vocabulary semantic understanding under varying input-view settings and on an unseen dataset demonstrate state-of-the-art performance, practical efficiency, and strong generalization.

cs.CV

Emergence of multiple quasi-ferromagnetic magnon modes induced by strong magnetoelastic coupling in $TmFeO_3$ single crystal

We investigate the magnetization dynamics of $TmFeO_3$ single crystals across the spin-reorientation phase transition using broadband microwave absorption spectroscopy up to 87.5 GHz. Temperature- and magnetic-field-dependent antiferromagnetic resonance measurements reveal the characteristic softening of the quasi-ferromagnetic (q-FM) resonance mode at the $Γ_2\rightarrowΓ_{24}$ and $Γ_{24}\rightarrowΓ_4$ transition points. The finite magnon gap observed at the transition points reflects the strong magnetoelastic coupling. In addition to the uniform q-FM mode, multiple magnon modes appear in the intermediate $Γ_{24}$ phase, separated by approximately 0.5--2 GHz and exhibiting similar field and temperature dependence. These additional modes are attributed to nonuniform spin-wave excitations arising from the periodic magnetic domain structure present in the intermediate phase and their hybridization with acoustic phonons mediated by strong magnetoelastic coupling. Our results demonstrate that the spin-reorientation transition in $TmFeO_3$ provides a natural platform for generating multiple hybridized magnon modes, offering new opportunities for tunable magnonic excitations in rare-earth orthoferrites.

cond-mat.mtrl-sci

Observation of cooperative strong coupling between optical phonon and crystal-field excitations in a pseudo Jahn-Teller system

Cooperative interactions between localized electronic excitations and crystal lattice are central to the emergence of complex structural phases in materials. However, the scaling relations governing these collective behaviors remain largely unexplored. Here, using magneto-Raman spectroscopy, we report the direct observation of the strongly coupled optical phonon and non-degenerate crystal-field excitations (CFEs) in ErFeO3. By independently tuning the effective population of Jahn-Teller-active erbium ions through temperature and chemical dilution with Jahn-Teller-inactive yttrium ions, we identify the coupling strength varies linearly with the square root of electronic excitations population. Notably, Y-doping reveals the hybridization gap reduces significantly faster than predicted by density scaling alone, indicating phonon coherence is essential for establishing this cooperative interaction. Our findings highlight the role of optical phonons in mediating short-range interactions that drive cooperative Jahn-Teller effect, evidencing the pathway for tailoring electronic and vibrational properties of Jahn-Teller materials through population control.

cond-mat.mtrl-sci

The PanAf-SBR Dataset: Social Behaviour Recognition for Wild Great Apes

Behavioural shifts in wild great ape populations, particularly the breakdown of social structures, can serve as an early indicator of population decline. Automating the detection of behaviours indicative of these shifts is therefore a critical task for conservation. Several valuable datasets have recently been introduced for the automated recognition of great ape behaviour, yet few include fine-grained social behaviour annotations, and those that do are captured either in captive settings or via aerial platforms such as UAVs. We address this gap by introducing PanAf-SBR, the first wild great ape camera trap dataset annotated with social behaviours. PanAf-SBR extends PanAf500 with 100 additional videos covering 36,063 frames. These come with 81,096 annotations including bounding boxes, segmentation masks, intra-video identities, and seven social behaviour classes defined under the action giver and receiver convention of ChimpACT. We use this data together with the AlphaChimp architecture to establish the first benchmarks for fine-grained social behaviour recognition in wild great apes from camera trap footage. We further conduct bidirectional transfer learning experiments between PanAf-SBR and the captive ChimpACT dataset, finding that cross-dataset pre-training is highly beneficial for specific classes rather than of uniform benefit. Finally, we examine the role of background context by inverting the segmentation masks to suppress non-ape pixels.

cs.CV

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Language Models excel at dialogue but lack embodied expression, while diffusion-based talking head models achieve visual fidelity but ignore social cognition. To bridge this gap, we propose a closed-loop dual-agent framework integrating perception, social reasoning, and expression into a continuous interaction cycle. The perception module analyzes partners' multimodal behaviors from video, while the social reasoning module infers hidden mental states through Theory of Mind and selects responses via an ensemble mechanism. The expression module then generates emotion-controllable videos that jointly synthesize speaker speech and facial expressions with listener reactive behaviors, capturing bidirectional dynamics absent in prior work. We further construct a hierarchical Persona-Scenario dataset with psychologically grounded personas and private social goals to support evaluation under information asymmetry. Experiments on this dataset demonstrate competitive or superior performance on both dialogue quality and video generation metrics. Notably, our method surpasses even the full-information Script mode on key dialogue quality dimensions, suggesting that explicit mental state inference under uncertainty can elicit more thoughtful dialogue than unrestricted information access. Project page: https://resonantminds.github.io/.

cs.CV

Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild

Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geometry and appearance. Prior approaches either learn to directly predict 4D representations from visual input or initialize a 3D representation that is subsequently deformed and refined based on video evidence. However, the former are constrained by the scarcity of 4D training data, while the latter leverage priors only for the initial reconstruction and rely solely on video supervision thereafter; neither handles complex in-the-wild scenarios with large deformations and occlusions well. We present Lift4D, a test-time optimization framework that addresses both limitations. First, we adapt an existing single-view 3D reconstruction model to yield temporally consistent per-frame predictions via causal latent conditioning, providing a coherent initialization for a deformable 3D Gaussian Splatting representation. We then ``sculpt'' this representation to match the input video through an occlusion-aware optimization that faithfully recovers visible surface details while completing unobserved regions using a view-conditioned diffusion prior. We demonstrate that Lift4D clearly improves over prior 4D reconstruction methods, particularly on challenging in-the-wild sequences with severe occlusions and non-rigid motion.

cs.CV

RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation

Text-to-image (T2I) diffusion models have shown remarkable success in generating high-quality images from text prompts. Recent efforts extend these models to incorporate conditional images (e.g., canny edge) for fine-grained spatial control. Among them, feature injection methods have emerged as a training-free alternative to traditional fine-tuning-based approaches. However, they often suffer from structural misalignment, condition leakage, and visual artifacts, especially when the condition image diverges significantly from natural RGB distributions. Through an analysis of existing methods, we identify a key limitation: the sampling schedule of condition features, previously unexplored, fails to account for the evolving interplay between structure preservation and domain alignment throughout diffusion steps. Inspired by this observation, we propose a flexible training-free framework that decouples the sampling schedule of condition features from the denoising process, and systematically investigate the spectrum of feature injection schedules to achieve a better balance between structural alignment and appearance quality. We further enhance the sampling process by introducing a restart refinement schedule, and improve the visual quality with an appearance-rich prompting strategy. Together, these designs enable training-free controllable generation that is both structure-rich and appearance-rich. Extensive experiments demonstrate that our method achieves state-of-the-art performance under complex and diverse conditions. Owing to its generality, our framework naturally supports compositional conditional generation and generalizes across architectures in a plug-and-play manner, from UNet-based diffusion models to modern DiT backbones such as FLUX.

cs.CV

How Label Imbalance Shapes Geometry: A General Spectral Analysis of Multi-Label Neural Collapse

This work investigates the phenomenon of Neural Collapse (NC) in multi-label classification, extending its conceptual framework from multi-class learning to general correlated and imbalanced multi-label settings. Although recent studies have identified a ''tag-wise averaging'' structure for multi-label features, this view relies on implicit assumptions of label balance and combinatorial symmetry. Consequently, it fails to account for the geometrical distortions caused by intrinsic label correlations and data imbalance, which are common in practice. We resolve the multiplicity-one imbalance conjecture raised by Li et al. (2024), showing that higher-multiplicity prototypes obey a class-frequency-weighted synthesis rule rather than uniform averaging. To address this, we propose a rigorous spectral-control framework to analyze the terminal phase of multi-label learning under general imbalanced conditions. We introduce the label covariance spectrum $κ_m$, a scalar controlling the distribution-dependent lower-bound geometry, derived from the second-order moment matrix of the label distribution. Contrary to the averaging perspective, our analysis reveals that the centered label covariance spectrum controls the stability of terminal geometry by quantifying the weakest centered inter-class contrast directions. We prove that the classical Tag-wise Averaging emerges only as a special case under perfect orthogonality. Numerical experiments on synthetic distributions validate our theoretical bounds. This work resolves the scaled-average aspect of the imbalance conjecture and establishes a unifying theoretical framework that extends Neural Collapse to complex, imbalanced multi-label settings.

cs.LG

Visually-grounded Humanoid Agents

Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to novel environments. We instead ask: how can digital humans actively behave using only visual observations and specified goals in novel scenes? Achieving this would enable populating any 3D environments with digital humans at scale that exhibit spontaneous, natural, goal-directed behaviors. To this end, we introduce Visually-grounded Humanoid Agents, a coupled two-layer (world-agent) paradigm that replicates humans at multiple levels: they look, perceive, reason, and behave like real people in real-world 3D scenes. The World Layer reconstructs semantically rich 3D Gaussian scenes from real-world videos via an occlusion-aware pipeline and accommodates animatable Gaussian-based human avatars. The Agent Layer transforms these avatars into autonomous humanoid agents, equipping them with first-person RGB-D perception and enabling them to perform accurate, embodied planning with spatial awareness and iterative reasoning, which is then executed at the low level as full-body actions to drive their behaviors in the scene. We further introduce a benchmark to evaluate humanoid-scene interaction in diverse reconstructed environments. Experiments show our agents achieve robust autonomous behavior, yielding higher task success rates and fewer collisions than ablations and state-of-the-art planning methods. This work enables active digital human population and advances human-centric embodied AI. Data, code, and models will be open-sourced.

cs.CV

Electromagnetic Inverse Scattering from a Single Transmitter

Electromagnetic Inverse Scattering Problems (EISP) seek to reconstruct relative permittivity from scattered fields and are fundamental to applications like medical imaging. This inverse process is inherently ill-posed and highly nonlinear, making it particularly challenging, especially under sparse transmitter setups, e.g., with only one transmitter. While recent machine learning-based approaches have shown promising results, they often rely on time-consuming, case-specific optimization and perform poorly under sparse transmitter setups. To address these limitations, we revisit EISP from a data-driven perspective. The scarcity of transmitters leads to an insufficient amount of measured data, which fails to capture adequate physical information for stable inversion. Accordingly, we propose a fully end-to-end and data-driven framework that predicts the relative permittivity of scatterers from measured fields, leveraging data distribution priors to compensate for the incomplete information from sparse measurements. This design enables data-driven training and feed-forward prediction of relative permittivity while maintaining strong robustness to transmitter sparsity. Extensive experiments show that our method outperforms state-of-the-art approaches in reconstruction accuracy and robustness. Notably, we demonstrate, for the first time, high-quality reconstruction from a single transmitter. This work advances practical electromagnetic imaging by providing a new, cost-effective paradigm to inverse scattering. Code and models are released at https://gomenei.github.io/SingleTX-EISP/.

cs.CV

Electrically Accessible Metamagnetic Transition via a Doping-Induced Low-Energy Magnetic State in Antiferromagnetic Insulator RFeO3

Low-energy antiferromagnetic phase transitions offer an appealing platform for low-power spintronic functionalities, yet their direct electrical access in insulating antiferromagnets remains challenging, particularly in the low-field regime where subtle Neeel vector reorientations dominate. Here, we demonstrate that targeted rare-earth-site engineering enables an electrically accessible metamagnetic transition in the insulating orthoferrite Ho0.5Dy0.5FeO3. By combining the distinct spin-reorientation sequences of DyFeO3 and HoFeO3, Dy substitution stabilizes a dual spin-reorientation pathway, hosting an intermediate state with a reduced energy barrier. This low-energy antiferromagnetic state can be tuned into the weak-ferromagnetic state under low magnetic fields. The critical field decreases with increasing temperature, providing a favorable window for functional manipulation. Both longitudinal and transverse spin Hall magnetoresistance channels exhibit clear and reproducible signatures of the metamagnetic transitions. Owing to the enhanced sensitivity of the transverse channel, additional low-field features are resolved, reflecting the projection of the Neel vector onto the spin-accumulation direction. Electrical transport measurements correlate directly with the magnetically determined phase boundaries, establishing a purely electrical access to low-energy phase transitions and to illustrate a viable pathway for exploring low-power spin dynamics in insulating oxide antiferromagnets.

cond-mat.mtrl-sci

Efficient Action Counting with Dynamic Queries

Temporal repetition counting aims to quantify the repeated action cycles within a video. The majority of existing methods rely on the similarity correlation matrix to characterize the repetitiveness of actions, but their scalability is hindered due to the quadratic computational complexity. In this work, we introduce a novel approach that employs an action query representation to localize repeated action cycles with linear computational complexity. Based on this representation, we further develop two key components to tackle the essential challenges of temporal repetition counting. Firstly, to facilitate open-set action counting, we propose the dynamic update scheme on action queries. Unlike static action queries, this approach dynamically embeds video features into action queries, offering a more flexible and generalizable representation. Secondly, to distinguish between actions of interest and background noise actions, we incorporate inter-query contrastive learning to regularize the video representations corresponding to different action queries. As a result, our method significantly outperforms previous works, particularly in terms of long video sequences, unseen actions, and actions at various speeds. On the challenging RepCountA benchmark, we outperform the state-of-the-art method TransRAC by 26.5% in OBO accuracy, with a 22.7% mean error decrease and 94.1% computational burden reduction. Code is available at https://github.com/lizishi/DeTRC.

cs.CV

Seeing My Future: Predicting Situated Interaction Behavior in Virtual Reality

Virtual and augmented reality systems increasingly demand intelligent adaptation to user behaviors for enhanced interaction experiences. Achieving this requires accurately understanding human intentions and predicting future situated behaviors - such as gaze direction and object interactions - which is vital for creating responsive VR/AR environments and applications like personalized assistants. However, accurate behavioral prediction demands modeling the underlying cognitive processes that drive human-environment interactions. In this work, we introduce a hierarchical, intention-aware framework that models human intentions and predicts detailed situated behaviors by leveraging cognitive mechanisms. Given historical human dynamics and the observation of scene contexts, our framework first identifies potential interaction targets and forecasts fine-grained future behaviors. We propose a dynamic Graph Convolutional Network (GCN) to effectively capture human-environment relationships. Extensive experiments on challenging real-world benchmarks and live VR environment demonstrate the effectiveness of our approach, achieving superior performance across all metrics and enabling practical applications for proactive VR systems that anticipate user behaviors and adapt virtual environments accordingly.

cs.CV

Altermagnetic magnon transport in the \textit{d}-wave altermagnet \ch{LuFeO3}

Altermagnets exhibit a spin-split band structure despite having zero net magnetization, leading to special magnonic properties such as anisotropic magnon lifetimes and field-free spin transport. Here, we present a direct experimental demonstration of non-local magnon transport in the \textit{d}-wave altermagnet \ch{LuFeO3}, using both spin Seebeck and spin Hall effect-based injection and detection. We observe a non-local spin signal at zero magnetic field when the transport is along an altermagnetic direction, but not for transport along other directions. The observed sign reversal between two distinct altermagnetic directions in the spin Seebeck response demonstrates the altermagnetic nature of the magnon transport. In contrast, when transport is aligned along or perpendicular to the easy axis, both the first-harmonic signal and the sign-reversal effect vanish, consistent with symmetry-imposed suppression. These findings are supported by atomistic spin dynamics simulations, as well as linear spin wave theory calculations, which explain how our altermagnetic system hosts anisotropic spin Seebeck transport. Our results provide direct evidence of direction-dependent magnon splitting in altermagnets and highlight their potential for field-free magnonic spin transport, offering a promising pathway for low-power spintronic applications.

cond-mat.mtrl-sci

SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens

We propose a one-stage framework for real-time multi-person 3D human mesh estimation from a single RGB image. While current one-stage methods, which follow a DETR-style pipeline, achieve state-of-the-art (SOTA) performance with high-resolution inputs, we observe that this particularly benefits the estimation of individuals in smaller scales of the image (e.g., those far from the camera), but at the cost of significantly increased computation overhead. To address this, we introduce scale-adaptive tokens that are dynamically adjusted based on the relative scale of each individual in the image within the DETR framework. Specifically, individuals in smaller scales are processed at higher resolutions, larger ones at lower resolutions, and background regions are further distilled. These scale-adaptive tokens more efficiently encode the image features, facilitating subsequent decoding to regress the human mesh, while allowing the model to allocate computational resources more effectively and focus on more challenging cases. Experiments show that our method preserves the accuracy benefits of high-resolution processing while substantially reducing computational cost, achieving real-time inference with performance comparable to SOTA methods.

cs.CV