SearcharxivSearch

arXiv subjects

Yang Fu

Publications and source records attributed to Yang Fu.

At least 37 records · Page 2Linked to original sources

Nearly-room-temperature ferromagnetism and tunable anomalous Hall effect in atomically thin Fe4CoGeTe2

Itinerant ferromagnetism at room temperature is a key ingredient for spin transport and manipulation. Here, we report the realization of nearly-room-temperature itinerant ferromagnetism in Co doped Fe5GeTe2 thin flakes. The ferromagnetic transition temperature TC (~ 323 K - 337 K) is almost unchanged when thickness is down to 12 nm and is still about 284 K at 2 nm (bilayer thickness). Theoretical calculations further indicate that the ferromagnetism persists in monolayer Fe4CoGeTe2. In addition to the robust ferromagnetism down to the ultrathin limit, Fe4CoGeTe2 exhibits an unusual temperature- and thickness-dependent intrinsic anomalous Hall effect. We propose that it could be ascribed to the dependence of band structure on thickness that changes the Berry curvature near the Fermi energy level subtly. The nearly-room-temperature ferromagnetism and tunable anomalous Hall effect in atomically thin Fe4CoGeTe2 provide opportunities to understand the exotic transport properties of two-dimensional van der Waals magnetic materials and explore their potential applications in spintronics.

cond-mat.mtrl-sci

An extraction of the Collins-Soper kernel from a joint analysis of experimental and lattice data

We present a first joint extraction of the Collins-Soper kernel (CSK) combining experimental and lattice QCD data in the context of an analysis of transverse-momentum-dependent distributions (TMDs). Based on a neural-network parametrization, we perform a Bayesian reweighting of an existing fits of TMDs using lattice data, as well as a joint TMD fit to lattice and experimental data. We consistently find that the inclusion of lattice information shifts the central value of the CSK by approximately 10% and reduces its uncertainty by 40-50%, highlighting the potential of lattice inputs to improve TMD extractions.

hep-ph

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality human speech videos with diverse camera shots, including close-up, half-body, and full-body views. The dataset includes detailed textual descriptions, 2D keypoints and 3D SMPL-X motion annotations, covering over 10k identities, enabling multimodal learning and evaluation. As a first attempt to showcase the value of the dataset, we present Orator, an LLM-guided multi-modal generation framework as a simple baseline, where the language model functions as a multi-faceted director, orchestrating detailed specifications for camera transitions, speaker gesticulations, and vocal modulation. This architecture enables the synthesis of coherent long-form videos through our integrated multi-modal video generation module. Extensive experiments in both pose-guided and audio-driven settings show that training on TalkCuts significantly enhances the cinematographic coherence and visual appeal of generated multi-shot speech videos. We believe TalkCuts provides a strong foundation for future work in controllable, multi-shot speech video generation and broader multimodal learning.

cs.CV

Superconductivity in kagome metal YRu3Si2 with strong electron correlations

We report the detailed physical properties of YRu3Si2 with the Ru kagome lattice at normal and superconducting states. The results of resistivity and magnetization show that YRu3Si2 is a type-II bulk superconductor with Tc ~ 3.0 K. The specific heat measurement further suggests that this superconductivity could originate from the weak or moderate electron-phonon coupling. On the other hand, both large Kadawaki-Woods ratio and Wilson ratio indicate that there is a strong electron correlation effect in this system, which may have a connection with the featured flat band of kagome lattice.

cond-mat.supr-con

Fine-Grained AI Model Caching and Downloading With Coordinated Multipoint Broadcasting in Multi-Cell Edge Networks

6G networks are envisioned to support on-demand AI model downloading to accommodate diverse inference requirements of end users. By proactively caching models at edge nodes, users can retrieve the requested models with low latency for on-device AI inference. However, the substantial size of contemporary AI models poses significant challenges for edge caching under limited storage capacity, as well as for the concurrent delivery of heterogeneous models over wireless channels. To address these challenges, we propose a fine-grained AI model caching and downloading system that exploits parameter reusability, stemming from the common practice of fine-tuning task-specific models from a shared pre-trained model with frozen parameters. This system selectively caches model parameter blocks (PBs) at edge nodes, eliminating redundant storage of reusable parameters across different cached models. Additionally, it incorporates coordinated multipoint (CoMP) broadcasting to simultaneously deliver reusable PBs to multiple users, thereby enhancing downlink spectrum utilization. Under this arrangement, we formulate a model downloading delay minimization problem to jointly optimize PB caching, migration (among edge nodes), and broadcasting beamforming. To tackle this intractable problem, we develop a distributed multi-agent learning framework that enables edge nodes to explicitly learn mutual influence among their actions, thereby facilitating cooperation. Furthermore, a data augmentation approach is proposed to adaptively generate synthetic training samples through a predictive model, boosting sample efficiency and accelerating policy learning. Both theoretical analysis and simulation experiments validate the superior convergence performance of the proposed learning framework.

cs.NI

Highly Efficient Room-Temperature Nonvolatile Magnetic Switching by Current in Fe3GaTe2 Thin Flakes

Effectively tuning magnetic state by using current is essential for novel spintronic devices. Magnetic van der Waals (vdW) materials have shown superior properties for the applications of magnetic information storage based on the efficient spin torque effect. However, for most of known vdW ferromagnets, the ferromagnetic transition temperatures lower than room temperature strongly impede their applications and the room-temperature vdW spintronic device with low energy consumption is still a long-sought goal. Here, we realize the highly efficient room-temperature nonvolatile magnetic switching by current in a single-material device based on vdW ferromagnet Fe3GaTe2. Moreover, the switching current density and power dissipation are about 300 and 60000 times smaller than conventional spin-orbit-torque devices of magnet/heavy-metal heterostructures. These findings make an important progress on the applications of magnetic vdW materials in the fields of spintronics and magnetic information storage.

cond-mat.mtrl-sci

3D Aware Region Prompted Vision Language Model

We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or directly in 3D, without the need for exhaustive multi-frame labeling. We achieve this by enriching 2D visual features with 3D positional embeddings, which allows the 3D model to draw upon strong 2D priors for more accurate spatial reasoning across frames, even when objects of interest do not co-occur within the same view. Extensive experiments on both general 2D vision language and specialized 3D spatial benchmarks demonstrate that SR-3D achieves state-of-the-art performance, underscoring its effectiveness for unifying 2D and 3D representation space on scene understanding. Moreover, we observe applicability to in-the-wild videos without sensory 3D inputs or ground-truth 3D annotations, where SR-3D accurately infers spatial relationships and metric measurements.

cs.CV

Free-form conformal metasurfaces robustly generating topological skyrmions

Skyrmions are topologically stable vector textures as potential information carriers for high-density data storage and communications, especially boosted by the recently emerging meta-generators of skyrmions in electromagnetic fields. However, these implementations always rely on planar, rigid designs with stringent fabrication requirements. Here, we propose the free-form conformal metasurface generating skyrmions towards future wearable and flexible devises for topological resilience light fields. Furthermore, we experimentally tested the outstanding topological robustness of the skyrmion number under different disorder degrees on the metasurface. This work promotes the development of flexible compact skyrmion-based communication devices and demonstrates their potential to improve the quality of space information transmission.

physics.optics

Long-range angular correlations of particle displacements at a plastic-to-elastic transition in jammed amorphous solids

Understanding how a flow turns into an amorphous solid is a fundamental challenge in statistical physics, during which no apparent structural ordering appears. In the athermal limit, the two states are connected by a well-defined jamming transition, near which the solid is marginally stable. A recent mechanical response screening theory proposes an additional transition above jamming, called a plastic-to-elastic transition here, separating anomalous and quasi-elastic mechanical behavior. Through numerical inflation simulations in two dimensions, we show that the onsets of long-range radial and angular correlations of particle displacements decouple, occurring respectively at the jamming and plastic-to-elastic transitions. The latter is characterized by a power-law diverging correlation angle and a power-law spectrum of the displacements along a circle. This work establishes two-step transitions on the mechanical properties during ``decompression melting'' of an athermal over-jammed amorphous solid, reminiscent of the two-step structural melting of a crystal in two dimensions. In contradistinction with the latter, the plastic-to-elastic transition exists also in three dimensions.

cond-mat.soft

Lattice QCD calculation of the subtraction function in forward Compton amplitude

The subtraction function plays a pivotal role in calculations involving the forward Compton amplitude, which is crucial for predicting the Lamb shift in muonic atom, as well as the proton-neutron mass difference. In this work, we present a lattice QCD calculation of the subtraction function using two domain wall fermion gauge ensembles at the physical pion mass. We utilize a recently proposed subtraction point, demonstrating its advantage in mitigating statistical and systematic uncertainties by eliminating the need for ground-state subtraction. Our results reveal significant contributions from $Nπ$ intermediate states to the subtraction function. Incorporating these contributions, we compute the proton, neutron and nucleon isovector subtraction functions at photon momentum transfer $Q^2\in[0,2]$ GeV$^2$. For the proton subtraction function, we compare our lattice results with chiral perturbation theory prediction at low $Q^2$ and with the results from the perturbative operator-product expansion at high $Q^2$. Finally, using these subtraction functions as input, we determine their contribution to two-photon exchange effects in the Lamb shift and isovector nucleon electromagnetic self-energy.

hep-lat

Learning Generalizable Feature Fields for Mobile Manipulation

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the complexity inherent at an expansive physical scale. In this work, we present GeFF (Generalizable Feature Fields), a scene-level generalizable neural feature field that acts as a unified representation for both navigation and manipulation that performs in real-time. To do so, we treat generative novel view synthesis as a pre-training task, and then align the resulting rich scene priors with natural language via CLIP feature distillation. We demonstrate the effectiveness of this approach by deploying GeFF on a quadrupedal robot equipped with a manipulator. We quantitatively evaluate GeFF's ability for open-vocabulary object-/part-level manipulation and show that GeFF outperforms point-based baselines in runtime and storage-accuracy trade-offs, with qualitative examples of semantics-aware navigation and articulated object manipulation.

cs.RO

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs' spatial perception and reasoning capabilities. SpatialRGPT advances VLMs' spatial understanding through two key innovations: (1) a data curation pipeline that enables effective learning of regional representation from 3D scene graphs, and (2) a flexible plugin module for integrating depth information into the visual encoder of existing VLMs. During inference, when provided with user-specified region proposals, SpatialRGPT can accurately perceive their relative directions and distances. Additionally, we propose SpatialRGBT-Bench, a benchmark with ground-truth 3D annotations encompassing indoor, outdoor, and simulated environments, for evaluating 3D spatial cognition in VLMs. Our results demonstrate that SpatialRGPT significantly enhances performance in spatial reasoning tasks, both with and without local region prompts. The model also exhibits strong generalization capabilities, effectively reasoning about complex spatial relations and functioning as a region-aware dense reward annotator for robotic tasks. Code, dataset, and benchmark are released at https://www.anjiecheng.me/SpatialRGPT

cs.CV

COLMAP-Free 3D Gaussian Splatting

While neural rendering has led to impressive advances in scene reconstruction and novel view synthesis, it relies heavily on accurately pre-computed camera poses. To relax this constraint, multiple efforts have been made to train Neural Radiance Fields (NeRFs) without pre-processed camera poses. However, the implicit representations of NeRFs provide extra challenges to optimize the 3D structure and camera poses at the same time. On the other hand, the recently proposed 3D Gaussian Splatting provides new opportunities given its explicit point cloud representations. This paper leverages both the explicit geometric representation and the continuity of the input video stream to perform novel view synthesis without any SfM preprocessing. We process the input frames in a sequential manner and progressively grow the 3D Gaussians set by taking one input frame at a time, without the need to pre-compute the camera poses. Our method significantly improves over previous approaches in view synthesis and camera pose estimation under large motion changes. Our project page is https://oasisyang.github.io/colmap-free-3dgs

cs.CV

RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos

We introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing, the direct capture of the depth channel allows better 3D annotations and broader downstream applications. WildRGB-D comprises large-scale category-level RGB-D object videos, which are taken using an iPhone to go around the objects in 360 degrees. It contains around 8500 recorded objects and nearly 20000 RGB-D videos across 46 common object categories. These videos are taken with diverse cluttered backgrounds with three setups to cover as many real-world scenarios as possible: (i) a single object in one video; (ii) multiple objects in one video; and (iii) an object with a static hand in one video. The dataset is annotated with object masks, real-world scale camera poses, and reconstructed aggregated point clouds from RGBD videos. We benchmark four tasks with WildRGB-D including novel view synthesis, camera pose estimation, object 6d pose estimation, and object surface reconstruction. Our experiments show that the large-scale capture of RGB-D objects provides a large potential to advance 3D object learning. Our project page is https://wildrgbd.github.io/.

cs.CV

Odd Dipole Screening in Radial Inflation

The inflation of an inner radial (or spherical) cavity in an amorphous solids confined in a disk (or a sphere), served as a fruitful case model for studying the effects of plastic deformations on the mechanical response. It was shown that when the field associated with Eshelby quadrupolar charges is non-uniform, the displacement field is riddled with dipole charges that screen elasticity, reminiscent of Debye monopoles screening in electrostatics. In this paper we look deeper into the screening phenomenon, taking into account the consequences of irreversibility that are associated with the breaking of Chiral symmetry. We consider the equations for the displacement field with the presence of ``Odd Dipole Screening", solve them analytically and compare with numerical simulations. Suggestions how to test the theory in experiments are provided.

cond-mat.dis-nn

Macroscopic Tunneling Probe of Moiré Spin Textures in Twisted CrI$_3$

Various noncollinear spin textures and magnetic phases have been predicted in twisted two-dimensional CrI$_3$ due to competing ferromagnetic (FM) and antiferromagnetic (AFM) interlayer exchange from moiré stacking - with potential spintronic applications even when the underlying material possesses a negligible Dzyaloshinskii-Moriya or dipole-dipole interaction. Recent measurements have shown evidence of coexisting FM and AFM layer order in small-twist-angle CrI$_3$ bilayers and double bilayers. Yet, the nature of the magnetic textures remains unresolved and possibilities for their manipulation and electrical readout are unexplored. Here, we use tunneling magnetoresistance to investigate the collective spin states of twisted double-bilayer CrI$_3$ under both out-of-plane and in-plane magnetic fields together with detailed micromagnetic simulations of domain dynamics based on magnetic circular dichroism. Our results capture hysteretic and anisotropic field evolutions of the magnetic states and we further uncover two distinct non-volatile spin textures (out-of-plane and in-plane domains) at $\approx$ 1° twist angle, with a different global tunneling resistance that can be switched by magnetic field.

cond-mat.mtrl-sci

A Construct-Optimize Approach to Sparse View Synthesis without Camera Pose

Novel view synthesis from a sparse set of input images is a challenging problem of great practical interest, especially when camera poses are absent or inaccurate. Direct optimization of camera poses and usage of estimated depths in neural radiance field algorithms usually do not produce good results because of the coupling between poses and depths, and inaccuracies in monocular depth estimation. In this paper, we leverage the recent 3D Gaussian splatting method to develop a novel construct-and-optimize method for sparse view synthesis without camera poses. Specifically, we construct a solution progressively by using monocular depth and projecting pixels back into the 3D world. During construction, we optimize the solution by detecting 2D correspondences between training views and the corresponding rendered images. We develop a unified differentiable pipeline for camera registration and adjustment of both camera poses and depths, followed by back-projection. We also introduce a novel notion of an expected surface in Gaussian splatting, which is critical to our optimization. These steps enable a coarse solution, which can then be low-pass filtered and refined using standard optimization methods. We demonstrate results on the Tanks and Temples and Static Hikes datasets with as few as three widely-spaced views, showing significantly better quality than competing methods, including those with approximate camera pose information. Moreover, our results improve with more views and outperform previous InstantNGP and Gaussian Splatting algorithms even when using half the dataset. Project page: https://raymondjiangkw.github.io/cogs.github.io/

cs.CV

HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-object geometric structure and text description as inputs for image synthesis. This offers a more controllable and realistic synthesis as we can specify the structure and style inputs in a disentangled manner. HOIDiffusion is trained by leveraging a diffusion model pre-trained on large-scale natural images and a few 3D human demonstrations. Beyond controllable image synthesis, we adopt the generated 3D data for learning 6D object pose estimation and show its effectiveness in improving perception systems. Project page: https://mq-zhang1.github.io/HOIDiffusion

cs.CV