SearcharxivSearch

arXiv subjects

Yichen Jin

Publications and source records attributed to Yichen Jin.

18 recordsLinked to original sources

Writing and erasing skyrmions by single ultrafast laser pulses in monolayer Janus 2D magnets

Skyrmions in 2D magnets are promising candidates for nonvolatile, low-power, and high-density spintronic memories. However, their experimental realization at the 2D limit remains challenging, owing to the difficulty in engineering the required chiral magnetic interactions. Here, we report the creation and direct imaging of N\'eel-type skyrmions in Janus 2D chromium chalcogenides using synchrotron X-ray photoemission electron microscopy, and scanning nitrogen-vacancy magnetometry, which exhibit field-free stability, nonvolatility, and size tunability. First-principles calculations and micromagnetic simulations reveal that Janus-surface-induced inversion-symmetry breaking enhances the Dzyaloshinskii-Moriya interaction, providing the microscopic mechanism for skyrmion stabilization and tunability. We further achieve reversible skyrmion writing and erasing using a single ultrafast laser pulse in a magnetic field as low as 300 Oe, demonstrating the excellent manipulability of this 2D magnetic system. These results establish Janus engineering as a route to creating and manipulating nonvolatile skyrmions in atomically thin magnets, with implications for skyrmion-based low-power spintronic devices.

cond-mat.mtrl-sci

ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting

Recent advancements in 3D Gaussian Splatting (3DGS) have enabled language-guided scene understanding. However, existing Referring 3D Gaussian Splatting (R3DGS) methods are fundamentally restricted to single-target queries. To reflect the ambiguity of real-world instructions, we introduce the Generalized Referring 3D Gaussian Splatting Segmentation (GR3DGS) task, which requires dynamically segmenting an arbitrary number of targets (0, 1, or $N$). To facilitate comprehensive evaluation of this new task, we construct two new benchmarks: GR-LERF and GR-ScanNet. Crucially, existing R3DGS paradigms exhibit fundamental technical bottlenecks that severely limit their performance on the GR3DGS task: they lack intrinsic 3D point-level understanding by operating merely on 2D rendered pixels, and they incur prohibitive computational overhead by requiring per-scene optimization to embed heavy semantic features. To dismantle these bottlenecks, we propose ZeroSplat, a novel training-free and zero-feature framework. ZeroSplat lifts 2D Vision-Language Model (VLM) priors into 3D space through robust multi-view geometric constraints. This strategy enables intrinsic point-level understanding without incurring any additional feature storage. Extensive experiments demonstrate that ZeroSplat significantly outperforms state-of-the-art methods across generalized and single-target scenarios while maintaining exceptional efficiency. Project Page: https://inkmind-ai.github.io/ZeroSplat

cs.CV

VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition

Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or naive curriculum strategies based on scalar difficulty metrics. We argue that scalar metrics fail to disentangle two orthogonal challenges in video understanding: Visual Temporal Perception Load and Cognitive Reasoning Depth. To address this, we propose VideoCuRL, a novel framework that decomposes difficulty into these two axes. We employ efficient, training-free proxies, optical flow and keyframe entropy for visual complexity, Calibrated Surprisal for cognitive complexity, to map data onto a 2D curriculum grid. A competence aware Diagonal Wavefront strategy then schedules training from base alignment to complex reasoning. Furthermore, we introduce Dynamic Sparse KL and Structured Revisiting to stabilize training against reward collapse and catastrophic forgetting. Extensive experiments show that VideoCuRL surpasses strong RL baselines on reasoning (+2.5 on VSI-Bench) and perception (+2.9 on VideoMME) tasks. Notably, VideoCuRL eliminates the prohibitive inference overhead of generation-based curricula, offering a scalable solution for robust video post-training.

cs.CV

MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are limited in terms of plausibility or alignment with the tasks. In this paper, we formulate a novel task, namely task-oriented tabletop scene generation, which poses significant challenges due to the substantial gap between high-level task instructions and the tabletop scenes. To support research on such a challenging task, we introduce MesaTask-10K, a large-scale dataset comprising approximately 10,700 synthetic tabletop scenes with manually crafted layouts that ensure realistic layouts and intricate inter-object relations. To bridge the gap between tasks and scenes, we propose a Spatial Reasoning Chain that decomposes the generation process into object inference, spatial interrelation reasoning, and scene graph construction for the final 3D layout. We present MesaTask, an LLM-based framework that utilizes this reasoning chain and is further enhanced with DPO algorithms to generate physically plausible tabletop scenes that align well with given task descriptions. Exhaustive experiments demonstrate the superior performance of MesaTask compared to baselines in generating task-conforming tabletop scenes with realistic layouts. Project page is at https://mesatask.github.io/

cs.CV

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collisions. To address these shortcomings, we introduce \textbf{InternScenes}, a novel large-scale simulatable indoor scene dataset comprising approximately 40,000 diverse scenes by integrating three disparate scene sources, real-world scans, procedurally generated scenes, and designer-created scenes, including 1.96M 3D objects and covering 15 common scene types and 288 object classes. We particularly preserve massive small items in the scenes, resulting in realistic and complex layouts with an average of 41.5 objects per region. Our comprehensive data processing pipeline ensures simulatability by creating real-to-sim replicas for real-world scans, enhances interactivity by incorporating interactive objects into these scenes, and resolves object collisions by physical simulations. We demonstrate the value of InternScenes with two benchmark applications: scene layout generation and point-goal navigation. Both show the new challenges posed by the complex and realistic layouts. More importantly, InternScenes paves the way for scaling up the model training for both tasks, making the generation and navigation in such complex scenes possible. We commit to open-sourcing the data, models, and benchmarks to benefit the whole community.

cs.CV

Balancing Latency and Model Accuracy for Fluid Antenna-Assisted LM-Embedded MIMO Network

This paper addresses the challenge of large model (LM)-embedded wireless network for handling the trade-off problem of model accuracy and network latency. To guarantee a high-quality of users' service, the network latency should be minimized while maintaining an acceptable inference accuracy. To meet this requirement, LM quantization is proposed to reduce the latency. However, the excessive quantization may destroy the accuracy of LM inference. To this end, a promising fluid antenna (FA) technology is investigated for enhancing the transmission capacity, leading to a lower network latency in the LM-embedded multiple-input multiple-output (MIMO) network. To design the FA-assisted LM-embedded network with the lower latency and higher accuracy requirements, the latency and peak signal to noise ratio (PSNR) are considered in the objective function. Then, an efficient optimization algorithm is proposed under the block coordinate descent framework. Simulation results are provided to show the convergence behavior of the proposed algorithm, and the performance gains from the proposed FA-assisted LMembedded network over the other benchmark networks in terms of network latency and PSNR.

eess.SP

Zero-energy band observation in an interfacial chalcogen-organic network

Structurally-defined molecule-based lattices such as covalent organic or metal-organic networks on substrates, have emerged as highly tunable, modular platforms for two-dimensional band structure engineering. The ability to grow molecule-based lattices on diverse platforms, such as metal dichalcogenides, would further enable band structure tuning and alignment to the Fermi level, which is crucial for the exploration and design of quantum matter. In this work, we study the emergence of a zero-energy band in a triarylamine-based network on semiconducting 1T-TiSe2 at low temperatures, by means of scanning probe microscopy and photoemission spectroscopy, together with density-functional theory. Hybridization between the position-selective nitrogens and selenium p-states results in CN-Se interfacial coordination motifs, leading to a hybrid molecule-semiconductor band at the Fermi level. Our findings introduce chalcogen-organic networks and showcase an approach for the engineering of organic-inorganic quantum matter.

physics.chem-ph

AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views

We introduce AnySplat, a feed forward network for novel view synthesis from uncalibrated image collections. In contrast to traditional neural rendering pipelines that demand known camera poses and per scene optimization, or recent feed forward methods that buckle under the computational weight of dense views, our model predicts everything in one shot. A single forward pass yields a set of 3D Gaussian primitives encoding both scene geometry and appearance, and the corresponding camera intrinsics and extrinsics for each input image. This unified design scales effortlessly to casually captured, multi view datasets without any pose annotations. In extensive zero shot evaluations, AnySplat matches the quality of pose aware baselines in both sparse and dense view scenarios while surpassing existing pose free approaches. Moreover, it greatly reduce rendering latency compared to optimization based neural fields, bringing real time novel view synthesis within reach for unconstrained capture settings.Project page: https://city-super.github.io/anysplat/

cs.CV

Large-Scale Bayesian Tensor Reconstruction via Approximate Message Passing

While CANDECOMP/PARAFAC (CP) decomposition (CPD) is fundamental for tensor reconstruction, Bayesian CPD often scales poorly because variational updates require repeated matrix inversions. We develop CP generalized approximate message passing (CP-GAMP) for incomplete noisy Bayesian CPD. The algorithm uses Gaussian message approximations to avoid high-dimensional inversions, and it combines a Bernoulli-Gaussian prior with expectation-maximization updates to estimate effective CP rank and noise variance. We also give a formal state evolution (SE) recursion and relate its fixed points to replica-symmetric saddle points, so CP-GAMP's SE-predicted error can be compared with the formal replica-symmetric minimum mean-squared error (MMSE) benchmark in the matched limit. Synthetic and image-inpainting experiments show that CP-GAMP substantially reduces runtime relative to variational Bayesian CPD while maintaining competitive reconstruction accuracy.

cs.LG

Deep Unfolding with Kernel-based Quantization in MIMO Detection

The development of edge computing places critical demands on energy-efficient model deployment for multiple-input multiple-output (MIMO) detection tasks. Deploying deep unfolding models such as PGD-Nets and ADMM-Nets into resource-constrained edge devices using quantization methods is challenging. Existing quantization methods based on quantization aware training (QAT) suffer from performance degradation due to their reliance on parametric distribution assumption of activations and static quantization step sizes. To address these challenges, this paper proposes a novel kernel-based adaptive quantization (KAQ) framework for deep unfolding networks. By utilizing a joint kernel density estimation (KDE) and maximum mean discrepancy (MMD) approach to align activation distributions between full-precision and quantized models, the need for prior distribution assumptions is eliminated. Additionally, a dynamic step size updating method is introduced to adjust the quantization step size based on the channel conditions of wireless networks. Extensive simulations demonstrate that the accuracy of proposed KAQ framework outperforms traditional methods and successfully reduces the model's inference latency.

cs.LG

iMacHSR: Intermediate Multi-Access Heterogeneous Supervision and Regularization Scheme Toward Architecture-Agnostic Training

While deep supervision is a powerful training strategy by supervising intermediate layers with auxiliary losses, it faces three underexplored problems: (I) Existing deep supervision techniques are generally bond with specific model architectures strictly, lacking generality. (II) The identical loss function for intermediate and output layers causes intermediate layers to prioritize output-specific features prematurely, limiting generalizable representations. (III) Lacking regularization on hidden activations risks overconfident predictions, reducing generalization to unseen scenarios. To tackle these challenges, we propose an architecture-agnostic, intermediate Multi-access Heterogeneous Supervision and Regularization (iMacHSR) scheme. Specifically, the proposed iMacHSR introduces below integral strategies: (I) we select multiple intermediate layers based on predefined architecture-agnostic standards; (II) loss functions (different from output-layer loss) are applied to those selected intermediate layers, which can guide intermediate layers to learn diverse and hierarchical representations; and (III) negative entropy regularization on selected layers' hidden features discourages overconfident predictions and mitigates overfitting. These intermediate terms are combined into the output-layer training loss to form a unified optimization objective, enabling comprehensive optimization across the network hierarchy. We then take the semantic understanding task as an example to assess iMacHSR and apply iMacHSR to several model architectures. Extensive experiments on multiple datasets demonstrate that iMacHSR outperforms conventional output-layer single-point supervision method up to 9.19% in mIoU.

cs.RO

A General Optimization Framework for Tackling Distance Constraints in Movable Antenna-Aided Systems

The recently emerged movable antenna (MA) shows great promise in leveraging spatial degrees of freedom to enhance the performance of wireless systems. However, resource allocation in MA-aided systems faces challenges due to the nonconvex and coupled constraints on antenna positions. This paper systematically reveals the challenges posed by the minimum antenna separation distance constraints. Furthermore, we propose a penalty optimization framework for resource allocation under such new constraints for MA-aided systems. Specifically, the proposed framework separates the non-convex and coupled antenna distance constraints from the movable region constraints by introducing auxiliary variables. Subsequently, the resulting problem is efficiently solved by alternating optimization, where the optimization of the original variables resembles that in conventional resource allocation problem while the optimization with respect to the auxiliary variables is achieved in closedform solutions. To illustrate the effectiveness of the proposed framework, we present three case studies: capacity maximization, latency minimization, and regularized zero-forcing precoding. Simulation results demonstrate that the proposed optimization framework consistently outperforms state-of-the-art schemes.

eess.SP

Handling Distance Constraint in Movable Antenna Aided Systems: A General Optimization Framework

The movable antenna (MA) is a promising technology to exploit more spatial degrees of freedom for enhancing wireless system performance. However, the MA-aided system introduces the non-convex antenna distance constraints, which poses challenges in the underlying optimization problems. To fill this gap, this paper proposes a general framework for optimizing the MA-aided system under the antenna distance constraints. Specifically, we separate the non-convex antenna distance constraints from the objective function by introducing auxiliary variables. Then, the resulting problem can be efficiently solved under the alternating optimization framework. For the subproblems with respect to the antenna position variables and auxiliary variables, the proposed algorithms are able to obtain at least stationary points without any approximations. To verify the effectiveness of the proposed optimization framework, we present two case studies: capacity maximization and regularized zero-forcing precoding. Simulation results demonstrate the proposed optimization framework outperforms the existing baseline schemes under both cases.

eess.SP

Electronic Correlations in Multiferroic van der Waals CuCrP$_2$S6: Insights From X-Ray Spectroscopy and DFT

The electronic structure of high-quality van der Waals multiferroic CuCrP$_2$S6 crystals was investigated applying photoelectron spectroscopy methods in combination with DFT analysis. Using X-ray photoelectron and near-edge X-ray absorption fine structure (NEXAFS) spectroscopy at the Cu L2,3 and Cr L2,3 absorption edges we determine the charge states of ions in the studied compound. Analyzing the systematic NEXAFS and resonant photoelectron spectroscopy data at the Cu/Cr L2,3 absorption edges allowed us to assign the CuCrP$_2$S6 material to a Mott-Hubbard type insulator and identify different Auger-decay channels (participator vs. spectator) during absorption and autoionization processes. Spectroscopic and theoretical data obtained for CuCrP$_2$S6 are very important for the detailed understanding of the electronic structure and electron-correlations phenomena in different layered materials, that will drive their further applications in different areas, like electronics, spintronics, sensing, and catalysis.

cond-mat.mtrl-sci

Realization of the electric-field driven "one-material"-based magnetic tunnel junction using van der Waals antiferromagnetic MnPX3 (X: S, Se)

Presently a lot of efforts are devoted to the investigation of new two-dimensional magnetic materials, which are considered as promising for the realization of the future electronics and spintronics devices. However, the utilization of these materials in different junctions requires complicated processing that in many cases leads to unwanted parasitic effects influencing the performance of the junctions. Here, we propose the new elegant approach for the realization of the "one-material"-based magnetic tunnel junction. The several layers of 2D van der Waals MnPX3 (X: S, Se), which is insulating antiferromagnet in its ground state, are used and the effect of the applied external electric filed leads to the half-metallic ferromagnetic states for the outermost layers of the MnPX3 stack. The rich states diagram of such magnetic tunnel junction permits to precisely control its tunneling conductivity. The realized "one-material"-based magnetic tunnel junction allows to avoid all effects connected with the lattice mismatches and carriers scattering effects at the materials interfaces, giving high perspectives for the application of such systems in electronics and spintronics.

cond-mat.mtrl-sci

Mott-Hubbard Insulating State for the Layered van der Waals FePX$_3$ (X:S, Se) As Revealed by NEXAFS and Resonant Photoelectron Spectroscopy

A broad family of the nowadays studied low-dimensional systems, including 2D materials, demonstrate many fascinating properties, which however depend on the atomic composition as well as on the system dimensionality. Therefore, the studies of the electronic correlation effects in the new 2D materials is of paramount importance for the understanding of their transport, optical and catalytic properties. Here, by means of electron spectroscopy methods in combination with density functional theory calculations we investigate the electronic structure of a new layered van der Waals FePX$_3$ (X: S, Se) materials. Using systematic resonant photoelectron spectroscopy studies we observed strong resonant behavior for the peaks associated with the $3d^{n-1}$ final state at low binding energies for these materials. Such observations clearly assign FePX$_3$ to the class of Mott-Hubbard type insulators for which the top of the valence band is formed by the hybrid Fe-S/Se electronic states. These observations are important for the deep understanding of this new class of materials and draw perspectives for their further applications in different application areas, like (opto)spintronics and catalysis.

cond-mat.mtrl-sci

Correlations in the Electronic Structure of van der Waals NiPS$_3$ Crystals: An X-Ray Absorption and Resonant Photoelectron Spectroscopy Study

The electronic structure of high-quality van der Waals NiPS$_3$ crystals was studied using near-edge x-ray absorption spectroscopy (NEXAFS) and resonant photoelectron spectroscopy (ResPES) in combination with density functional theory (DFT) approach. The experimental spectroscopic methods, being element specific, allow to discriminate between atomic contributions in the valence and conduction band density of states and give direct comparison with the results of DFT calculations. Analysis of the NEXAFS and ResPES data allows to identify the NiPS$_3$ material as a charge-transfer insulator. Obtained spectroscopic and theoretical data are very important for the consideration of possible correlated-electron phenomena in such transition-metal layered materials, where the interplay between different degrees of freedom for electrons defines their electronic properties, allowing to understand their optical and transport properties and to propose further possible applications in electronics, spintronics and catalysis.

cond-mat.mtrl-sci

Topological Quasi-2D Semimetal Co$_3$Sn$_2$S$_2$: Insights To Electronic Structure From NEXAFS and Resonant Photoelectron Spectroscopy

The electronic structure of the natural topological semimetal Co$_3$Sn$_2$S$_2$ crystals was studied using near-edge x-ray absorption spectroscopy (NEXAFS) and resonant photoelectron spectroscopy (ResPES). Although, the significant increase of the Co\,$3d$ valence band emission is observed at the Co\,$2p$ absorption edge in the ResPES experiments, the spectral weight at these photon energies is dominated by the normal Auger contribution. This observation indicates the delocalised character of photoexcited Co\,$3d$ electrons and is supported by the first-principle calculations. Our results on the investigations of the element- and orbital-specific electronic states near the Fermi level of Co$_3$Sn$_2$S$_2$ are of importance for the comprehensive description of the electronic structure of this materials, which is significant for future applications of this material in different areas of science and technology, including catalysis and water splitting applications.

cond-mat.mtrl-sci