Searcharxiv⌕ Search

arXiv subjects

Kun Han

Publications and source records attributed to Kun Han.

At least 37 records · Page 2Linked to original sources

CVTHead: One-shot Controllable Head Avatar with Vertex-feature Transformer

Reconstructing personalized animatable head avatars has significant implications in the fields of AR/VR. Existing methods for achieving explicit face control of 3D Morphable Models (3DMM) typically rely on multi-view images or videos of a single subject, making the reconstruction process complex. Additionally, the traditional rendering pipeline is time-consuming, limiting real-time animation possibilities. In this paper, we introduce CVTHead, a novel approach that generates controllable neural head avatars from a single reference image using point-based neural rendering. CVTHead considers the sparse vertices of mesh as the point set and employs the proposed Vertex-feature Transformer to learn local feature descriptors for each vertex. This enables the modeling of long-range dependencies among all the vertices. Experimental results on the VoxCeleb dataset demonstrate that CVTHead achieves comparable performance to state-of-the-art graphics-based methods. Moreover, it enables efficient rendering of novel human heads with various expressions, head poses, and camera views. These attributes can be explicitly controlled using the coefficients of 3DMMs, facilitating versatile and realistic animation in real-time scenarios.

cs.CV↗

Self-Sampling Meta SAM: Enhancing Few-shot Medical Image Segmentation with Meta-Learning

While the Segment Anything Model (SAM) excels in semantic segmentation for general-purpose images, its performance significantly deteriorates when applied to medical images, primarily attributable to insufficient representation of medical images in its training dataset. Nonetheless, gathering comprehensive datasets and training models that are universally applicable is particularly challenging due to the long-tail problem common in medical images. To address this gap, here we present a Self-Sampling Meta SAM (SSM-SAM) framework for few-shot medical image segmentation. Our innovation lies in the design of three key modules: 1) An online fast gradient descent optimizer, further optimized by a meta-learner, which ensures swift and robust adaptation to new tasks. 2) A Self-Sampling module designed to provide well-aligned visual prompts for improved attention allocation; and 3) A robust attention-based decoder specifically designed for medical few-shot learning to capture relationship between different slices. Extensive experiments on a popular abdominal CT dataset and an MRI dataset demonstrate that the proposed method achieves significant improvements over state-of-the-art methods in few-shot segmentation, with an average improvements of 10.21% and 1.80% in terms of DSC, respectively. In conclusion, we present a novel approach for rapid online adaptation in interactive image segmentation, adapting to a new organ in just 0.83 minutes. Code is publicly available on GitHub upon acceptance.

cs.CV↗

Multifunctional magnetic oxide-MoS$_2$ heterostructures on silicon

Correlated oxides and related heterostructures are intriguing for developing future multifunctional devices by exploiting their exotic properties, but their integration with other materials, especially on Si-based platforms, is challenging. Here, van der Waals heterostructures of La$_{0.7}$Sr$_{0.3}$MnO$_3$ (LSMO), a correlated manganite perovskite, and MoS$_2$ are demonstrated on Si substrates with multiple functions. To overcome the problems due to the incompatible growth process, technologies involving freestanding LSMO membranes and van der Waals force-mediated transfer are used to fabricate the LSMO-MoS$_2$ heterostructures. The LSMO-MoS$_2$ heterostructures exhibit a gate-tunable rectifying behavior, based on which metal-semiconductor field-effect transistors (MESFETs) with on-off ratios of over 104 can be achieved. The LSMO-MoS$_2$ heterostructures can function as photodiodes displaying considerable open-circuit voltages and photocurrents. In addition, the colossal magnetoresistance of LSMO endows the LSMO-MoS$_2$ heterostructures with an electrically tunable magnetoresponse at room temperature. This work not only proves the applicability of the LSMO-MoS$_2$ heterostructure devices on Si-based platform but also demonstrates a paradigm to create multifunctional heterostructures from materials with disparate properties.

cond-mat.str-el↗

Competition of electronic correlation and reconstruction in La1-xSrxTiO3/SrTiO3 heterostructures

Electronic correlation and reconstruction are two important factors that play a critical role in shaping the magnetic and electronic properties of correlated low-dimensional systems. Here, we report a competition between the electronic correlation and structural reconstruction in La1-xSrxTiO3/SrTiO3 heterostructures by modulating material polarity and interfacial strain, respectively. The heterostructures exhibit a critical thickness (tc) at which a metal-to-insulator transition (MIT) abruptly occurs at certain thickness, accompanied by the coexistence of two- and three-dimensional (2D and 3D) carriers. Intriguingly, the tc exhibits a V-shaped dependence on the doping concentration of Sr, with the smallest tc value at x = 0.5. We attribute this V-shaped dependence to the competition between the electronic reconstruction (modulated by the polarity) and the electronic correlation (modulated by strain), which are borne out by the experimental results, including strain-dependent electronic properties and the evolution of 2D and 3D carriers. Our findings underscore the significance of the interplay between electronic reconstruction and correlation in the realization and utilization of emergent electronic functionalities in low-dimensional correlated systems.

cond-mat.str-el↗

Hybrid-CSR: Coupling Explicit and Implicit Shape Representation for Cortical Surface Reconstruction

We present Hybrid-CSR, a geometric deep-learning model that combines explicit and implicit shape representations for cortical surface reconstruction. Specifically, Hybrid-CSR begins with explicit deformations of template meshes to obtain coarsely reconstructed cortical surfaces, based on which the oriented point clouds are estimated for the subsequent differentiable poisson surface reconstruction. By doing so, our method unifies explicit (oriented point clouds) and implicit (indicator function) cortical surface reconstruction. Compared to explicit representation-based methods, our hybrid approach is more friendly to capture detailed structures, and when compared with implicit representation-based methods, our method can be topology aware because of end-to-end training with a mesh-based deformation module. In order to address topology defects, we propose a new topology correction pipeline that relies on optimization-based diffeomorphic surface registration. Experimental results on three brain datasets show that our approach surpasses existing implicit and explicit cortical surface reconstruction methods in numeric metrics in terms of accuracy, regularity, and consistency.

cs.CV↗

MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation

Acquiring and annotating sufficient labeled data is crucial in developing accurate and robust learning-based models, but obtaining such data can be challenging in many medical image segmentation tasks. One promising solution is to synthesize realistic data with ground-truth mask annotations. However, no prior studies have explored generating complete 3D volumetric images with masks. In this paper, we present MedGen3D, a deep generative framework that can generate paired 3D medical images and masks. First, we represent the 3D medical data as 2D sequences and propose the Multi-Condition Diffusion Probabilistic Model (MC-DPM) to generate multi-label mask sequences adhering to anatomical geometry. Then, we use an image sequence generator and semantic diffusion refiner conditioned on the generated mask sequences to produce realistic 3D medical images that align with the generated masks. Our proposed framework guarantees accurate alignment between synthetic images and segmentation maps. Experiments on 3D thoracic CT and brain MRI datasets show that our synthetic data is both diverse and faithful to the original data, and demonstrate the benefits for downstream segmentation tasks. We anticipate that MedGen3D's ability to synthesize paired 3D medical images and masks will prove valuable in training deep learning models for medical imaging tasks.

eess.IV↗

Hybrid Neural Diffeomorphic Flow for Shape Representation and Generation via Triplane

Deep Implicit Functions (DIFs) have gained popularity in 3D computer vision due to their compactness and continuous representation capabilities. However, addressing dense correspondences and semantic relationships across DIF-encoded shapes remains a critical challenge, limiting their applications in texture transfer and shape analysis. Moreover, recent endeavors in 3D shape generation using DIFs often neglect correspondence and topology preservation. This paper presents HNDF (Hybrid Neural Diffeomorphic Flow), a method that implicitly learns the underlying representation and decomposes intricate dense correspondences into explicitly axis-aligned triplane features. To avoid suboptimal representations trapped in local minima, we propose hybrid supervision that captures both local and global correspondences. Unlike conventional approaches that directly generate new 3D shapes, we further explore the idea of shape generation with deformed template shape via diffeomorphic flows, where the deformation is encoded by the generated triplane features. Leveraging a pre-existing 2D diffusion model, we produce high-quality and diverse 3D diffeomorphic flows through generated triplanes features, ensuring topological consistency with the template shape. Extensive experiments on medical image organ segmentation datasets evaluate the effectiveness of HNDF in 3D shape representation and generation.

cs.CV↗

Localized Region Contrast for Enhancing Self-Supervised Learning in Medical Image Segmentation

Recent advancements in self-supervised learning have demonstrated that effective visual representations can be learned from unlabeled images. This has led to increased interest in applying self-supervised learning to the medical domain, where unlabeled images are abundant and labeled images are difficult to obtain. However, most self-supervised learning approaches are modeled as image level discriminative or generative proxy tasks, which may not capture the finer level representations necessary for dense prediction tasks like multi-organ segmentation. In this paper, we propose a novel contrastive learning framework that integrates Localized Region Contrast (LRC) to enhance existing self-supervised pre-training methods for medical image segmentation. Our approach involves identifying Super-pixels by Felzenszwalb's algorithm and performing local contrastive learning using a novel contrastive sampling loss. Through extensive experiments on three multi-organ segmentation datasets, we demonstrate that integrating LRC to an existing self-supervised method in a limited annotation setting significantly improves segmentation performance. Moreover, we show that LRC can also be applied to fully-supervised pre-training methods to further boost performance.

cs.CV↗

Diffeomorphic Image Registration with Neural Velocity Field

Diffeomorphic image registration, offering smooth transformation and topology preservation, is required in many medical image analysis tasks.Traditional methods impose certain modeling constraints on the space of admissible transformations and use optimization to find the optimal transformation between two images. Specifying the right space of admissible transformations is challenging: the registration quality can be poor if the space is too restrictive, while the optimization can be hard to solve if the space is too general. Recent learning-based methods, utilizing deep neural networks to learn the transformation directly, achieve fast inference, but face challenges in accuracy due to the difficulties in capturing the small local deformations and generalization ability. Here we propose a new optimization-based method named DNVF (Diffeomorphic Image Registration with Neural Velocity Field) which utilizes deep neural network to model the space of admissible transformations. A multilayer perceptron (MLP) with sinusoidal activation function is used to represent the continuous velocity field and assigns a velocity vector to every point in space, providing the flexibility of modeling complex deformations as well as the convenience of optimization. Moreover, we propose a cascaded image registration framework (Cas-DNVF) by combining the benefits of both optimization and learning based methods, where a fully convolutional neural network (FCN) is trained to predict the initial deformation, followed by DNVF for further refinement. Experiments on two large-scale 3D MR brain scan datasets demonstrate that our proposed methods significantly outperform the state-of-the-art registration methods.

cs.CV↗

Identity-Aware Hand Mesh Estimation and Personalization from RGB Images

Reconstructing 3D hand meshes from monocular RGB images has attracted increasing amount of attention due to its enormous potential applications in the field of AR/VR. Most state-of-the-art methods attempt to tackle this task in an anonymous manner. Specifically, the identity of the subject is ignored even though it is practically available in real applications where the user is unchanged in a continuous recording session. In this paper, we propose an identity-aware hand mesh estimation model, which can incorporate the identity information represented by the intrinsic shape parameters of the subject. We demonstrate the importance of the identity information by comparing the proposed identity-aware model to a baseline which treats subject anonymously. Furthermore, to handle the use case where the test subject is unseen, we propose a novel personalization pipeline to calibrate the intrinsic shape parameters using only a few unlabeled RGB images of the subject. Experiments on two large scale public datasets validate the state-of-the-art performance of our proposed method.

cs.CV↗

Critical behavior in the Mn$_{5}$Ge$_{3}$ ferromagnet

High-Curie-temperature ferromagnets are promising candidates for designing new spintronic devices. Here we have successfully synthesized a single-crystal sample of the itinerant ferromagnet Mn$ _{5}$Ge$_{3}$ used flux method and its critical properties were investigated by means of bulk dc-magnetization at the boundary between the ferromagnetic (FM) and paramagnetic (PM) phase. Critical exponents $ β=0.336 \pm 0.001 $ with a critical temperature $ T_{c}=300.29 \pm 0.01 $ K and $ γ=1.193 \pm 0.003 $ with $ T_{c} = 300.15 \pm 0.05 $ K are obtained by the modified Arrott plot, whereas $ δ= 4.61 \pm 0.03 $ is deduced by a critical isotherm analysis at $ T_{c} = 300 $ K. The self-consistency and reliability of these critical exponents are verified by the Widom scaling law and the scaling equations. Further analysis reveals that the spin coupling in Mn$ _{5}$Ge$_{3}$ exhibits three-dimensional Ising-like behavior. The magnetic exchange is found to decay as $ J(r)\approx r^{-4.855} $ and the spin interactions are extended beyond the nearest neighbors, which may be related to different set of Mn--Mn interactions with unequal magnitude of exchange strengths. Additionally, the existence of noncollinear spin configurations in Mn$ _{5} $Ge$ _{3} $ results in a small deviation of obtained critical exponents from those for standard 3D-Ising model.

cond-mat.mtrl-sci↗

Large anomalous Hall effect in layered antiferromagnet Co$_{0.29}$TaS$_2$

We present a study on the magnetization, anomalous Hall effect (AHE) and novel longitudinal resistivity in layered antiferromagnet Co$_{0.29}$TaS$_{2}$. Of particular interests in Co$_{0.29}$TaS$_{2}$ are abundant magnetic transitions, which show that the magnetic structures are tuned by temperature or magnetic field. With decreasing temperature, Co$_{0.29}$TaS$_{2}$ undergoes two transitions at T$_{t1}\sim$ 38.3 K and T$_{t2}\sim$ 24.3 K. Once the magnetic field is applied, another transition T$_{t3}\sim$ 34.3 K appears between 0.3 T and 5 T. At 2 K, an obvious ferromagnetic hysteresis loop within H$_{t1}\sim\pm$ 6.9 T is observed, which decreases with increasing temperature and eventually disappears at T$_{t2}$. Besides, Co$_{0.29}$TaS$_{2}$ displays step-like behavior as another magnetic transition around H$_{t2}\sim\pm$ 4 T, which exists until $\sim$ T$_{t1}$. These characteristic temperatures and magnetic fields mark complex magnetic phase transitions in Co$_{0.29}$TaS$_{2}$, which are also evidenced in transport results. Large AHE dominates in the Hall resistivity with the conspicuous value of R$_{s}$/R$_{0}\sim 10^{5}$, considering that the tiny net magnetization (0.0094$μ_{B}$/Co) alone would not lead to this value, thus the contribution of Berry curvature is necessary. The longitudinal resistivity illustrates a prominent irreversible behavior within H$_{t1}$. The abrupt change at H$_{t2}$ below T$_{t1}$, corresponding to the step-like magnetic transitions, is also observed. Synergy between the magnetism and topological properties, both playing a crucial role, may be the key factor of large AHE in antiferromagnet, which also offers a new perspective in magnetic topological materials with the platform of Co$_{0.29}$TaS$_{2}$.

cond-mat.mtrl-sci↗

Quantum oscillations and weak anisotropic resistivity in the chiral Fermion semimetal PdGa

We perform a detailed analysis of the magnetotransport and de Haas-van Alphen (dHvA) oscillations in crystal PdGa which is predicted to be a typical chiral Fermion semimetal from CoSi family holding a large Chern number. The unsaturated quadratic magnetoresistance (MR) and nonlinear Hall resistivity indicate that PdGa is a multi-band system without electron-hole compensation. Angle-dependent resistivity in PdGa shows weak anisotropy with twofold or threefold symmetry when the magnetic field rotates within the (1$\bar{1}$0) or (111) plane perpendicular to the current. Nine or three frequencies are extracted after the fast Fourier-transform analysis (FFT) of the dHvA oscillations with B//[001] or B//[011], respectively, which is confirmed to be consistent with the Fermi surfaces (FSs) obtained from first-principles calculations with spin-orbit coupling (SOC) considered.

cond-mat.mtrl-sci↗

Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic Flow

Deep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shapes of the same class by learning a DIFs-modeled shape template. These methods provide great flexibility and accuracy in reconstructing 3D shapes and inferring correspondences. However, the point correspondences built from these methods do not intrinsically preserve the topology of the shapes, unlike mesh-based template matching methods. This limits their applications on 3D geometries where underlying topological structures exist and matter, such as anatomical structures in medical images. In this paper, we propose a new model called Neural Diffeomorphic Flow (NDF) to learn deep implicit shape templates, representing shapes as conditional diffeomorphic deformations of templates, intrinsically preserving shape topologies. The diffeomorphic deformation is realized by an auto-decoder consisting of Neural Ordinary Differential Equation (NODE) blocks that progressively map shapes to implicit templates. We conduct extensive experiments on several medical image organ segmentation datasets to evaluate the effectiveness of NDF on reconstructing and aligning shapes. NDF achieves consistently state-of-the-art organ shape reconstruction and registration results in both accuracy and quality. The source code is publicly available at https://github.com/Siwensun/Neural_Diffeomorphic_Flow--NDF.

cs.CV↗

Tailoring magnetic order via atomically stacking 3d/5d electrons

The ability to tune magnetic orders, such as magnetic anisotropy and topological spin texture, is desired in order to achieve high-performance spintronic devices. A recent strategy has been to employ interfacial engineering techniques, such as the introduction of spin-correlated interfacial coupling, to tailor magnetic orders and achieve novel magnetic properties. We chose a unique polar-nonpolar LaMnO3/SrIrO3 superlattice because Mn (3d)/Ir (5d) oxides exhibit rich magnetic behaviors and strong spin-orbit coupling through the entanglement of their 3d and 5d electrons. Through magnetization and magnetotransport measurements, we found that the magnetic order is interface-dominated as the superlattice period is decreased. We were able to then effectively modify the magnetization, tilt of the ferromagnetic easy axis, and symmetry transition of the anisotropic magnetoresistance of the LaMnO3/SrIrO3 superlattice by introducing additional Mn (3d) and Ir (5d) interfaces. Further investigations using in-depth first-principles calculations and numerical simulations revealed that these magnetic behaviors could be understood by the 3d/5d electron correlation and Rashba spin-orbit coupling. The results reported here demonstrate a new route to synchronously engineer magnetic properties through the atomic stacking of different electrons, contributing to future applications.

cond-mat.str-el↗

Reversible modulation of metal-insulator transition in VO2 via chemically-induced oxygen migration

Metal-insulator transitions (MIT),an intriguing correlated phenomenon induced by the subtle competition of the electrons' repulsive Coulomb interaction and kinetic energy, is of great potential use for electronic applications due to the dramatic change in resistivity. Here, we demonstrate a reversible control of MIT in VO2 films via oxygen stoichiometry engineering. By facilely depositing and dissolving a water-soluble yet oxygen-active Sr3Al2O6 capping layer atop the VO2 at room temperature, oxygen ions can reversibly migrate between VO2 and Sr3Al2O6, resulting in a gradual suppression and a complete recovery of MIT in VO2. The migration of the oxygen ions is evidenced in a combination of transport measurement, structural characterization and first-principles calculations. This approach of chemically-induced oxygen migration using a water-dissolvable adjacent layer could be useful for advanced electronic and iontronic devices and studying oxygen stoichiometry effects on the MIT.

cond-mat.mtrl-sci↗

Enhanced metal-insulator transition in freestanding VO2 down to 5 nm thickness

Ultrathin freestanding membranes with a pronounced metal-insulator transition (MIT) provides huge potential in future flexible electronic applications as well as a unique aspect of the study of lattice-electron interplay. However, the reduction of the thickness to an ultrathin region (a few nm) is typically detrimental to the MIT in epitaxial films, and even catastrophic for their freestanding form. Here, we report an enhanced MIT in VO2-based freestanding membranes, with a lateral size up to millimetres and VO2 thickness down to 5 nm. The VO2-membranes were detached by dissolving a Sr3Al2O6 sacrificial layer between the VO2 thin film and c-Al2O3(0001) substrate, allowing a transfer onto arbitrary surfaces. Furthermore, the MIT in the VO2-membrane was greatly enhanced by inserting an intermediate Al2O3 buffer layer. In comparison to the best available ultrathin VO2-membranes, the enhancement of MIT is over 400% at 5 nm VO2 thickness and more than one order of magnitude for VO2 above 10 nm. Our study widens the spectrum of functionality in ultrathin and large-scale membranes, and enables the potential integration of MIT into flexible electronics and photonics.

cond-mat.mtrl-sci↗

DiDiSpeech: A Large Scale Mandarin Speech Corpus

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech processing tasks, such as voice conversion, multi-speaker text-to-speech and automatic speech recognition. We conduct experiments with multiple speech tasks and evaluate the performance, showing that it is promising to use the corpus for both academic research and practical application. The corpus is available at https://outreach.didichuxing.com/research/opendata/.

eess.AS↗