SearcharxivSearch

arXiv subjects

Jingxuan Kang

Publications and source records attributed to Jingxuan Kang.

15 recordsLinked to original sources

Two-step growth of (In,Ga)N pseudo-substrates on GaN templates by plasma-assisted molecular beam epitaxy

(In,Ga)N layers are grown by plasma-assisted molecular beam epitaxy on GaN templates. We introduce a two-step protocol that involves switching the growth conditions from initially N-stable to metal-stable. Reflection high-energy electron diffraction as well as scanning electron and atomic force microscopy reveal that the first step results in a rough intermediate surface with open pits, whereas the final surface is smooth. The narrow linewidth of the photoluminescence band indicates an excellent compositional homogeneity of the upper layer. Its in-plane lattice constant is determined to be $\approx$3.26 Åfrom X-ray diffraction measurements. This combination of favorable properties makes these layers attractive as pseudo-substrates for the growth of red-emitting (In,Ga)N light-emitting diodes. In particular, the approach presented here does not require any complex external processing and is, thus, scalable and economical.

cond-mat.mtrl-sci

Simultaneously monitoring Ga adsorption and desorption kinetics on GaN(0001) using four in situ techniques

We present a systematic investigation of Ga adsorption and desorption kinetics on the wurtzite GaN(0001) surface using four in situ techniques operated simultaneously: reflection high-energy electron diffraction, laser reflectometry, line-of-sight quadrupole mass spectrometry, and optical pyrometry. Flux- and temperature-dependent experiments are performed for Ga coverages ranging from the submonolayer to the droplet regime. Despite their distinct transient responses, the signals from all four techniques and their trends with surface coverage are quantitatively reproduced by a unified kinetic model of Ga adsorption, diffusion, and desorption. An Arrhenius analysis of the Ga adlayer desorption yields an activation energy of (2.87 $\pm$ 0.04) eV.

physics.app-ph

Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM

Segmentation is central to clinical diagnosis and monitoring, yet the reliability of modern foundation models in medical imaging still depends on the availability of precise prompts. The Segment Anything Model (SAM) offers powerful zero-shot capabilities, although it collapses under the weak, generic, and noisy prompts that dominate real clinical workflows. In practice, annotations such as centerline points are coarse and ambiguous, often drifting across neighboring anatomy and misguiding SAM toward inconsistent or incomplete masks. We introduce SPD, a Saliency-Guided Prompt Distillation framework that converts these unreliable cues into robust guidance. SPD first learns data-driven anatomical priors through a lightweight saliency head to obtain confident localization maps. These priors then drive Contextual Prompt Distillation, which validates and enriches noisy prompts using cues from anatomically adjacent slices, producing a consensus prompt set that matches the behavior of expert reasoning. A Pairwise Slice Consistency objective further enforces local anatomical coherence during segmentation. Experiments on four challenging MRI and CT benchmarks demonstrate that SPD consistently outperforms existing SAM adaptations and supervised baselines, delivering large gains in both region-based and boundary-based metrics. SPD provides a practical and principled path toward reliable foundation model deployment in clinical environments where only imperfect prompts are available.

cs.CV

Fabrication of (In,Ga)N pseudo-substrates by a three-step growth protocol without ex-situ processing

We fabricate (In,Ga)N pseudo-substrates with a total thickness of ~1 um grown on GaN templates using plasma-assisted molecular beam epitaxy. In a three-step process, we change growth conditions from N-rich to metal-rich in order to sequentially form a roughened GaN layer, relaxed (In,Ga)N nanostructures, and a coalesced, smooth (In,Ga)N layer. Samples are analyzed by scanning electron and atomic force microscopy, X-ray diffraction, as well as photo- and cathodoluminescence spectroscopy. Compared to a reference layer grown directly on GaN, the pseudo-substrate exhibits a higher In content (~0.3), strain relaxation degree (~80%), narrower photoluminescence linewidth, and larger area fraction of bright regions in cathodoluminescence maps, showing the benefits of the three-step growth protocol. This straightforward approach does not necessitate any ex-situ processing and could enable the scalable fabrication of (In,Ga)N pseudo-substrates for high-efficiency red-emitting (In,Ga)N devices.

cond-mat.mtrl-sci

Combining metal dewetting and lateral etching for the scalable top-down fabrication of GaN nanowire arrays with independently tunable diameter and spacing

The top-down fabrication of nanowires based on patterning via metal dewetting is a cost-effective and scalable approach that is particularly suited for applications requiring large arrays of nanowires. Advantageously, the nanowire diameter can be tailored by the initial metal film thickness. However, we show here that metal dewetting inherently leads to a coupling between the nanowire diameter and spacing. To overcome this limitation, we introduce two strategies that are exemplified for GaN nanowires: (i) modification of the surface and interface energies within the dewetting system, and (ii) thinning of the nanowires by lateral etching. In the first strategy, GaN(0001), SiOx, and SiNx substrate surfaces are combined with Au, Pt, and Pt-Au alloy dewetting metals to tune the dewetting behavior. The differences in interface energies affect the relation between nanowire diameter and spacing, albeit within a limited range. The second strategy adds a lateral etching step to the conventional top-down nanowire fabrication process. This step at the same time reduces the nanowire diameter and increases the spacing, thus enabling combinations beyond the constraints of metal dewetting alone. When in addition different initial nanowire diameters are employed, it is possible to independently control diameter and spacing over a substantially extended range. Therefore, the inherent limitation of conventional dewetting-based patterning approaches for the top-down fabrication of nanowires is overcome.

physics.app-ph

From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection

Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language models (LLMs), significantly enhancing zero-shot performance by incorporating multi-view information. However, the inherent randomness of these augmentations can inevitably introduce background artifacts and cause models to overly focus on local details, compromising global semantic understanding. To address these issues, we propose an \textbf{A}ttention-\textbf{B}ased \textbf{S}election (\textbf{ABS}) method from local details to global context, which applies attention-guided cropping in both raw images and feature space, supplement global semantic information through strategic feature selection. Additionally, we introduce a soft matching technique to effectively filter LLM descriptions for better alignment. \textbf{ABS} achieves state-of-the-art performance on out-of-distribution generalization and zero-shot classification tasks. Notably, \textbf{ABS} is training-free and even rivals few-shot and test-time adaptation methods. Our code is available at \href{https://github.com/BIT-DA/ABS}{\textcolor{darkgreen}{https://github.com/BIT-DA/ABS}}.

cs.CV

Growth of compositionally uniform $\mathrm{In}_{x}\mathrm{Ga}_{1-x}\mathrm{N}$ layers with low relaxation degree on GaN by molecular beam epitaxy

500-nm-thick $\mathrm{In}_{x}\mathrm{Ga}_{1-x}\mathrm{N}$ layers with $x=$ 0.05-0.14 are grown using plasma-assisted molecular beam epitaxy, and their properties are assessed by a comprehensive analysis involving x-ray diffraction, secondary ion mass spectrometry, and cathodoluminescence as well as photoluminescence spectroscopy. We demonstrate low degrees of strain relaxation (10% for $x=0.12$), low threading dislocation densities ($\mathrm{1\times10^{9}\,cm^{-2}}$ for $x=0.12$), uniform composition both in the growth and lateral direction, and a narrow emission band. The unique sum of excellent materials properties make these layers an attractive basis for the top-down fabrication of ternary nanowires.

cond-mat.mtrl-sci

Learning Modality Knowledge Alignment for Cross-Modality Transfer

Cross-modality transfer aims to leverage large pretrained models to complete tasks that may not belong to the modality of pretraining data. Existing works achieve certain success in extending classical finetuning to cross-modal scenarios, yet we still lack understanding about the influence of modality gap on the transfer. In this work, a series of experiments focusing on the source representation quality during transfer are conducted, revealing the connection between larger modality gap and lesser knowledge reuse which means ineffective transfer. We then formalize the gap as the knowledge misalignment between modalities using conditional distribution P(Y|X). Towards this problem, we present Modality kNowledge Alignment (MoNA), a meta-learning approach that learns target data transformation to reduce the modality knowledge discrepancy ahead of the transfer. Experiments show that out method enables better reuse of source modality knowledge in cross-modality transfer, which leads to improvements upon existing finetuning methods.

cs.CV

Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scarcity of labeled data. In this paper, we propose an end-to-end method, PaRe, to enhance cross-modal fine-tuning, aiming to transfer a large-scale pretrained model to various target modalities. PaRe employs a gating mechanism to select key patches from both source and target data. Through a modality-agnostic Patch Replacement scheme, these patches are preserved and combined to construct data-rich intermediate modalities ranging from easy to hard. By gradually intermediate modality generation, we can not only effectively bridge the modality gap to enhance stability and transferability of cross-modal fine-tuning, but also address the challenge of limited data in the target modality by leveraging enriched intermediate modality data. Compared with hand-designed, general-purpose, task-specific, and state-of-the-art cross-modal fine-tuning approaches, PaRe demonstrates superior performance across three challenging benchmarks, encompassing more than ten modalities.

cs.CV

Uniform large-area surface patterning achieved by metal dewetting for the top-down fabrication of GaN nanowire ensembles

The dewetting of thin Pt films on different surfaces is investigated as a means to provide the patterning for the top-down fabrication of GaN nanowire ensembles. The transformation from a thin film to an ensemble of nanoislands upon annealing proceeds in good agreement with the void growth model. With increasing annealing duration, the size and shape uniformity of the nanoislands improves. This improvement speeds up for higher annealing temperature. After an optimum annealing duration, the size uniformity deteriorates due to the coalescence of neighboring islands. By changing the Pt film thickness, the nanoisland diameter and density can be quantitatively controlled in a way predicted by a simple thermodynamic model. We demonstrate the uniformity of the nanoisland ensembles for an area larger than 1 cm$^2$. GaN nanowires are fabricated by a sequence of dry and wet etching steps, and these nanowires inherit the diameters and density of the Pt nanoisland ensemble used as a mask. Our study achieves advancements in size uniformity and range of obtainable diameters compared to previous works. This simple, economical, and scalable approach to the top-down fabrication of nanowires is useful for applications requiring large and uniform nanowire ensembles with controllable dimensions.

cond-mat.mtrl-sci

Shape-Sensitive Loss for Catheter and Guidewire Segmentation

We introduce a shape-sensitive loss function for catheter and guidewire segmentation and utilize it in a vision transformer network to establish a new state-of-the-art result on a large-scale X-ray images dataset. We transform network-derived predictions and their corresponding ground truths into signed distance maps, thereby enabling any networks to concentrate on the essential boundaries rather than merely the overall contours. These SDMs are subjected to the vision transformer, efficiently producing high-dimensional feature vectors encapsulating critical image attributes. By computing the cosine similarity between these feature vectors, we gain a nuanced understanding of image similarity that goes beyond the limitations of traditional overlap-based measures. The advantages of our approach are manifold, ranging from scale and translation invariance to superior detection of subtle differences, thus ensuring precise localization and delineation of the medical instruments within the images. Comprehensive quantitative and qualitative analyses substantiate the significant enhancement in performance over existing baselines, demonstrating the promise held by our new shape-sensitive loss function for improving catheter and guidewire segmentation.

eess.IV

Autonomous Catheterization with Open-source Simulator and Expert Trajectory

Endovascular robots have been actively developed in both academia and industry. However, progress toward autonomous catheterization is often hampered by the widespread use of closed-source simulators and physical phantoms. Additionally, the acquisition of large-scale datasets for training machine learning algorithms with endovascular robots is usually infeasible due to expensive medical procedures. In this chapter, we introduce CathSim, the first open-source simulator for endovascular intervention to address these limitations. CathSim emphasizes real-time performance to enable rapid development and testing of learning algorithms. We validate CathSim against the real robot and show that our simulator can successfully mimic the behavior of the real robot. Based on CathSim, we develop a multimodal expert navigation network and demonstrate its effectiveness in downstream endovascular navigation tasks. The intensive experimental results suggest that CathSim has the potential to significantly accelerate research in the autonomous catheterization field. Our project is publicly available at https://github.com/airvlab/cathsim.

cs.RO

Language Semantic Graph Guided Data-Efficient Learning

Developing generalizable models that can effectively learn from limited data and with minimal reliance on human supervision is a significant objective within the machine learning community, particularly in the era of deep neural networks. Therefore, to achieve data-efficient learning, researchers typically explore approaches that can leverage more related or unlabeled data without necessitating additional manual labeling efforts, such as Semi-Supervised Learning (SSL), Transfer Learning (TL), and Data Augmentation (DA). SSL leverages unlabeled data in the training process, while TL enables the transfer of expertise from related data distributions. DA broadens the dataset by synthesizing new data from existing examples. However, the significance of additional knowledge contained within labels has been largely overlooked in research. In this paper, we propose a novel perspective on data efficiency that involves exploiting the semantic information contained in the labels of the available data. Specifically, we introduce a Language Semantic Graph (LSG) which is constructed from labels manifest as natural language descriptions. Upon this graph, an auxiliary graph neural network is trained to extract high-level semantic relations and then used to guide the training of the primary model, enabling more adequate utilization of label knowledge. Across image, video, and audio modalities, we utilize the LSG method in both TL and SSL scenarios and illustrate its versatility in significantly enhancing performance compared to other data-efficient learning approaches. Additionally, our in-depth analysis shows that the LSG method also expedites the training process.

cs.CV

Translating Simulation Images to X-ray Images via Multi-Scale Semantic Matching

Endovascular intervention training is increasingly being conducted in virtual simulators. However, transferring the experience from endovascular simulators to the real world remains an open problem. The key challenge is the virtual environments are usually not realistically simulated, especially the simulation images. In this paper, we propose a new method to translate simulation images from an endovascular simulator to X-ray images. Previous image-to-image translation methods often focus on visual effects and neglect structure information, which is critical for medical images. To address this gap, we propose a new method that utilizes multi-scale semantic matching. We apply self-domain semantic matching to ensure that the input image and the generated image have the same positional semantic relationships. We further apply cross-domain matching to eliminate the effects of different styles. The intensive experiment shows that our method generates realistic X-ray images and outperforms other state-of-the-art approaches by a large margin. We also collect a new large-scale dataset to serve as the new benchmark for this task. Our source code and dataset will be made publicly available.

eess.IV

Abnormal Staebler-Wronski effect of amorphous silicon

Great achievements in last five years, such as record-efficient amorphous/crystalline silicon heterojunction (SHJ) solar cells and cutting-edge perovskite/SHJ tandem solar cells, place hydrogenated amorphous silicon (a-Si:H) at the forefront of emerging photovoltaics. Due to the extremely low doping efficiency of trivalent boron (B) in amorphous tetravalent silicon, light harvesting of aforementioned devices are limited by their fill factors (FF), which is a direct metric of the charge carrier transport. It is challenging but crucial to develop highly conductive doped a-Si:H for minimizing the FF losses. Here we report intensive light soaking can efficiently boost the dark conductance of B-doped a-Si:H "thin" films, which is an abnormal Staebler-Wronski effect. By implementing this abnormal effect to SHJ solar cells, we achieve a certificated power conversion efficiency (PCE) of 25.18% (26.05% on designated area) with FF of 85.42% on a 244.63-cm2 wafer. This PCE is one of the highest reported values for total-area "top/rear" contact silicon solar cells. The FF reaches 98.30 per cent of its Shockley-Queisser limit.

physics.app-ph