SearcharxivSearch

arXiv subjects

Zhichuan Wang

Publications and source records attributed to Zhichuan Wang.

13 recordsLinked to original sources

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to build view-based 3D descriptors. Despite CLIP's strong generalization ability, its lack of fine-grainedness prompted us to explore the potential of a more recent self-supervised encoder-DINO. To address this, we propose DINO Eats CLIP (DEC), a novel framework for dynamic multi-view integration that is regularized by synthesizing data for unseen classes. We first find that simply mean-pooling over view features from a frozen DINO backbone gives decent performance. Yet, further adaptation causes severe overfitting on average view patterns of known classes. To combat it, we then design a module named Chunking and Adapting Module (CAM). It segments multi-view images into chunks and dynamically integrates local view relations, yielding more robust features than the standard pooling strategy. Finally, we propose Virtual Feature Synthesis (VFS) module to mitigate bias towards known categories explicitly. Under the hood, VFS leverages CLIP's broad, pre-aligned vision-language space to synthesize virtual features for unseen classes. By exposing DEC to these virtual features, we greatly enhance its open-set discrimination capacity. Extensive experiments on standard open-set 3DOR benchmarks demonstrate its superior efficacy.

cs.CV

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

Open-set 3D object retrieval (3DOR) is an emerging task aiming to retrieve 3D objects of unseen categories beyond the training set. Existing methods typically utilize all modalities (i.e., voxels, point clouds, multi-view images) and train specific backbones before fusion. However, they still struggle to produce generalized representations due to insufficient 3D training data. Being contrastively pre-trained on web-scale image-text pairs, CLIP inherently produces generalized representations for a wide range of downstream tasks. Building upon it, we present a simple yet effective framework named Describe, Adapt and Combine (DAC) by taking only multi-view images for open-set 3DOR. DAC innovatively synergizes a CLIP model with a multi-modal large language model (MLLM) to learn generalized 3D representations, where the MLLM is used for dual purposes. First, it describes the seen category information to align with CLIP's training objective for adaptation during training. Second, it provides external hints about unknown objects complementary to visual cues during inference. To improve the synergy, we introduce an Additive-Bias Low-Rank adaptation (AB-LoRA), which alleviates overfitting and further enhances the generalization to unseen categories. With only multi-view images, DAC significantly surpasses prior arts by an average of +10.01\% mAP on four open-set 3DOR datasets. Moreover, its generalization is also validated on image-based and cross-dataset setups. Code is available at https://github.com/wangzhichuan123/DAC.

cs.CV

TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment

Learning discriminative 3D representations that generalize well to unknown testing categories is an emerging requirement for many real-world 3D applications. Existing well-established methods often struggle to attain this goal due to insufficient 3D training data from broader concepts. Meanwhile, pre-trained large vision-language models (e.g., CLIP) have shown remarkable zero-shot generalization capabilities. Yet, they are limited in extracting suitable 3D representations due to substantial gaps between their 2D training and 3D testing distributions. To address these challenges, we propose Testing-time Distribution Alignment (TeDA), a novel framework that adapts a pretrained 2D vision-language model CLIP for unknown 3D object retrieval at test time. To our knowledge, it is the first work that studies the test-time adaptation of a vision-language model for 3D feature learning. TeDA projects 3D objects into multi-view images, extracts features using CLIP, and refines 3D query embeddings with an iterative optimization strategy by confident query-target sample pairs in a self-boosting manner. Additionally, TeDA integrates textual descriptions generated by a multimodal language model (InternVL) to enhance 3D object understanding, leveraging CLIP's aligned feature space to fuse visual and textual cues. Extensive experiments on four open-set 3D object retrieval benchmarks demonstrate that TeDA greatly outperforms state-of-the-art methods, even those requiring extensive training. We also experimented with depth maps on Objaverse-LVIS, further validating its effectiveness. Code is available at https://github.com/wangzhichuan123/TeDA.

cs.CV

Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray

Medical foundation models have the potential to revolutionize healthcare by providing robust and generalized representations of medical data. Medical vision-language pre-training has emerged as a promising approach for learning domain-general representations of medical image and text. Current algorithms that exploit global and local alignment between medical image and text could however be marred by redundant information in medical data. To address this issue, we propose a grounded knowledge-enhanced medical vision-language pre-training (GK-MVLP) framework for chest X-ray. In this framework, medical knowledge was grounded to the appropriate anatomical regions by using a transformer-based grounded knowledge-enhanced module for fine-grained alignment between textural features of medical knowledge and the corresponding anatomical region-level visual features. The performance of GK-MVLP was competitive with or exceeded the state of the art on downstream image understanding tasks (chest X-ray disease classification, disease localization), generative task (report generation), and vision-language understanding task (medical visual question-answering). Our results demonstrate the advantage of incorporating grounding mechanism to remove biases and improve the alignment between chest X-ray image and radiology report.

cs.CV

Expert Insight-Enhanced Follow-up Chest X-Ray Summary Generation

A chest X-ray radiology report describes abnormal findings not only from X-ray obtained at current examination, but also findings on disease progression or change in device placement with reference to the X-ray from previous examination. Majority of the efforts on automatic generation of radiology report pertain to reporting the former, but not the latter, type of findings. To the best of the authors' knowledge, there is only one work dedicated to generating summary of the latter findings, i.e., follow-up summary. In this study, we therefore propose a transformer-based framework to tackle this task. Motivated by our observations on the significance of medical lexicon on the fidelity of summary generation, we introduce two mechanisms to bestow expert insight to our model, namely expert soft guidance and masked entity modeling loss. The former mechanism employs a pretrained expert disease classifier to guide the presence level of specific abnormalities, while the latter directs the model's attention toward medical lexicon. Extensive experiments were conducted to demonstrate that the performance of our model is competitive with or exceeds the state-of-the-art.

cs.MM

In situ tuning of dynamical Coulomb blockade on Andreev bound states in hybrid nanowire devices

Electron interactions in quantum devices can exhibit intriguing phenomena. One example is assembling an electronic device in series with an on-chip resistor. The quantum laws of electricity of the device is modified at low energies and temperatures by dissipative interactions induced by the resistor, a phenomenon known as dynamical Coulomb blockade (DCB). The DCB strength is usually non-adjustable in a fixed environment defined by the resistor. Here, we design an on-chip circuit for InAs-Al hybrid nanowires where the DCB strength can be gate-tuned in situ. InAs-Al nanowires could host Andreev or Majorana zero-energy states. This technique enables tracking the evolution of the same state while tuning the DCB strength from weak to strong. We observe the transition from a zero-bias conductance peak to split peaks for Andreev zero-energy states. Our technique opens the door to in situ tuning interaction strength on zero-energy states.

cond-mat.mes-hall

Observation of Rydberg moiré excitons

Rydberg excitons, the solid-state counterparts of Rydberg atoms, have sparked considerable interest in harnessing their quantum application potentials, whereas a major challenge is realizing their spatial confinement and manipulation. Lately, the rise of two-dimensional moiré superlattices with highly tunable periodic potentials provides a possible pathway. Here, we experimentally demonstrate this capability through the observation of Rydberg moiré excitons (XRM), which are moiré trapped Rydberg excitons in monolayer semiconductor WSe2 adjacent to twisted bilayer graphene. In the strong coupling regime, the XRM manifest as multiple energy splittings, pronounced redshift, and narrowed linewidth in the reflectance spectra, highlighting their charge-transfer character where electron-hole separation is enforced by the strongly asymmetric interlayer Coulomb interactions. Our findings pave the way for pursuing novel physics and quantum technology exploitation based on the excitonic Rydberg states.

cond-mat.mes-hall

Gatemon qubit based on a thin InAs-Al hybrid nanowire

We study a gate-tunable superconducting qubit (gatemon) based on a thin InAs-Al hybrid nanowire. Using a gate voltage to control its Josephson energy, the gatemon can reach the strong coupling regime to a microwave cavity. In the dispersive regime, we extract the energy relaxation time $T_1\sim$0.56 $μ$s and the dephasing time $T_2^* \sim$0.38 $μ$s. Since thin InAs-Al nanowires can have fewer or single sub-band occupation and recent transport experiment shows the existence of nearly quantized zero-bias conductance peaks, our result holds relevancy for detecting Majorana zero modes in thin InAs-Al nanowires using circuit quantum electrodynamics.

cond-mat.mes-hall

BDG-Net: Boundary Distribution Guided Network for Accurate Polyp Segmentation

Colorectal cancer (CRC) is one of the most common fatal cancer in the world. Polypectomy can effectively interrupt the progression of adenoma to adenocarcinoma, thus reducing the risk of CRC development. Colonoscopy is the primary method to find colonic polyps. However, due to the different sizes of polyps and the unclear boundary between polyps and their surrounding mucosa, it is challenging to segment polyps accurately. To address this problem, we design a Boundary Distribution Guided Network (BDG-Net) for accurate polyp segmentation. Specifically, under the supervision of the ideal Boundary Distribution Map (BDM), we use Boundary Distribution Generate Module (BDGM) to aggregate high-level features and generate BDM. Then, BDM is sent to the Boundary Distribution Guided Decoder (BDGD) as complementary spatial information to guide the polyp segmentation. Moreover, a multi-scale feature interaction strategy is adopted in BDGD to improve the segmentation accuracy of polyps with different sizes. Extensive quantitative and qualitative evaluations demonstrate the effectiveness of our model, which outperforms state-of-the-art models remarkably on five public polyp datasets while maintaining low computational complexity. Code: https://github.com/zihuanqiu/BDG-Net

eess.IV

Large Andreev bound state zero bias peaks in a weakly dissipative environment

We study Andreev bound states in hybrid InAs-Al nanowire devices. The energy of these states can be tuned to zero by gate voltage or magnetic field, revealing large zero bias peaks (ZBPs) near 2e^2/h in tunneling conductance. Probing these large ZBPs using a weakly dissipative lead reveals non-Fermi liquid temperature (T) dependence due to environmental Coulomb blockade (ECB), an interaction effect from the lead acting on the nanowire junction. By increasing T, these large ZBPs either show a height increase or a transition from split peaks to a ZBP, both deviate significantly from non-dissipative devices where a Fermi-liquid T dependence is revealed. Our result demonstrates the competing effect between ECB and thermal broadening on Andreev bound states.

cond-mat.mes-hall

A new quasi-one-dimensional superconductor parent compound NaMn$_6$Bi$_5$ with lower antiferromagnetic transition temperatures

Mn-based superconductor is rare and recently reported in quasi-one-dimensional KMn$_6$Bi$_5$ with [Mn$_6$Bi$_5$]-columns under high pressure. Here we report the synthesis, magnetic properties, electrical resistivity, and specific heat capacity of the newly-discovered quasi-one-dimensional NaMn$_6$Bi$_5$ single crystal. Compared with other AMn$_6$Bi$_5$ (A = K, Rb, and Cs), NaMn$_6$Bi$_5$ has larger intra-column Bi-Bi bond length, which may result in the two decoupled antiferromagnetic transitions at 47.3 K and 52.3 K. The relatively lower antiferromagnetic transition temperatures make NaMn$_6$Bi$_5$ a more suitable platform to explore Mn-based superconductors. Anisotropic resistivity and non-Fermi liquid behavior at low temperature are observed. Heat capacity measurement reveals that NaMn$_6$Bi$_5$ has similar Debye temperature with those of AMn$_6$Bi$_5$ (A = K and Rb), whereas the Sommerfeld coefficient is unusually large. Using first-principles calculations, the quite different density of states and an unusual enhancement near the Fermi level are observed for NaMn$_6$Bi$_5$, when compared with those of other AMn$_6$Bi$_5$ (A = K, Rb, and Cs) compounds.

cond-mat.supr-con

Suppressing Andreev bound state zero bias peaks using a strongly dissipative lead

Hybrid semiconductor-superconductor nanowires are predicted to host Majorana zero modes, manifested as zero-bias peaks (ZBPs) in tunneling conductance. ZBPs alone, however, are not sufficient evidence due to the ubiquitous presence of Andreev bound states in the same system. Here, we implement a strongly resistive normal lead in our InAs-Al nanowire devices and show that most of the expected ZBPs, corresponding to zero-energy Andreev bound states, can be suppressed, a phenomenon known as environmental Coulomb blockade. Our result is the first experimental demonstration of this dissipative interaction effect on Andreev bound states and can serve as a possible filter to narrow down ZBP phase diagram in future Majorana searches.

cond-mat.mes-hall

In Situ Epitaxy of Pure Phase Ultra-Thin InAs-Al Nanowires for Quantum Devices

Hybrid semiconductor-superconductor InAs-Al nanowires with uniform and defect-free crystal interfaces are one of the most promising candidates used in the quest for Majorana zero modes (MZMs). However, InAs nanowires often exhibit a high density of randomly distributed twin defects and stacking faults, which result in an uncontrolled and non-uniform InAs-Al interface. Furthermore, this type of disorder can create potential inhomogeneity in the wire, destroy the topological gap, and form trivial sub-gap states mimicking MZM in transport experiments. Further study shows that reducing the InAs nanowire diameter from growth can significantly suppress the formation of these defects and stacking faults. Here, we demonstrate the in situ growth of ultra-thin InAs nanowires with epitaxial Al film by molecular-beam epitaxy. Our InAs diameter (~ 30 nm) is only one-third of the diameters (~ 100 nm) commonly used in literatures. The ultra-thin InAs nanowires are pure phase crystals for various different growth directions, suggesting a low level of disorder. Transmission electron microscopy confirms an atomically sharp and uniform interface between the Al shell and the InAs wire. Quantum transport study on these devices resolves a hard induced superconducting gap and $2e^-$ periodic Coulomb blockade at zero magnetic field, a necessary step for future MZM experiments. A large zero bias conductance peak with a peak height reaching 80% of $2e^2/h$ is observed.

cond-mat.mtrl-sci