SearcharxivSearch

arXiv subjects

Kaixiang Chen

Publications and source records attributed to Kaixiang Chen.

7 recordsLinked to original sources

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models

Adapting large vision-language models (VLMs) such as CLIP to downstream tasks remains challenging, as full fine-tuning is computationally prohibitive and prone to overfitting in low-data regimes. Parameter-efficient fine-tuning (PEFT) alleviates these issues with lightweight prompt- or adapter-based modules, and cross-modal coupling has proven especially effective by strengthening interactions between vision and language. However, existing coupling mechanisms predominantly rely on external auxiliary modules, leading to indirect, coarse-grained interactions that are structurally decoupled from the original VLM and thus limit representational expressiveness. In this paper, we propose Multi-Modal Interactive Agent Layer (MAIL), a PEFT paradigm that embeds cross-modal coupling directly into the intrinsic computation modules of VLMs. MAIL freezes the backbone and inserts lightweight agent layers after core modules, such as LayerNorm, to approximate the parameter updates induced by full fine-tuning. To couple visual and textual streams at this level, we introduce a bottleneck-based text-to-image bridge that jointly optimizes paired agent layers across modalities, coordinating the adaptation of corresponding computation modules. We further present MAIL++, which enables bidirectional cross-modal exchange through a meta agent layer, a meta-text bridge, and a meta-image bridge. At inference time, all agent layers are re-parameterized into the frozen backbone, preserving the original computational efficiency. Extensive experiments on few-shot image classification and few-shot universal cross-domain retrieval demonstrate that MAIL and MAIL++ consistently outperform state-of-the-art PEFT methods.

cs.CV

Harish-Chandra Theorem for the Multi-Parameter Quantum Groups of Okado-Yamane Type

This paper is devoted to studying the centre of the multi-parameter quantum group $U_{q,G}(\mathfrak{g})$ introduced by Okado and Yamane, where $\mathfrak{g}$ is a complex simple Lie algebra, and all parameters lie in general position. We mainly establish the Harish-Chandra theorem, proving that the Harish-Chandra homomorphism is an isomorphism; in particular, we determine the centre $Z(U_{q,G})\cong (U^0_\flat)^W$ is isomorphic to a polynomial algebra or a quotient algebra of a polynomial algebra. The same result holds for the $(U^0_\flat)^W$ of the two-parameter quantum group $U_{r,s}(\mathfrak{g})$.

math.QA

Distributed Ranging SLAM for Multiple Robots with Ultra-WideBand and Odometry Measurements

To accomplish task efficiently in a multiple robots system, a problem that has to be addressed is Simultaneous Localization and Mapping (SLAM). LiDAR (Light Detection and Ranging) has been used for many SLAM solutions due to its superb accuracy, but its performance degrades in featureless environments, like tunnels or long corridors. Centralized SLAM solves the problem with a cloud server, which requires a huge amount of computational resources and lacks robustness against central node failure. To address these issues, we present a distributed SLAM solution to estimate the trajectory of a group of robots using Ultra-WideBand (UWB) ranging and odometry measurements. The proposed approach distributes the processing among the robot team and significantly mitigates the computation concern emerged from the centralized SLAM. Our solution determines the relative pose (also known as loop closure) between two robots by minimizing the UWB ranging measurements taken at different positions when the robots are in close proximity. UWB provides a good distance measure in line-of-sight conditions, but retrieving a precise pose estimation remains a challenge, due to ranging noise and unpredictable path traveled by the robot. To deal with the suspicious loop closures, we use Pairwise Consistency Maximization (PCM) to examine the quality of loop closures and perform outlier rejections. The filtered loop closures are then fused with odometry in a distributed pose graph optimization (DPGO) module to recover the full trajectory of the robot team. Extensive experiments are conducted to validate the effectiveness of the proposed approach.

cs.RO

Explicit gain equations for hybrid graphene-quantum-dot photodetectors

Graphene is an attractive material for broadband photodetection but suffers from weak light absorption. Coating graphene with quantum dots can significantly enhance light absorption and create extraordinarily high photo gain. This high gain is often explained by the classical gain theory which is unfortunately an implicit function and may even be questionable. In this work, we managed to derive explicit gain equations for hybrid graphene-quantum-dot photodetectors. Due to the work function mismatch, lead sulfide (PbS) quantum dots coated on graphene will form a surface depletion region near the interface of quantum dots and graphene. Light illumination narrows down the surface depletion region, creating a photovoltage that gates the graphene. As a result, high photo gain in graphene is observed. The explicit gain equations are derived from the theoretical gate transfer characteristics of graphene and the correlation of the photovoltage with the light illumination intensity. The derived explicit gain equations fit well with the experimental data, from which physical parameters are extracted.

cond-mat.mes-hall

Explicit Gain Equations for Single Crystalline Photoconductors

Photoconductors based on semiconducting thin films, nanowires and 2-dimensional atomic layers have been extensively investigated. But there is no explicit photogain equation that allows for fitting and designing photoresponses of these devices. In this work, we managed to derive explicit photogain equations for silicon nanowire photoconductors based on experimental observations. The silicon nanowires were fabricated by patterning the device layer of silicon-on-insulator wafers by standard lithography that were doped with boron. It was found that the as-fabricated silicon nanowires have a surface depletion region ~ 32 nm wide. This depletion region protects charge carriers in the channel from surface scatterings, resulting in the independence of charge carrier mobilities on nanowire size. It is consistent with our Hall effect measurements but in contradiction with the accepted conclusion in the past decades that charge carrier mobilities become smaller for smaller nanowires due to surface scatterings. Under light illumination, the depletion region logarithmically narrows down and the nanowire channel widens accordingly. Photo Hall effect measurements show that the nanowire photoconductance is not contributed by the increase of carrier concentrations but the widening of the nanowire channel. As a result, a nanowire photoconductor can be modeled as a resistor in connection with floating Schottky junctions near the nanowire surfaces. Based on the photoresponses of a Schottky junction, we derived explicit photogain equations for nanowire photoconductors that are a function of light intensity and device physical parameters. The gain equations fit well with the experimental data, from which we extracted the minority carrier lifetimes that are consistent with the minority carrier lifetime in nanowires reported in literature.

cond-mat.mes-hall

A photoconductor intrinsically has no gain

In the past 50 years, the high gain in quantum efficiency of photoconductors is often explained by a widely accepted theory in which the photogain is proportional to the minority carrier lifetime and inversely proportional to the carrier transit time across the photoconductor. It occasionally misleads scientists to believe that a high-speed and high-gain photodetector can be made simply by shortening the device length. The theory is derived on the assumption that the distribution of photogenerated excess carriers is spatially uniform. In this Letter, we find that this assumption is not valid for a photoconductive semiconductor due to the metal-semiconductor boundary at the two metal electrodes inducing carrier confinement. By solving the continuity equation and performing numerical simulations, we conclude that a photoconductor intrinsically has no gain or at least no high gain, no matter how short the transit time and how long the minority lifetime is. The high gain observed in experiments comes from other extrinsic effects such as defects, surface states and surface depletion regions that localize excess minority carriers, leaving a large number of excess majority carriers accumulated in the conduction channel for the photogain. Following the Ohm's Law, a universal equation governing the photogain in a photoconductor is established at the end of this Letter.

cond-mat.mes-hall

Self-aligned process for forming microlenses at the tips of vertical silicon nanowires by atomic layer deposition

The microlens is a key enabling technology in optoelectronics, permitting light to be efficiently coupled to and from devices such as image sensors and light-emitting diodes. Their ubiquitous nature motivates the development of new fabrication techniques, since existing methods face challenges as microlenses are scaled to smaller dimensions. Here, we demonstrate the formation of microlenses at the tips of vertically-oriented silicon nanowires via a rapid atomic layer deposition (ALD) process. The nature of the process is such that the microlenses are centered on the nanowires, and there is a self-limiting effect on the final sizes of the microlenses arising from the nanowire spacing. Finite difference time domain electromagnetic simulations are performed of microlens focusing properties, including showing their ability to enhance visible-wavelength absorption in silicon nanostructures.

cond-mat.mtrl-sci