Searcharxiv⌕ Search

arXiv subjects

Yang Tan

Publications and source records attributed to Yang Tan.

52 records · Page 3Linked to original sources

Transferability-Guided Cross-Domain Cross-Task Transfer Learning

We propose two novel transferability metrics F-OTCE (Fast Optimal Transport based Conditional Entropy) and JC-OTCE (Joint Correspondence OTCE) to evaluate how much the source model (task) can benefit the learning of the target task and to learn more transferable representations for cross-domain cross-task transfer learning. Unlike the existing metric that requires evaluating the empirical transferability on auxiliary tasks, our metrics are auxiliary-free such that they can be computed much more efficiently. Specifically, F-OTCE estimates transferability by first solving an Optimal Transport (OT) problem between source and target distributions, and then uses the optimal coupling to compute the Negative Conditional Entropy between source and target labels. It can also serve as a loss function to maximize the transferability of the source model before finetuning on the target task. Meanwhile, JC-OTCE improves the transferability robustness of F-OTCE by including label distances in the OT problem, though it may incur additional computation cost. Extensive experiments demonstrate that F-OTCE and JC-OTCE outperform state-of-the-art auxiliary-free metrics by 18.85% and 28.88%, respectively in correlation coefficient with the ground-truth transfer accuracy. By eliminating the training cost of auxiliary tasks, the two metrics reduces the total computation time of the previous method from 43 minutes to 9.32s and 10.78s, respectively, for a pair of tasks. When used as a loss function, F-OTCE shows consistent improvements on the transfer accuracy of the source model in few-shot classification experiments, with up to 4.41% accuracy gain.

cs.CV↗

PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications

Large protein language models are adept at capturing the underlying evolutionary information in primary structures, offering significant practical value for protein engineering. Compared to natural language models, protein amino acid sequences have a smaller data volume and a limited combinatorial space. Choosing an appropriate vocabulary size to optimize the pre-trained model is a pivotal issue. Moreover, despite the wealth of benchmarks and studies in the natural language community, there remains a lack of a comprehensive benchmark for systematically evaluating protein language model quality. Given these challenges, PETA trained language models with 14 different vocabulary sizes under three tokenization methods. It conducted thousands of tests on 33 diverse downstream datasets to assess the models' transfer learning capabilities, incorporating two classification heads and three random seeds to mitigate potential biases. Extensive experiments indicate that vocabulary sizes between 50 and 200 optimize the model, whereas sizes exceeding 800 detrimentally affect the model's representational performance. Our code, model weights and datasets are available at https://github.com/ginnm/ProteinPretraining.

cs.CL↗

Non-zero Integral Spin of Acoustic Vortices and Spin-orbit Interaction in Longitudinal Acoustics

Spin and orbital angular momenta (AM) are of fundamental interest in wave physics. Acoustic wave, as a typical longitudinal wave, has been well studied in terms of orbital AM, but still considered unable to carry non-zero integral spin AM or spin-orbital interaction in homogeneous media due to its spin-0 nature. Here we give the first self-consistent analytical calculations of spin, orbital and total AM of guided vortices under different boundary conditions, revealing that vortex field can carry non-zero integral spin AM. We also introduce for acoustic waves the canonical-Minkowski and kinetic-Abraham AM, which has aroused long-lasting debate in optics, and prove that only the former is conserved with the corresponding symmetries. Furthermore, we present the theoretical and experimental observation of the spin-orbit interaction of vortices in longitudinal acoustics, which is thought beyond attainable in longitudinal waves in the absence of spin degree of freedom. Our work provides a solid platform for future studies of the spin and orbital AM of guided acoustic waves and may open up a new dimension for acoustic vortex-based applications such as underwater communications and object manipulations.

physics.flu-dyn↗

MedChatZH: a Better Medical Adviser Learns from Better Instructions

Generative large language models (LLMs) have shown great success in various applications, including question-answering (QA) and dialogue systems. However, in specialized domains like traditional Chinese medical QA, these models may perform unsatisfactorily without fine-tuning on domain-specific datasets. To address this, we introduce MedChatZH, a dialogue model designed specifically for traditional Chinese medical QA. Our model is pre-trained on Chinese traditional medical books and fine-tuned with a carefully curated medical instruction dataset. It outperforms several solid baselines on a real-world medical dialogue dataset. We release our model, code, and dataset on https://github.com/tyang816/MedChatZH to facilitate further research in the domain of traditional Chinese medicine and LLMs.

cs.CL↗

Single-sensor and real-time ultrasonic imaging using an AI-driven disordered metasurface

Non-destructive testing and medical diagnostic techniques using ultrasound has become indispensable in evaluating the state of materials or imaging the internal human body, respectively. To conduct spatially resolved high-quality observations, conventionally, sophisticated phased arrays are used both at the emitting and receiving ends of the setup. In comparison, single-sensor imaging techniques offer significant benefits including compact physical dimensions and reduced manufacturing expenses. However, recent advances such as compressive sensing have shown that this improvement comes at the cost of additional time-consuming dynamic spatial scanning or multi-mode mask switching, which severely hinders the quest for real-time imaging. Consequently, real-time single-sensor imaging, at low cost and simple design, still represents a demanding and largely unresolved challenge till this day. Here, we bestow on ultrasonic metasurface with both disorder and artificial intelligence (AI). The former ensures strong dispersion and highly complex scattering to encode the spatial information into frequency spectra at an arbitrary location, while the latter is used to decode instantaneously the amplitude and spectral component of the sample under investigation. Thus, thanks to this symbiosis, we demonstrate that a single fixed sensor suffices to recognize complex ultrasonic objects through the random scattered field from an unpretentious metasurface, which enables real-time and low-cost imaging, easily extendable to 3D.

physics.app-ph↗

Multi-level Protein Representation Learning for Blind Mutational Effect Prediction

Directed evolution plays an indispensable role in protein engineering that revises existing protein sequences to attain new or enhanced functions. Accurately predicting the effects of protein variants necessitates an in-depth understanding of protein structure and function. Although large self-supervised language models have demonstrated remarkable performance in zero-shot inference using only protein sequences, these models inherently do not interpret the spatial characteristics of protein structures, which are crucial for comprehending protein folding stability and internal molecular interactions. This paper introduces a novel pre-training framework that cascades sequential and geometric analyzers for protein primary and tertiary structures. It guides mutational directions toward desired traits by simulating natural selection on wild-type proteins and evaluates the effects of variants based on their fitness to perform the function. We assess the proposed approach using a public database and two new databases for a variety of variant effect prediction tasks, which encompass a diverse set of proteins and assays from different taxa. The prediction results achieve state-of-the-art performance over other zero-shot learning methods for both single-site mutations and deep mutations.

q-bio.QM↗

Finding the Most Transferable Tasks for Brain Image Segmentation

Although many studies have successfully applied transfer learning to medical image segmentation, very few of them have investigated the selection strategy when multiple source tasks are available for transfer. In this paper, we propose a prior knowledge guided and transferability based framework to select the best source tasks among a collection of brain image segmentation tasks, to improve the transfer learning performance on the given target task. The framework consists of modality analysis, RoI (region of interest) analysis, and transferability estimation, such that the source task selection can be refined step by step. Specifically, we adapt the state-of-the-art analytical transferability estimation metrics to medical image segmentation tasks and further show that their performance can be significantly boosted by filtering candidate source tasks based on modality and RoI characteristics. Our experiments on brain matter, brain tumor, and white matter hyperintensities segmentation datasets reveal that transferring from different tasks under the same modality is often more successful than transferring from the same task under different modalities. Furthermore, within the same modality, transferring from the source task that has stronger RoI shape similarity with the target task can significantly improve the final transfer performance. And such similarity can be captured using the Structural Similarity index in the label space.

eess.IV↗

Mistakes of A Popular Protocol Calculating Private Set Intersection and Union Cardinality and Its Corrections

In 2012, De Cristofaro et al. proposed a protocol to calculate the Private Set Intersection and Union cardinality(PSI-CA and PSU-CA). This protocol's security is based on the famous DDH assumption. Since its publication, it has gained lots of popularity because of its efficiency(linear complexity in computation and communication) and concision. So far, it's still considered one of the most efficient PSI-CA protocols and the most cited(more than 170 citations) PSI-CA paper based on the Google Scholar search. However, when we tried to implement this protocol, we couldn't get the correct result of the test data. Since the original paper lacks of experimental results to verify the protocol's correctness, we looked deeper into the protocol and found out it made a fundamental mistake. Needless to say, its correctness analysis and security proof are also wrong. In this paper, we will point out this PSI-CA protocol's mistakes, and provide the correct version of this protocol as well as the PSI protocol developed from this protocol. We also present a new security proof and some experimental results of the corrected protocol.

cs.CR↗

OTCE: A Transferability Metric for Cross-Domain Cross-Task Representations

Transfer learning across heterogeneous data distributions (a.k.a. domains) and distinct tasks is a more general and challenging problem than conventional transfer learning, where either domains or tasks are assumed to be the same. While neural network based feature transfer is widely used in transfer learning applications, finding the optimal transfer strategy still requires time-consuming experiments and domain knowledge. We propose a transferability metric called Optimal Transport based Conditional Entropy (OTCE), to analytically predict the transfer performance for supervised classification tasks in such cross-domain and cross-task feature transfer settings. Our OTCE score characterizes transferability as a combination of domain difference and task difference, and explicitly evaluates them from data in a unified framework. Specifically, we use optimal transport to estimate domain difference and the optimal coupling between source and target distributions, which is then used to derive the conditional entropy of the target task (task difference). Experiments on the largest cross-domain dataset DomainNet and Office31 demonstrate that OTCE shows an average of 21% gain in the correlation with the ground truth transfer accuracy compared to state-of-the-art methods. We also investigate two applications of the OTCE score including source model selection and multi-source feature fusion.

cs.LG↗

Four-Dimensional Higher-Order Chern Insulator and Its Acoustic Realization

We present a theoretical study and experimental realization of a system that is simultaneously a four-dimensional (4D) Chern insulator and a higher-order topological insulator (HOTI). The system sustains the coexistence of (4-1)-dimensional chiral topological hypersurface modes (THMs) and (4-2)-dimensional chiral topological surface modes (TSMs). Our study reveals that the THMs are protected by second Chern numbers, and the TSMs are protected by a topological invariant composed of two first Chern numbers, each belonging a Chern insulator existing in sub-dimensions. With the synthetic coordinates fixed, the THMs and TSMs respectively manifest as topological edge modes (TEMs) and topological corner modes (TCMs) in the real space, which are experimentally observed in a 2D acoustic lattice. These TCMs are not related to quantized polarizations, making them fundamentally distinctive from existing examples. We further show that our 4D topological system offers an effective way for the manipulation of the frequency, location, and the number of the TCMs, which is highly desirable for applications.

physics.app-ph↗

Smart Cameras

We review camera architecture in the age of artificial intelligence. Modern cameras use physical components and software to capture, compress and display image data. Over the past 5 years, deep learning solutions have become superior to traditional algorithms for each of these functions. Deep learning enables 10-100x reduction in electrical sensor power per pixel, 10x improvement in depth of field and dynamic range and 10-100x improvement in image pixel count. Deep learning enables multiframe and multiaperture solutions that fundamentally shift the goals of physical camera design. Here we review the state of the art of deep learning in camera operations and consider the impact of AI on the physical design of cameras.

eess.IV↗

Justlookup: One Millisecond Deep Feature Extraction for Point Clouds By Lookup Tables

Deep models are capable of fitting complex high dimensional functions while usually yielding large computation load. There is no way to speed up the inference process by classical lookup tables due to the high-dimensional input and limited memory size. Recently, a novel architecture (PointNet) for point clouds has demonstrated that it is possible to obtain a complicated deep function from a set of 3-variable functions. In this paper, we exploit this property and apply a lookup table to encode these 3-variable functions. This method ensures that the inference time is only determined by the memory access no matter how complicated the deep function is. We conduct extensive experiments on ModelNet and ShapeNet datasets and demonstrate that we can complete the inference process in 1.5 ms on an Intel i7-8700 CPU (single core mode), 32x speedup over the PointNet architecture without any performance degradation.

cs.CV↗

Face Recognition from Sequential Sparse 3D Data via Deep Registration

Previous works have shown that face recognition with high accurate 3D data is more reliable and insensitive to pose and illumination variations. Recently, low-cost and portable 3D acquisition techniques like ToF(Time of Flight) and DoE based structured light systems enable us to access 3D data easily, e.g., via a mobile phone. However, such devices only provide sparse(limited speckles in structured light system) and noisy 3D data which can not support face recognition directly. In this paper, we aim at achieving high-performance face recognition for devices equipped with such modules which is very meaningful in practice as such devices will be very popular. We propose a framework to perform face recognition by fusing a sequence of low-quality 3D data. As 3D data are sparse and noisy which can not be well handled by conventional methods like the ICP algorithm, we design a PointNet-like Deep Registration Network(DRNet) which works with ordered 3D point coordinates while preserving the ability of mining local structures via convolution. Meanwhile we develop a novel loss function to optimize our DRNet based on the quaternion expression which obviously outperforms other widely used functions. For face recognition, we design a deep convolutional network which takes the fused 3D depth-map as input based on AMSoftmax model. Experiments show that our DRNet can achieve rotation error 0.95° and translation error 0.28mm for registration. The face recognition on fused data also achieves rank-1 accuracy 99.2% , FAR-0.001 97.5% on Bosphorus dataset which is comparable with state-of-the-art high-quality data based recognition performance.

cs.CV↗

Ultrasensitive biosensor based on Nd:YAG waveguide laser: Tumor cell and Dextrose solution

This work demonstrates the Nd:YAG waveguide laser as an efficient platform for the bio-sensing. The waveguide was fabricated in the Nd:YAG crystal by the cooperation of the ultrafast laser writing and ion irradiation. As the laser oscillation in the Nd:YAG waveguide is ultra-sensitivity to the external environment of the waveguide. Even a weak disturbance would induce a large variation of the output power of the laser. According to this feature, the Nd:YAG waveguide coated with Graphene and WSe2 layers is used as substrate for the microfluidic channel. When the microflow crosses the Nd:YAG waveguide, the laser oscillation in the waveguide is disturbed, and induces the fluctuation of the output laser. Through the analysis of the fluctuation, the concentration of the dextrose solution and the size of the tumor cell are distinguished

physics.app-ph↗

Tuning of Interlayer Coupling in Large-Area Graphene/WSe2 van der Waals Heterostructure via Ion Irradiation: Optical Evidences and Photonic Applications

Van der Waals (vdW) heterostructures are receiving great attentions due to their intriguing properties and potentials in many research fields. The flow of charge carriers in vdW heterostructures can be efficiently rectified by the inter-layer coupling between neighboring layers, offering a rich collection of functionalities and a mechanism for designing atomically thin devices. Nevertheless, non-uniform contact in larger-area heterostructures reduces the device efficiency. In this work, ion irradiation had been verified as an efficient technique to enhance the contact and interlayer coupling in the newly developed graphene/WSe2 hetero-structure with a large area of 10 mm x 10 mm. During the ion irradiation process, the morphology of monolayer graphene had been modified, promoting the contact with WSe2. Experimental evidences of the tunable interlayer electron transfer are displayed by investigation of photoluminescence and ultrafast absorption of the irradiated heterostructure. Besides, we have found that in graphene/WSe2 heterostructure, graphene serves as a fast channel for the photo-excited carriers to relax in WSe2, and the nonlinear absorption of WSe2 could be effectively tuned by the carrier transfer process in graphene, enabling specific optical absorption of the heterostructure in comparison with separated graphene or WSe2. On the basis of these new findings, by applying the ion beam modified graphene/WSe2 heterostructure as a saturable absorber, Q-switched pulsed lasing with optimized performance has been realized in a Nd:YAG waveguide cavity. This work paves the way towards developing novel devices based on large-area heterostructures by using ion beam irradiation.

cond-mat.mes-hall↗

Layer compression and enhanced optical properties of few-layer graphene nanosheets induced by ion irradiation

Graphene has been recognized as an attractive two-dimensional material for fundamental research and wide applications in electronic and photonic devices owing to its unique properties. The technologies to modulate the properties of graphene are of continuous interest to researchers in multidisciplinary areas. Herein, we report on the first experimental observation of the layer-to-layer compression and enhanced optical properties of few-layer graphene nanosheets by applying the irradiation of energetic ion beams. After the irradiation, the space between the graphene layers was reduced, resulting in a tighter contact between the few-layer graphene nanosheet and the surface of the substrate. This processing also enhanced the interaction between the graphene nanosheets and the evanescent-field wave near the surface, thus reinforcing the polarization-dependent light absorption of the graphene layers (with 3-fold polarization extinction ratio increment). Utilizing the ion-irradiated graphene nanosheets as saturable absorbers, the passively Q-switched waveguide lasing with considerably improved performances was achieved, owing to the enhanced interactions between the graphene nanosheets and evanescent field of light. The obtained repetition rate of waveguide laser was up to 2.3 MHz with a pulse duration of 101 ns.

physics.optics↗