SearcharxivSearch

arXiv subjects

Linjun Li

Publications and source records attributed to Linjun Li.

At least 37 records · Page 2Linked to original sources

Fast Fabrication of WS2/Bi2Se3 Heterostructures for High Performance Photodetection

Two-dimensional (2D) material heterostructures have attracted considerable attention owing to their interesting and novel physical properties, which expand the possibilities for future optoelectronic, photovoltaic, and nanoelectronic applications. A portable, fast, and deterministic transfer technique is highly needed for the fabrication of heterostructures. Herein, we report a fast half wet poly(dimethylsiloxane) (PDMS) transfer process utilizing the change of adhesion energy with the help of micron-sized water droplets. Using this method, a vertical stacking of the WS2/Bi2Se3 heterostructure with a straddling band configuration is successfully assembled on a fluorophlogopite substrate. Thanks to the complementary band gaps and high efficiency of interfacial charge transfer, the photodetector based on the heterostructure exhibits a superior responsivity of 109.9 A/W for a visible incident light at 473 nm and 26.7 A/W for a 1064 nm near-infrared illumination. Such high photoresponsivity of the heterostructure demonstrates that our transfer method not only owns time efficiency but also ensures high quality of the heterointerface. Our study may open new pathways to the fast and massive fabrication of various vertical 2D heterostructures for applications in twistronics/valleytronics and other band engineering devices.

cond-mat.mes-hall

Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Grounding

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair annotations in 3D point clouds, which are both time-consuming and expensive. To address the problem that fine-grained annotated data is difficult to obtain, we propose to leverage weakly supervised annotations to learn the 3D visual grounding model, i.e., only coarse scene-sentence correspondences are used to learn object-sentence links. To accomplish this, we design a novel semantic matching model that analyzes the semantic similarity between object proposals and sentences in a coarse-to-fine manner. Specifically, we first extract object proposals and coarsely select the top-K candidates based on feature and class similarity matrices. Next, we reconstruct the masked keywords of the sentence using each candidate one by one, and the reconstructed accuracy finely reflects the semantic similarity of each candidate to the query. Additionally, we distill the coarse-to-fine semantic matching knowledge into a typical two-stage 3D visual grounding model, which reduces inference costs and improves performance by taking full advantage of the well-studied structure of the existing architectures. We conduct extensive experiments on ScanRefer, Nr3D, and Sr3D, which demonstrate the effectiveness of our proposed method.

cs.CV

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Speech Recognition builds a bridge between the multimedia streaming (audio-only, visual-only or audio-visual) and the corresponding text transcription. However, when training the specific model of new domain, it often gets stuck in the lack of new-domain utterances, especially the labeled visual utterances. To break through this restriction, we attempt to achieve zero-shot modality transfer by maintaining the multi-modality alignment in phoneme space learned with unlabeled multimedia utterances in the high resource domain during the pre-training \cite{shi2022learning}, and propose a training system Open-modality Speech Recognition (\textbf{OpenSR}) that enables the models trained on a single modality (e.g., audio-only) applicable to more modalities (e.g., visual-only and audio-visual). Furthermore, we employ a cluster-based prompt tuning strategy to handle the domain shift for the scenarios with only common words in the new domain utterances. We demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few- and full-shot), and achieves highly competitive zero-shot performance compared to the existing few-shot and full-shot lip-reading methods. To the best of our knowledge, OpenSR achieves the state-of-the-art performance of word error rate in LRS2 on audio-visual speech recognition and lip-reading with 2.7\% and 25.0\%, respectively. The code and demo are available at https://github.com/Exgc/OpenSR.

cs.CL

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e., the movement of lips and teeth). In this work, we present AV-TranSpeech, the first audio-visual speech-to-speech (AV-S2ST) translation model without relying on intermediate text. AV-TranSpeech complements the audio stream with visual information to promote system robustness and opens up a host of practical applications: dictation or dubbing archival films. To mitigate the data scarcity with limited parallel AV-S2ST data, we 1) explore self-supervised pre-training with unlabeled audio-visual data to learn contextual representation, and 2) introduce cross-modal distillation with S2ST models trained on the audio-only corpus to further reduce the requirements of visual data. Experimental results on two language pairs demonstrate that AV-TranSpeech outperforms audio-only models under all settings regardless of the type of noise. With low-resource audio-visual data (10h, 30h), cross-modal distillation yields an improvement of 7.6 BLEU on average compared with baselines. Audio samples are available at https://AV-TranSpeech.github.io

cs.CL

A simplified method characterizing magnetic ordering modulated photo-thermoelectric response in noncentrosymmetric semimetal Ca3Ru2O7

Photo-Thermoelectric (PTE) response is usually one of the main working mechanisms for photodetectors. However, as another fast and easier way to measure thermoelectric characteristics of materials, it can also reveal important physics such as electric-phonon coupling, electron-electron correlation, etc. Recently, the spin entropy related to magnetic order transition which contributes to thermoelectric power is attracting more and more attention. Here, we demonstrate the PTE response can be reshaped when Ca3Ru2O7 undergoes meta-magnetic phase (MMP) transition driven by both temperature and magnetic field. Firstly, a sign change is observed crossing TS = 48 K and the linear polarization angle dependent PTE current maximizes along a-axis above TS while maximizes along b-axis below TS, which indicates that the antiferromagnetic spin order contributes to such spatial anisotropy. Secondly, in the temperature range of around 40 ~ 50 K, the PTE current is found to be sharply suppressed when external magnetic field is applied in plane along a-axis but is only gradually suppressed when applied field is along b-axis which gives out two critical fields. We attribute such suppression of PTE current under magnetic field to the suppression of the spin entropy in the phase transition between the antiferromagnetic state and the MMP state and the H-T phase diagrams of Ca3Ru2O7 is redrawn accordingly. Compared to previously work which trying to understand the magnetic phase transition in Ca3Ru2O7, such as neutron scattering, specific heat, and other advanced transport measurements, our work provides a more convenient yet efficient method, which may also find applications in other correlated spin materials in general.

cond-mat.str-el

Highly sensitive photodetector based on two-dimensional ferroelectric semiconducting \{beta}-InSe/graphene heterostructure

2D ferroelectric \{beta}-InSe/graphene heterostructure was fabricated by mechanical exfoliation, and the carrier dynamics crossing the heterostructure interface has been systematically investigated by Raman, photoluminescence and transient absorption measurements. Due to the efficient interfacial photo excited electron transfer and photogating effect from trapped holes, the heterostructure devices demonstrate superior performance with maximum responsivity of 2.12*10e4 A/W, detectivity of 1.73*10e14 Jones and fast response time (241 us) under λ = 532 nm laser illumination. Furthermore, the photo responses influenced by ferroelectric polarization field are investigated. Our work confirms ferroelectric \{beta}-InSe/graphene heterostructure as an outstanding material platform for sensitive optoelectronic application.

cond-mat.mes-hall

Highly tunable lateral homojunction formed in 2D layered CuInP2S6 via in-plane ionic migration

As basic building blocks for next-generation information technologies devices, high-quality p-n junctions based on van der Waals (vdW) materials have attracted widespread interest.Compared to traditional two dimensional (2D) heterojunction diodes, the emerging homojunctions are more attractive owing to their intrinsic advantages, such as continuous band alignments and smaller carrier trapping. Here, utilizing the long-range migration of Cu + ions under in-plane electric field, a novel lateral p-n homojunction was constructed in the 2D layered copper indium thiophosphate (CIPS). The symmetric Au/CIPS/Au devices demonstrate an electric-field-driven resistance switching (RS) accompanying by a rectification behavior without any gate control. Moreover, such rectification behavior can be continuously modulated by poling voltage. We deduce that the reversable rectifying RS behavior is governed by the effective lateral build-in potential and the change of the interfacial barrier during the poling process. Furthermore, the CIPS p-n homojuction is evidenced by the photovoltaic effect, with the spectral response extending up to visible region due to the better photogenerated carrier separation efficiency. Our study provides a facile route to fabricate homojuctions through electric-field-driven ionic migration and paves the way towards the use of this method in other vdW materials.

cond-mat.mes-hall

Interfacial Charge Transfer and Ultrafast Photonics Application of 2D Graphene/InSe Heterostructure

Interface interactions in 2D vertically stacked heterostructures play an important role in optoe-lectronic applications, photodetectors based on graphene/InSe heterostructures had shown promising performance nowadays. However, nonlinear optical properties studies based on the graphene/InSe heterostructure was insufficient. Here, we fabricated graphene/InSe heterostruc-ture by mechanical exfoliation, and investigated the optically induced charge transfer between graphene/InSe heterostructures by taking photoluminescence and pump-probe measurements. The large built-in electric field at the interface is confirmed by Kelvin probe force microscopy. Furthermore, due to the efficient interfacial carrier transfer driven by built-in electric potential (~ 286 meV) and broadband nonlinear absorption, the application of graphene/InSe heterostruc-ture in mode-locked laser is realized. Our work not only provides a deeper understanding for the dipole orientation related interface interactions on the photoexcited charge transfer of gra-phene/InSe heterostructure, but also enrich the saturable absorber family for ultrafast photon-ics application.

physics.optics

MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present \textbf{AVMuST-TED}, the first dataset for \textbf{A}udio-\textbf{V}isual \textbf{Mu}ltilingual \textbf{S}peech \textbf{T}ranslation, derived from \textbf{TED} talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1\%), LRS2 (25.5\%), and LRS3 (28.0\%).

cs.CV

Motion Planning Transformers: A Motion Planning Framework for Mobile Robots

Fast and efficient sampling-based motion planning (SMP) is an integral component of many robotic systems, such as autonomous cars. A popular technique to improve the efficiency of these planners is to restrict search space in the planning domain. Existing algorithms define parametric functions to bound the search space, but these do not extend to non-holonomic robotic systems. Recent learning-based methods use a combination of convolutional and fully connected networks to encode the planning space. However, these methods are restricted to fixed map sizes, which are often not realistic in the real world. In this paper, we introduce a transformer-based approach, Motion Planning Transformer, to restrict the search space by learning to discern regions with a valid path from prior data. The model learns not only to restrict search spaces for simple 2D systems but also for non-holonomic robotic systems. We validate our method on various randomly generated environments with different map sizes and plan trajectories for a physical non-holonomic robot. We also provide a ROS2 plugin of our method for the Nav2 planning stack. The results show that our method reduces search space nodes by 2-12 times compared to traditional planners and has better generalizability than recent learning-based planners.

cs.RO

Organic metallic epsilon-near-zero materials with large ultrafast optical nonlinearity

Epsilon-near-zero (ENZ) materials have shown significant potential for nonlinear optical applications due to their ultrafast hot carriers and consequent optical nonlinearity enhancement. Modified poly(3,4-ethylenedioxythiophene) (PEDOT) films show metallic characteristics and a resultant ENZ wavelength near 1550nm through polar solvent treatment and annealing. The metallic PEDOT film exhibits an intrinsic optical nonlinear response that is comparable to gold and 100-fold higher than typical inorganic semiconductor ENZ materials due to π-conjugated delocalized electrons. Hot carriers generate a 22-fold increase in the optical nonlinearity coefficient of metallic PEDOT films at 1550 nm. Hot holes in metallic PEDOT films have a smaller enhancement multiple of carrier temperature and a longer relaxation time than hot electrons in inorganic ENZ materials due to the larger imaginary permittivity and hot-phonon bottleneck for carrier cooling. Our findings suggest that π-conjugated ENZ polymer may have unique ultrafast and nonlinear optical properties compared to inorganic ENZ materials, enabling new possibilities in on-chip nanophotonic devices, nonlinear optics, and plasmonics.

physics.optics

Anderson-Bernoulli localization at large disorder on the 2D lattice

We consider the Anderson model at large disorder on $\mathbb{Z}^2$ where the potential has a symmetric Bernoulli distribution. We prove that Anderson localization happens outside a small neighborhood of finitely many energies. These finitely many energies are Dirichlet eigenvalues of the minus Laplacian restricted on some finite subsets of $\mathbb{Z}^{2}$.

math.AP

Atomically smooth single-crystalline platform for low-loss plasmonic nanocavities

Nanoparticle-on-mirror plasmonic nanocavities, capable of extreme optical confinement and enhancement, have triggered state-of-the-art progress in nanophotonics and development of applications in enhanced spectroscopies and molecular detection. However, the optical quality factor and thus performance of these nanoconstructs are undermined by the granular polycrystalline metal films used as a mirror. Here, we report an atomically smooth single-crystalline platform for low-loss nanocavities using chemically-synthesized gold microflakes as a mirror. Nanocavities constructed using gold nanorods on such microflakes exhibit a rich structure of plasmonic modes, which are highly sensitive to the thickness of optically-thin (down to ~15 nm) microflakes. The atomically smooth single-crystalline microflakes endow nanocavities with significantly improved quality factor (~2 times) and scattering intensity (~3 times) compared with their counterparts based on deposited films. The developed low-loss nanocavities further allow for the integration with a mature platform of fiber optics, opening opportunities for realizing nanocavity-based miniaturized photonic devices with high performance.

physics.optics

Thickness dependent dark exciton emission in (PEA)2PbI4 nanoflake and its brightening by in-plane magnetic field

Halide perovskite materials raised tremendous interest in recent years since their cheap fabrication, superior performance in both solar cell and light emitting diode (LED). Due to the existence of layered quantum well structure, quasi two-dimensional(2D) halide perovskite has more intriguing spin related physics than its 3D counterpart. For instance, the detection and brightening of dark exciton (DX) in 2D halide perovskite attracts much attention since these species can be used in opto-spintronic and quantum computing devices. Here, we report the gradually brightened emission of the DX at 2.33 eV with the thickness decreases in (PEA)2PbI4 single crystalline nanoflake, which hitherto has not been reported. By coupling with in-plane (IP) magnetic field in Voigt configuration, the DX emission can be sharply enhanced, while for the out-of-plane (OP) magnetic field in Faraday configuration, the DX emission has no noticeable change, which can be reconciled with the theory interpretation of magnetic field dependent wave function mixing between the four exciton states fi1, fi2, fi3- , fi3+. The emission of DX fi2 at 2.335 eV and the fine splitting of all the four states are observed in static PL spectroscopy for the first time. Our work thus clarifies the debating questions regarding to previous research on DX behavior in 2D halide perovskite material and sheds light on the road of realizing opto-spintronic or quantum computing devices with these materials.

cond-mat.mtrl-sci

Anderson-Bernoulli Localization on the 3D lattice and discrete unique continuation principle

We consider the Anderson model with Bernoulli potential on the 3D lattice, and prove localization of eigenfunctions corresponding to eigenvalues near zero, the lower boundary of the spectrum. We follow the framework by Bourgain-Kenig and Ding-Smart, and our main contribution is a 3D discrete unique continuation, which says that any eigenfunction of the harmonic operator with bounded potential cannot be too small on a significant fractional portion of all the points. Its proof relies on geometric arguments about the 3D lattice.

math.AP

On the Manhattan pinball problem

We consider the periodic Manhattan lattice with alternating orientations going north-south and east-west. Place obstructions on vertices independently with probability $0 \frac{1}{2}-\varepsilon$ with some $\varepsilon>0$.

math.PR

MPC-MPNet: Model-Predictive Motion Planning Networks for Fast, Near-Optimal Planning under Kinodynamic Constraints

Kinodynamic Motion Planning (KMP) is to find a robot motion subject to concurrent kinematics and dynamics constraints. To date, quite a few methods solve KMP problems and those that exist struggle to find near-optimal solutions and exhibit high computational complexity as the planning space dimensionality increases. To address these challenges, we present a scalable, imitation learning-based, Model-Predictive Motion Planning Networks framework that quickly finds near-optimal path solutions with worst-case theoretical guarantees under kinodynamic constraints for practical underactuated systems. Our framework introduces two algorithms built on a neural generator, discriminator, and a parallelizable Model Predictive Controller (MPC). The generator outputs various informed states towards the given target, and the discriminator selects the best possible subset from them for the extension. The MPC locally connects the selected informed states while satisfying the given constraints leading to feasible, near-optimal solutions. We evaluate our algorithms on a range of cluttered, kinodynamically constrained, and underactuated planning problems with results indicating significant improvements in computation times, path qualities, and success rates over existing methods.

cs.RO

Shrinking braids and Left distributive monoid

We consider a natural generalization of braids which we call shrinking braids. We state the relations of shrinking braids and use them to define algebraically the monoid $R$. We endow a subset of $R$ with a \emph{left distributive monoid} structure and use it to extend the Dehornoy order on $B_{\infty}$ to an order on $R$. By using this order, we prove that $R$ is isomorphic to the monoid which is generated (geometrically) by shrinking braids.

math.GR