SearcharxivSearch

arXiv subjects

Jin Wei

Publications and source records attributed to Jin Wei.

At least 19 recordsLinked to original sources

Active Passivation Tunes Hotspot Locations in GaN Transistors with In Situ Thermal Mechanical Visualization

Efficient thermal dissipation has become critical in emerging electronic devices. However, most existing studies have primarily focused on engineering heat dissipation pathways, largely overlooking the intrinsic behavior of the heat source itself. We demonstrate an active passivation technology that proactively tunes hotspot locations in GaN transistors. By adjusting the active passivation layer length, the hotspot is shifted from the gate edge to the drain-side AP edge, establishing a clear one-to-one spatial correlation. In-situ thermal-mechanical visualization via micro-Raman thermography, combined with multi-physics electro-thermal-mechanical simulations, directly captures the spatial redistribution of both temperature and thermal stress profiles. Electrical analysis confirms that this hotspot migration is driven by the spatial shift of the peak electric field and localized Joule heating. This proactive heat-source tuning strategy provides critical design guidelines for power electronics.

physics.app-ph

Logical Metatheorems for Abstract Spaces axiomatized in Positive Bounded Logic II: Metric spaces and the model-theoretic uniformity principle

We extend the proof-theoretic treatment of uniform bound extraction from normed structures axiomatized in positive bounded logic [Advances in Mathematics, 290:503-551, 2016] (as developed for the model theory of Banach spaces) to the more general setting of abstract metric structures, including discrete structures viewed as classical first-order models. In particular, we establish uniform bound extraction theorems for our generalized framework for $\forall\exists$-sentences whose matrix is the negation of (an embedding of) a formula in positive bounded logic, whose proofs use saturation. In this way, we provide a formal explanation for the successes in the extraction of uniform bounds from nonstandard proofs given in [Advances in Mathematics, 343:567-623, 2019], which had informally followed the perspective of the monotone functional interpretation. As an application of the formal framework we develop, we provide novel explicit bounds for a structural theorem for stable subsets of groups given in [Mathematical Proceedings of the Cambridge Philosophical Society, 168(2):405-413, 2020].

math.LO

Empirical verification of principal mode orthogonality and relative phase calibration in photonic lanterns

Photonic lanterns efficiently map input spatial modes to single-mode outputs for applications like high angular resolution imaging and nulling interferometry. However, manufacturing limits prevent full control over the device's mode transfer matrix at the design stage, making empirical characterisation essential. In this work we further analyse a dataset of direct measurements of a photonic lantern's principal modes using digital off-axis holography over a 73 nm range near 1550 nm. By analysing the electric field directly, we find that the principal modes are significantly more orthogonal than random vectors in a space of the same size, as expected for near-adiabatic devices. We propose metrics for quantifying this effect, noting that mode converters with orthogonal principal modes provide better conditioned inverse solvers. We also simulate additional measurements that characterisation systems could take, where the orthogonality would be leveraged to determine the relative phase between principal modes.

physics.optics

Seidr update: photonic 'black magic' for high-contrast interferometry using kernel-nulling and photonic lanterns

Seidr is a new interferometric beam combiner within the Asgard Suite, utilizing infrastructure common to the BIFROST instrument at the Very Large Telescope Interferometer. Seidr combines hybrid mode-selective photonic lantern injection modules with a kernel-nulling photonic chip backend to enable deep H-band nulling for high-contrast studies of exoplanets, exomoons, and circumstellar dust. This instrument update summarizes Seidr's current design maturity and recent simulations of the point source - to - lantern outputs. We also outline progress on our neural network-based wavefront estimation scheme, which uses the photonic lantern outputs to sense phase fluctuations, designed to feed back to Baldr's deformable mirror, and improve nuller light injection.

physics.optics

Overcoming the low signal-to-noise problem for hybrid mode-selective photonic lantern-based wavefront correction using machine learning

Hybrid mode-selective photonic lanterns transform an input complex point-spread function into several single-mode outputs, where a selected core feeds the fundamental mode to a photonic science instrument, while the remaining cores are used for wavefront sensing in a closed-loop adaptive optics system. A neural network maps the intensities of the wavefront sensing cores to an estimated wavefront correction, which is applied to an upstream deformable mirror. However, there exists a trade between maximizing the amount of light reserved for the photonic instrument and the reduced signal-to-noise ratios for the wavefront sensing cores. We explore wavefront correction for the Seidr instrument, a part of the Asgard Suite for the Very Large Telescope Interferometer. We evaluate different neural network architectures, comparing wavefront estimation performance for different wavefront error types, as a first step toward addressing the signal-to-noise trade-off. Results show transformer neural networks as a promising solution for temporal photonic lantern-based wavefront estimation.

physics.optics

The light carrying orbital angular momentum through time-varying scattering media using dual orthogonal-polarization channels

Orbital angular momentum (OAM) of light can be used as a degree of freedom for optical communication. However, the reliable transmission of OAM beams through time-varying scattering media remains challenging. In this paper, we report an approach for OAM transmission through time-varying scattering media using dual orthogonal-polarization channels. A perfect vortex beam (PVB) is generated and transmitted at one channel, and the plane wave is employed as a reference beam at another channel. With the cross-correlation obtained between the recorded speckle patterns at each single-shot measurement, the data can be retrieved without prior knowledge about time-varying scattering media. It is revealed that the multiplexed PVB can be applied to generate recognizable coherent-superposition patterns in cross-correlation images. Experimental results demonstrate that the proposed method can be applied to transmit the data through time-varying scattering media, and high robustness can be achieved. The proposed method could open up an avenue for the development of OAM transmission in harsh environments.

physics.optics

Finding New Connections between Concepts from Medline Database Incorporating Domain Knowledge

In this digital world, data is everything and significantly impacts our everyday lives. Interestingly, in this small world, everything is part of an ecosystem, where everything is connected, directly or indirectly. The same thing happens to data as well. In most cases, it may seem like a particular topic does not have any connection with another one, but in reality, they are connected through a mutually related topic. Therefore, in this research, we will discuss an adaptive model modified from the ABC model by Don R. Swanson, a Literature-Based Discovery (LBD) Model, to find the hidden connections between Concepts of Interest. The model demonstrates that two topics, A and C are different and have no relationship. But they have a common topic, B that can be used to connect topics A and C This famous model will be used in this discussion to connect Medical Concepts.

cs.CL

Illuminating the lantern: coherent, spectro-polarimetric characterisation of a multimode converter

While photonic lanterns efficiently and uniquely map a set of input modes to single-mode outputs (or vice versa), the optical mode transfer matrix of any particular fabricated device cannot be constrained at the design stage due to manufacturing imperfections. Accurate knowledge of the mapping enables complex sensing or beam control applications that leverage multimode conversion. In this work, we present a characterisation system to directly measure the electric field from a photonic lantern using digital off-axis holography, following its evolution over a 73 nm range near 1550 nm and in two orthogonal, linear polarisations. We provide the first multi-wavelength, polarisation decomposed characterisation of the principal modes of a photonic lantern. Performance of our testbed is validated on a single-mode fibre then harnessed to characterise a 19-port, multicore fibre fed photonic lantern. We uncover the typical wavelength scale at which the modal mapping evolves and measure the relative dispersion in the device, finding significant differences with idealised simulations. In addition to detailing the system, we also share the empirical mode transfer matrices, enabling future work in astrophotonic design, computational imaging, device fabrication feedback loops and beam shaping.

physics.optics

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Although longer captions convey richer information, existing models struggle to process them effectively because of limited text-encoding capacity, and there remains a shortage of resources that align remote sensing images with both short text and long text captions. To address this gap, we introduce DGTRSD, a dual-granularity remote sensing image-text dataset, where each image is paired with both a short text caption and a long text description, providing a solid foundation for dual-granularity semantic modeling. Based on this, we further propose DGTRS-CLIP, a dual-granularity curriculum learning framework that combines short text and long text supervision to achieve dual-granularity semantic alignment. Extensive experiments on four typical zero-shot tasks: long text cross-modal retrieval, short text cross-modal retrieval, image classification, and semantic localization demonstrate that DGTRS-CLIP consistently outperforms existing methods across all tasks. The code has been open-sourced and is available at https://github.com/MitsuiChen14/DGTRS.

cs.CV

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information.

cs.CV

Approximate Completeness of Hypersequent Calculus for First-Order {\L}ukasiewicz Logic

Hypersequent calculus G{\L}$\forall$ for first-order {\L}ukasiewicz logic was first introduced by Baaz and Metcalfe, along with a proof of its approximate completeness with respect to standard $[0,1]$-semantics. The completeness result was later pointed out by Gerasimov that it only applies to prenex formulas. In this paper, we will present our proof of approximate completeness of G{\L}$\forall$ for arbitrary first-order formulas by generalizing the original completeness proof to hypersequents.

math.LO

Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval

Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be publicly available at https://github.com/Gyann-z/FDP.

cs.CV

SQLaser: Detecting DBMS Logic Bugs with Clause-Guided Fuzzing

Database Management Systems (DBMSs) are vital components in modern data-driven systems. Their complexity often leads to logic bugs, which are implementation errors within the DBMSs that can lead to incorrect query results, data exposure, unauthorized access, etc., without necessarily causing visible system failures. Existing detection employs two strategies: rule-based bug detection and coverage-guided fuzzing. In general, rule specification itself is challenging; as a result, rule-based detection is limited to specific and simple rules. Coverage-guided fuzzing blindly explores code paths or blocks, many of which are unlikely to contain logic bugs; therefore, this strategy is cost-ineffective. In this paper, we design SQLaser, a SQL-clause-guided fuzzer for detecting logic bugs in DBMSs. Through a comprehensive examination of most existing logic bugs across four distinct DBMSs, excluding those causing system crashes, we have identified 35 logic bug patterns. These patterns manifest as certain SQL clause combinations that commonly result in logic bugs, and behind these clause combinations are a sequence of functions. We therefore model logic bug patterns as error-prone function chains (ie, sequences of functions). We further develop a directed fuzzer with a new path-to-path distance-calculation mechanism for effectively testing these chains and discovering additional logic bugs. This mechanism enables SQLaser to swiftly navigate to target sites and uncover potential bugs emerging from these paths. Our evaluation, conducted on SQLite, MySQL, PostgreSQL, and TiDB, demonstrates that SQLaser significantly accelerates bug discovery compared to other fuzzing approaches, reducing detection time by approximately 60%.

cs.CR

HuntFUZZ: Enhancing Error Handling Testing through Clustering Based Fuzzing

Testing a program's capability to effectively handling errors is a significant challenge, given that program errors are relatively uncommon. To solve this, Software Fault Injection (SFI)-based fuzzing integrates SFI and traditional fuzzing, injecting and triggering errors for testing (error handling) code. However, we observe that current SFI-based fuzzing approaches have overlooked the correlation between paths housing error points. In fact, the execution paths of error points often share common paths. Nonetheless, Fuzzers usually generate test cases repeatedly to test error points on commonly traversed paths. This practice can compromise the efficiency of the fuzzer(s). Thus, this paper introduces HuntFUZZ, a novel SFI-based fuzzing framework that addresses the issue of redundant testing of error points with correlated paths. Specifically, HuntFUZZ clusters these correlated error points and utilizes concolic execution to compute constraints only for common paths within each cluster. By doing so, we provide the fuzzer with efficient test cases to explore related error points with minimal redundancy. We evaluate HuntFUZZ on a diverse set of 42 applications, and HuntFUZZ successfully reveals 162 known bugs, with 62 of them being related to error handling. Additionally, due to its efficient error point detection method, HuntFUZZ discovers 7 unique zero-day bugs, which are all missed by existing fuzzers. Furthermore, we compare HuntFUZZ with 4 existing fuzzing approaches, including AFL, AFL++, AFLGo, and EH-FUZZ. Our evaluation confirms that HuntFUZZ can cover a broader range of error points, and it exhibits better performance in terms of bug finding speed.

cs.CR

TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model

Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: 1) "Can machines spot texts without precise detection just like human beings?", and if yes, 2) "Is text block another alternative for scene text spotting other than word or character?" To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we finetune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs).

cs.CV

RESURF Ga$_{2}$O$_{3}$-on-SiC Field Effect Transistors for Enhanced Breakdown Voltage

Heterosubstrates have been extensively studied as a method to improve the heat dissipation of Ga$_{2}$O$_{3}$ devices. In this simulation work, we propose a novel role for $p$-type available heterosubstrates, as a component of a reduced surface field (RESURF) structure in Ga$_{2}$O$_{3}$ lateral field-effect transistors (FETs). The RESURF structure can eliminate the electric field crowding and contribute to higher breakdown voltage. Using SiC as an example, the designing strategy for doping concentration and dimensions of the $p$-type region is systematically studied using TCAD modeling. To mimic realistic devices, the impacts of interface charge and binding interlayer at the Ga$_{2}$O$_{3}$/SiC interface are also explored. Additionally, the feasibility of the RESURF structure for high-frequency switching operation is supported by the short time constant ($\sim$0.5 ns) of charging/discharging the $p$-SiC depletion region. This study demonstrates the great potential of utilizing the electrical properties of heat-dissipating heterosubstrates to achieve a uniform electric field distribution in Ga$_{2}$O$_{3}$ FETs.

physics.app-ph

Demonstration of a photonic lantern focal-plane wavefront sensor: measurement of atmospheric wavefront error modes and low wind effect in the non-linear regime

Here we present a laboratory analysis of the use of a 19-core photonic lantern (PL) in combination with neural network (NN) algorithms as an efficient focal plane wavefront sensor (FP-WFS) for adaptive optics (AO), measuring wavefront errors such as low wind effect (LWE), Zernike modes and Kolmogorov phase maps. The aberrated wavefronts were experimentally simulated using a Spatial Light Modulator (SLM) with combinations of different phase maps in both the linear regime (average incident RMS wavefront error (WFE) of 0.88 rad) and in the non-linear regime (average incident RMS WFE of 1.5 rad). Results were analysed using a NN to determine the transfer function of the relationship between the incident wavefront error (WFE) at the input modes at the multimode input of the PL and the intensity distribution output at the multicore fibre outputs end of the PL. The root mean square error (RMSE) of the reconstruction of petal and LWE modes were just $2.87\times10^{-2}$ rad and $2.07\times10^{-1}$ rad respectively, in the non-linear regime. The reconstruction RMSE for Zernike combinations ranged from $5.67\times10^{-2}$ rad to $8.43\times10^{-1}$ rad, depending on the number of Zernike terms and incident RMS WFE employed. These results demonstrate the promising potential of PLs as an innovative FP-WFS in conjunction with NNs.

physics.optics

Masked and Permuted Implicit Context Learning for Scene Text Recognition

Scene Text Recognition (STR) is difficult because of the variations in text styles, shapes, and backgrounds. Though the integration of linguistic information enhances models' performance, existing methods based on either permuted language modeling (PLM) or masked language modeling (MLM) have their pitfalls. PLM's autoregressive decoding lacks foresight into subsequent characters, while MLM overlooks inter-character dependencies. Addressing these problems, we propose a masked and permuted implicit context learning network for STR, which unifies PLM and MLM within a single decoder, inheriting the advantages of both approaches. We utilize the training procedure of PLM, and to integrate MLM, we incorporate word length information into the decoding process and replace the undetermined characters with mask tokens. Besides, perturbation training is employed to train a more robust model against potential length prediction errors. Our empirical evaluations demonstrate the performance of our model. It not only achieves superior performance on the common benchmarks but also achieves a substantial improvement of $9.1\%$ on the more challenging Union14M-Benchmark.

cs.CV