SearcharxivSearch

arXiv subjects

Yudi Zhao

Publications and source records attributed to Yudi Zhao.

5 recordsLinked to original sources

Photonic-chip-based generation of sub-100-femtosecond optical frequency combs

Sub-100-fs optical pulses and frequency comb sources have been revolutionizing a wide range of applications, from ultrafast optical science to optical frequency standard and measurement. To date, the leading techniques for generating such pulses in practical systems rely on tabletop mode-locked lasers, which inherently suffer from high system complexity, limited long-term reliability, and pronounced environmental sensitivity. Meanwhile, driven by advances in photonic integration, chip-scale approaches have sought to realize miniaturized pulse sources. However, simultaneously achieving sub-100-fs duration, ideal pulse shape, and a broadband flat-topped spectrum remains a significant challenge. Here, we address these challenges by combining two key photonic chip technologies: TFLN EO modulators for picosecond seed pulse generation, and highly nonlinear optical loop mirrors (NOLM) based on AlGaAsOI nanowaveguides for efficient temporal pulse cleaning and spectral broadening. In theoretical simulation and experiment, we show that for an input seed pulse centred at ~1550nm, a single-stage AlGaAs NOLM with a loop length of 1cm can produce flat-topped, nearly tenfold spectral broadening and over tenfold compression of pulse width, and more than 10dB suppression of pulse pedestals. Using initial EO comb pulses with ps-level durations at repetition rates of 10-20GHz, we demonstrate photonic-chip-enabled pulses with an unprecedented duration of 55fs and a flat-topped comb spectrum whose 10dB optical bandwidth exceeds 90nm. Our results highlight the remarkable potential of photonic chip technologies to realize high-repetition-rate, miniaturized sub-100-fs optical pulse generators with the prospect of superior stability and operability. The demonstrated photonic-chip-based sub-100-fs optical frequency comb sources may establish a new paradigm for both scientific research and practical applications.

physics.optics

Higher-Order Flexible Configurations of Planar Parallel Manipulators Constructed by Averaging

This paper investigates singular configurations of planar 3-RPR parallel manipulators, which result from applying the averaging technique to solution pairs of their direct kinematic problem. Without computing the zeros of the corresponding degree 6 polynomial we parametrize the input pairs and determine their relative orientation in a way that the flexion order of the averaged configurations increases. Moreover, the obtained results are visualized for concrete examples. The presented methodology can also be used for studying the spherical and spatial analogues of planar 3-RPR parallel manipulators.

cs.RO

Self-supervised Implicit Glyph Attention for Text Recognition

The attention mechanism has become the \emph{de facto} module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e., implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SIGA). SIGA delineates the glyph structures of text images by jointly self-supervised text segmentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks.

cs.CV

Distribution Learning Based on Evolutionary Algorithm Assisted Deep Neural Networks for Imbalanced Image Classification

To address the trade-off problem of quality-diversity for the generated images in imbalanced classification tasks, we research on over-sampling based methods at the feature level instead of the data level and focus on searching the latent feature space for optimal distributions. On this basis, we propose an iMproved Estimation Distribution Algorithm based Latent featUre Distribution Evolution (MEDA_LUDE) algorithm, where a joint learning procedure is programmed to make the latent features both optimized and evolved by the deep neural networks and the evolutionary algorithm, respectively. We explore the effect of the Large-margin Gaussian Mixture (L-GM) loss function on distribution learning and design a specialized fitness function based on the similarities among samples to increase diversity. Extensive experiments on benchmark based imbalanced datasets validate the effectiveness of our proposed algorithm, which can generate images with both quality and diversity. Furthermore, the MEDA_LUDE algorithm is also applied to the industrial field and successfully alleviates the imbalanced issue in fabric defect classification.

cs.CV

Visual Sensation and Perception Computational Models for Deep Learning: State of the art, Challenges and Prospects

Visual sensation and perception refers to the process of sensing, organizing, identifying, and interpreting visual information in environmental awareness and understanding. Computational models inspired by visual perception have the characteristics of complexity and diversity, as they come from many subjects such as cognition science, information science, and artificial intelligence. In this paper, visual perception computational models oriented deep learning are investigated from the biological visual mechanism and computational vision theory systematically. Then, some points of view about the prospects of the visual perception computational models are presented. Finally, this paper also summarizes the current challenges of visual perception and predicts its future development trends. Through this survey, it will provide a comprehensive reference for research in this direction.

cs.AI