SearcharxivSearch

arXiv subjects

Yifan Ma

Publications and source records attributed to Yifan Ma.

16 recordsLinked to original sources

TransDex: Pre-training Visuo-Tactile Policy with Point Cloud Reconstruction for Dexterous Manipulation of Transparent Objects

Dexterous manipulation enables complex tasks but suffers from self-occlusion, severe depth noise, and depth information loss when manipulating transparent objects. To solve this problem, this paper proposes TransDex, a 3D visuo-tactile fusion motor policy based on point cloud reconstruction pre-training. Specifically, we first propose a self-supervised point cloud reconstruction pre-training approach based on Transformer. This method accurately recovers the 3D structure of objects from interactive point clouds of dexterous hands, even when random noise and large-scale masking are added. Building on this, TransDex is constructed in which perceptual encoding adopts a fine-grained hierarchical scheme and multi-round attention mechanisms adaptively fuse features of the robotic arm and dexterous hand to enable differentiated motion prediction. Results from transparent object manipulation experiments conducted on a real robotic system demonstrate that TransDex outperforms existing baseline methods. Further analysis validates the generalization capabilities of TransDex and the effectiveness of its individual components.

cs.RO

Repetition-Rate-Difference Tunable Dual-Comb Fiber Laser Using Bidirectional Lyot filtering

Single cavity dual-comb fiber lasers adopting different multiplexing configurations are benefited from the natures of common-mode noise suppression and superior coherence. Particularly, repetition-rate tunable dual-combs enable non-ambiguous ranging and aliasing-free spectroscopy. However, their sampling rate and spectral resolution is severely restricted by the mechanical delay lines. In a previous work, as rapid as 500 kHz/s tuning rate was realized to address this issue, while the minimum comb frequency difference remained large under the inaccuracy of mechanical DLL. In this work, a dual-comb prototype incorporated with a thermally controlled bidirectional lyot filter is demonstrated with 870-times enhanced tuning precision compared with mechanical schemes. Linear correlation between temperature and repetition-rate-difference of this tuning mechanism is revealed. We achieve a tuning efficiency of 4.4 Hz/{\deg}C and a control accuracy of 0.44 Hz/K, denoting a significant advance in operating Hz-scale differential comb lines. This design offers an optimal playground for extending non-ambiguous distance in dead-zone-free dual-comb ranging and eliminating aliasing in spectroscopy.

physics.optics

Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark

Large models have achieved remarkable performance across various tasks, yet they incur significant computational costs and privacy concerns during both training and inference. Distributed deployment has emerged as a potential solution, but it necessitates the exchange of intermediate information between model segments, with feature representations serving as crucial information carriers. To optimize information exchange, feature coding is required to reduce transmission and storage overhead. Despite its importance, feature coding for large models remains an under-explored area. In this paper, we draw attention to large model feature coding and make three fundamental contributions. First, we introduce a comprehensive dataset encompassing diverse features generated by three representative types of large models. Second, we establish unified test conditions, enabling standardized evaluation pipelines and fair comparisons across future feature coding studies. Third, we introduce two baseline methods derived from widely used image coding techniques and benchmark their performance on the proposed dataset. These contributions aim to provide a foundation for future research and inspire broader engagement in this field. To support a long-term study, all source code and the dataset are made available at \href{https://github.com/chansongoal/LaMoFC}{https://github.com/chansongoal/LaMoFC}.

cs.MM

GHz fundamental mode-locking of a highly integrated Er-doped all-fiber ring laser

High repetition rate ultrafast fiber lasers are important tools for both fundamental science and industry applications. However, achieving over GHz repetition rate in passively mode-locked fiber ring lasers is still challenging. Here, we demonstrate the first ring-cavity Er-doped fiber laser that achieves over GHz fundamental repetition rate by using an all-integration cavity design. In the proposed laser oscillator, all functions are integrated into one device, making it an ultra-compact laser cavity. The laser is mode-locked by carbon nanotubes (CNTs) film that is directly deposited on the pigtail active fiber connectors. The laser produces ultrafast optical pulses at 1562 nm, with a pulse width of 682 fs and a fundamental repetition rate of 1.028 GHz with improved performance. Stable and low-noise mode-locking is characterized by high signal-to-noise ratio (SNR) radiofrequency signal and low relative intensity noise (RIN). The proposed all-integration laser design may serve as a reference for compact fiber ring lasers using other mode-locking mechanisms or at diverse wavelengths.

physics.optics

Deep Learning for Near-Field XL-MIMO Transceiver Design: Principles and Techniques

Massive multiple-input multiple-output (MIMO) has been a critical enabling technology in 5th generation (5G) wireless networks. With the advent of 6G, a natural evolution is to employ even more antennas, potentially an order of magnitude more, to meet the ever-increasing demand for spectral efficiency. This is beyond a mere quantitative scale-up. The enlarged array aperture brings a paradigm shift towards near-field communications, departing from traditional far-field approaches. However, designing advanced transceiver algorithms for near-field systems is extremely challenging because of the enormous system scale, the complicated channel characteristics, and the uncertainties in the propagation environments. Hence, it is important to develop scalable, low-complexity, and robust algorithms that can efficiently characterize and leverage the properties of the near-field channel. In this article, we discuss the principles and advocate two general frameworks to design deep learning-based near-field transceivers covering both iterative and non-iterative algorithms. Case studies on channel estimation and beam focusing are presented to provide a hands-on tutorial. Finally, we discuss open issues and shed light on future directions.

eess.SP

Low-Complexity CSI Feedback for FDD Massive MIMO Systems via Learning to Optimize

In frequency-division duplex (FDD) massive multiple-input multiple-output (MIMO) systems, the growing number of base station antennas leads to prohibitive feedback overhead for downlink channel state information (CSI). To address this challenge, state-of-the-art (SOTA) fully data-driven deep learning (DL)-based CSI feedback schemes have been proposed. However, the high computational complexity and memory requirements of these methods hinder their practical deployment on resource-constrained devices like mobile phones. To solve the problem, we propose a model-driven DL-based CSI feedback approach by integrating the wisdom of compressive sensing and learning to optimize (L2O). Specifically, only a linear learnable projection is adopted at the encoder side to compress the CSI matrix, thereby significantly cutting down the user-side complexity and memory expenditure. On the other hand, the decoder incorporates two specially designed components, i.e., a learnable sparse transformation and an element-wise L2O reconstruction module. The former is developed to learn a sparse basis for CSI within the angular domain, which explores channel sparsity effectively. The latter shares the same long short term memory (LSTM) network across all elements of the optimization variable, eliminating the retraining cost when problem scale changes. Simulation results show that the proposed method achieves a comparable performance with the SOTA CSI feedback scheme but with much-reduced complexity, and enables multiple-rate feedback.

eess.SP

783-MHz fundamental repetition rate all-fiber ring laser mode-locked by carbon nanotubes

We demonstrate a 783-MHz fundamental repetition rate mode-locked Er-doped all-fiber ring laser with a pulse width of 623 fs. By using carbon nanotubes (CNT) saturable absorber (SA), a relatively low self-starting pump threshold of 108 mW is achieved. The laser has a very compact footprint less than 10 cm * 10 cm, benefiting from the all-active-fiber cavity design. The robust mode-locking is confirmed by the low relative intensity noise (RIN) and a long-term stability test. We propose a new scheme for generating high repetition rate femtosecond optical pulses from a compact and stable all-active-fiber ring oscillator.

physics.optics

Mode-Locked Fiber Laser with up to 19 kHz Wavelength Sweep Rate via External Pump LD Modulation

For the first time, we introduce a rapid wavelength-swept, passively mode-locked fiber laser in an all-polarization-maintaining and all-fiber configuration. Achieving an exceptional wavelength sweep rate of up to 19 kHz through external modulation of the LD driver pump current, this laser offers a high sweep rate, simple cavity design, cost-effectiveness, and excellent repeatability.

physics.optics

Rapid-scanned and self-corrected repetition rates enabled in a bidirectional polarization-multiplexed fiber laser

Repetition-rate-scanned lasers are practical in accordion frequency comb generation that serves as a variable gearbox connecting optical and radio wave domains. Rapid and wide-range scanned repetition rate can benefit versatile purposes, however scanning robustness remains unsecured that typically requires complicated feedback loops. Recently, multiplexed lasers have been demonstrated with the nature of common-noise rejection among simultaneously emitted combs. Here, we propose a bidirectional polarization-multiplexed fiber laser that delivers synchronized pulses with rapid-scanned and reference-free repetition rates. Benefiting from the all polarization-maintaining fiber configuration, the laser shows good robustness and inter-comb coherence. As rapid as 493.5 kHz/s scanning rate over 329-kHz scanning range of fundamental repetition rate is realized. The 1-hour and 1-day maximal variations of difference frequency are merely 0.52 Hz and 5.46 Hz. The capability to rebuilt steady state after mode hopping is also demonstrated. These results provide a promising solution for developing high-performance accordion-frequency laser sources.

physics.optics

Pump-power-controlled L-band wavelength-tunable mode-locked fiber laser utilizing all polarization maintaining nonlinear polarization rotation

For the first time, we present the pump power-controlled wavelength-tunable mode-locked fiber laser in the L-band (1565 nm to 1625 nm), achieved by all-polarization maintaining (all-PM) nonlinear polarization rotation (NPR). The wavelength of the laser can be tuned over 20 nm, from 1568.2 nm to 1588.9 nm simply by controlling the pump power from 45 mW to 115 mW. In contrast to conventional wavelength tuning mechanisms such as optical bandpass filters, our tuning method is non-mechanical and electrically controllable, featuring simplicity and cost-effectiveness in a superior all-fiber design.

physics.optics

Lightweight and Flexible Deep Equilibrium Learning for CSI Feedback in FDD Massive MIMO

In frequency-division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems, downlink channel state information (CSI) needs to be sent back to the base station (BS) by the users, which causes prohibitive feedback overhead. In this paper, we propose a lightweight and flexible deep learning-based CSI feedback approach by capitalizing on deep equilibrium models. Different from existing deep learning-based methods that stack multiple explicit layers, we propose an implicit equilibrium block to mimic the behavior of an infinite-depth neural network. In particular, the implicit equilibrium block is defined by a fixed-point iteration and the trainable parameters in different iterations are shared, which results in a lightweight model. Furthermore, the number of forward iterations can be adjusted according to users' computation capability, enabling a flexible accuracy-efficiency trade-off. Simulation results will show that the proposed design obtains a comparable performance as the benchmarks but with much-reduced complexity and permits an accuracy-efficiency trade-off at runtime.

cs.IT

Augmented Deep Unfolding for Downlink Beamforming in Multi-cell Massive MIMO With Limited Feedback

In limited feedback multi-user multiple-input multiple-output (MU-MIMO) cellular networks, users send quantized information about the channel conditions to the associated base station (BS) for downlink beamforming. However, channel quantization and beamforming have been treated as two separate tasks conventionally, which makes it difficult to achieve global system optimality. In this paper, we propose an augmented deep unfolding (ADU) approach that jointly optimizes the beamforming scheme at the BSs and the channel quantization scheme at the users. In particular, the classic WMMSE beamformer is unrolled and a deep neural network (DNN) is leveraged to pre-process its input to enhance the performance. The variational information bottleneck technique is adopted to further improve the performance when the feedback capacity is strictly restricted. Simulation results demonstrate that the proposed ADU method outperforms all the benchmark schemes in terms of the system average rate.

eess.SP

Learn to Communicate with Neural Calibration: Scalability and Generalization

The conventional design of wireless communication systems typically relies on established mathematical models that capture the characteristics of different communication modules. Unfortunately, such design cannot be easily and directly applied to future wireless networks, which will be characterized by large-scale ultra-dense networks whose design complexity scales exponentially with the network size. Furthermore, such networks will vary dynamically in a significant way, which makes it intractable to develop comprehensive analytical models. Recently, deep learning-based approaches have emerged as potential alternatives for designing complex and dynamic wireless systems. However, existing learning-based methods have limited capabilities to scale with the problem size and to generalize with varying network settings. In this paper, we propose a scalable and generalizable neural calibration framework for future wireless system design, where a neural network is adopted to calibrate the input of conventional model-based algorithms. Specifically, the backbone of a traditional time-efficient algorithm is integrated with deep neural networks to achieve a high computational efficiency, while enjoying enhanced performance. The permutation equivariance property, carried out by the topological structure of wireless systems, is furthermore utilized to develop a generalizable neural network architecture. The proposed neural calibration framework is applied to solve challenging resource management problems in massive multiple-input multiple-output (MIMO) systems. Simulation results will show that the proposed neural calibration approach enjoys significantly improved scalability and generalization compared with the existing learning-based methods.

eess.SP

Neural Calibration for Scalable Beamforming in FDD Massive MIMO with Implicit Channel Estimation

Channel estimation and beamforming play critical roles in frequency-division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems. However, these two modules have been treated as two stand-alone components, which makes it difficult to achieve a global system optimality. In this paper, we propose a deep learning-based approach that directly optimizes the beamformers at the base station according to the received uplink pilots, thereby, bypassing the explicit channel estimation. Different from the existing fully data-driven approach where all the modules are replaced by deep neural networks (DNNs), a neural calibration method is proposed to improve the scalability of the end-to-end design. In particular, the backbone of conventional time-efficient algorithms, i.e., the least-squares (LS) channel estimator and the zero-forcing (ZF) beamformer, is preserved and DNNs are leveraged to calibrate their inputs for better performance. The permutation equivariance property of the formulated resource allocation problem is then identified to design a low-complexity neural network architecture. Simulation results will show the superiority of the proposed neural calibration method over benchmark schemes in terms of both the spectral efficiency and scalability in large-scale wireless networks.

eess.SP

Automated Prostate Cancer Diagnosis Based on Gleason Grading Using Convolutional Neural Network

The Gleason grading system using histological images is the most powerful diagnostic and prognostic predictor of prostate cancer. The current standard inspection is evaluating Gleason H&E-stained histopathology images by pathologists. However, it is complicated, time-consuming, and subject to observers. Deep learning (DL) based-methods that automatically learn image features and achieve higher generalization ability have attracted significant attention. However, challenges remain especially using DL to train the whole slide image (WSI), a predominant clinical source in the current diagnostic setting, containing billions of pixels, morphological heterogeneity, and artifacts. Hence, we proposed a convolutional neural network (CNN)-based automatic classification method for accurate grading of PCa using whole slide histopathology images. In this paper, a data augmentation method named Patch-Based Image Reconstruction (PBIR) was proposed to reduce the high resolution and increase the diversity of WSIs. In addition, a distribution correction (DC) module was developed to enhance the adaption of pretrained model to the target dataset by adjusting the data distribution. Besides, a Quadratic Weighted Mean Square Error (QWMSE) function was presented to reduce the misdiagnosis caused by equal Euclidean distances. Our experiments indicated the combination of PBIR, DC, and QWMSE function was necessary for achieving superior expert-level performance, leading to the best results (0.8885 quadratic-weighted kappa coefficient).

eess.IV

A Low-Complexity Algorithmic Framework for Large-Scale IRS-Assisted Wireless Systems

Intelligent reflecting surfaces (IRSs) are revolutionary enablers for next-generation wireless communication networks, with the ability to customize the radio propagation environment. To fully exploit the potential of IRS-assisted wireless systems, reflective elements have to be jointly optimized with conventional communication techniques. However, the resulting optimization problems pose significant algorithmic challenges, mainly due to the large-scale non-convex constraints induced by the passive hardware implementations. In this paper, we propose a low-complexity algorithmic framework incorporating alternating optimization and gradient-based methods for large-scale IRS-assisted wireless systems. The proposed algorithm provably converges to a stationary point of the optimization problem. Extensive simulation results demonstrate that the proposed framework provides significant speedups compared with existing algorithms, while achieving a comparable or better performance.

cs.IT