SearcharxivSearch

arXiv subjects

Mingxing Li

Publications and source records attributed to Mingxing Li.

13 recordsLinked to original sources

Searth Transformer: A Transformer Architecture Incorporating Earth's Geospheric Physical Priors for Global Mid-Range Weather Forecasting

Accurate global medium-range weather forecasting is fundamental to Earth system science. Most existing Transformer-based forecasting models adopt vision-centric architectures that neglect the Earth's spherical geometry and zonal periodicity. In addition, conventional autoregressive training is computationally expensive and limits forecast horizons due to error accumulation. To address these challenges, we propose the Shifted Earth Transformer (Searth Transformer), a physics-informed architecture that incorporates zonal periodicity and meridional boundaries into window-based self-attention for physically consistent global information exchange. We further introduce a Relay Autoregressive (RAR) fine-tuning strategy that enables learning long-range atmospheric evolution under constrained memory and computational budgets. Based on these methods, we develop YanTian, a global medium-range weather forecasting model. YanTian achieves higher accuracy than the high-resolution forecast of the European Centre for Medium-Range Weather Forecasts and performs competitively with state-of-the-art AI models at one-degree resolution, while requiring roughly 200 times lower computational cost than standard autoregressive fine-tuning. Furthermore, YanTian attains a longer skillful forecast lead time for Z500 (10.3 days) than HRES (9 days). Beyond weather forecasting, this work establishes a robust algorithmic foundation for predictive modeling of complex global-scale geophysical circulation systems, offering new pathways for Earth system science.

cs.LG

From Editor to Dense Geometry Estimator

Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image task, suggesting that image editing models, rather than T2I generative models, may be a more suitable foundation for fine-tuning. Motivated by this, we conduct a systematic analysis of the fine-tuning behaviors of both editors and generators for dense geometry estimation. Our findings show that editing models possess inherent structural priors, which enable them to converge more stably by ``refining" their innate features, and ultimately achieve higher performance than their generative counterparts. Based on these findings, we introduce \textbf{FE2E}, a framework that pioneeringly adapts an advanced editing model based on Diffusion Transformer (DiT) architecture for dense geometry prediction. Specifically, to tailor the editor for this deterministic task, we reformulate the editor's original flow matching loss into the ``consistent velocity" training objective. And we use logarithmic quantization to resolve the precision conflict between the editor's native BFloat16 format and the high precision demand of our tasks. Additionally, we leverage the DiT's global attention for a cost-free joint estimation of depth and normals in a single forward pass, enabling their supervisory signals to mutually enhance each other. Without scaling up the training data, FE2E achieves impressive performance improvements in zero-shot monocular depth and normal estimation across multiple datasets. Notably, it achieves over 35\% performance gains on the ETH3D dataset and outperforms the DepthAnything series, which is trained on 100$\times$ data. The project page can be accessed \href{https://amap-ml.github.io/FE2E/}{here}.

cs.CV

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Traditional visual grounding methods primarily focus on single-image scenarios with simple textual references. However, extending these methods to real-world scenarios that involve implicit and complex instructions, particularly in conjunction with multiple images, poses significant challenges, which is mainly due to the lack of advanced reasoning ability across diverse multi-modal contexts. In this work, we aim to address the more practical universal grounding task, and propose UniVG-R1, a reasoning guided multimodal large language model (MLLM) for universal visual grounding, which enhances reasoning capabilities through reinforcement learning (RL) combined with cold-start data. Specifically, we first construct a high-quality Chain-of-Thought (CoT) grounding dataset, annotated with detailed reasoning chains, to guide the model towards correct reasoning paths via supervised fine-tuning. Subsequently, we perform rule-based reinforcement learning to encourage the model to identify correct reasoning chains, thereby incentivizing its reasoning capabilities. In addition, we identify a difficulty bias arising from the prevalence of easy samples as RL training progresses, and we propose a difficulty-aware weight adjustment strategy to further strengthen the performance. Experimental results demonstrate the effectiveness of UniVG-R1, which achieves state-of-the-art performance on MIG-Bench with a 9.1% improvement over the previous method. Furthermore, our model exhibits strong generalizability, achieving an average improvement of 23.4% in zero-shot performance across four image and video reasoning grounding benchmarks. The project page can be accessed at https://amap-ml.github.io/UniVG-R1-page/.

cs.CV

FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing

Scene text editing aims to modify or add texts on images while ensuring text fidelity and overall visual quality consistent with the background. Recent methods are primarily built on UNet-based diffusion models, which have improved scene text editing results, but still struggle with complex glyph structures, especially for non-Latin ones (\eg, Chinese, Korean, Japanese). To address these issues, we present \textbf{FLUX-Text}, a simple and advanced multilingual scene text editing DiT method. Specifically, our FLUX-Text enhances glyph understanding and generation through lightweight Visual and Text Embedding Modules, while preserving the original generative capability of FLUX. We further propose a Regional Text Perceptual Loss tailored for text regions, along with a matching two-stage training strategy to better balance text editing and overall image quality. Benefiting from the DiT-based architecture and lightweight feature injection modules, FLUX-Text can be trained with only $0.1$M training examples, a \textbf{97\%} reduction compared to $2.9$M required by popular methods. Extensive experiments on multiple public datasets, including English and Chinese benchmarks, demonstrate that our method surpasses other methods in visual quality and text fidelity. All the code is available at https://github.com/AMAP-ML/FluxText.

cs.CV

Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model

The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models (MLLMs) have shown great potential in image quality assessment (IQA) and image aesthetic assessment (IAA). Despite this progress, effectively scoring the quality and aesthetics of UGC images still faces two main challenges: 1) A single score is inadequate to capture the hierarchical human perception. 2) How to use MLLMs to output numerical scores, such as mean opinion scores (MOS), remains an open question. To address these challenges, we introduce a novel dataset, named Realistic image Quality and Aesthetic (RealQA), including 14,715 UGC images, each of which is annoted with 10 fine-grained attributes. These attributes span three levels: low level (e.g., image clarity), middle level (e.g., subject integrity) and high level (e.g., composition). Besides, we conduct a series of in-depth and comprehensive investigations into how to effectively predict numerical scores using MLLMs. Surprisingly, by predicting just two extra significant digits, the next token paradigm can achieve SOTA performance. Furthermore, with the help of chain of thought (CoT) combined with the learnt fine-grained attributes, the proposed method can outperform SOTA methods on five public datasets for IQA and IAA with superior interpretability and show strong zero-shot generalization for video quality assessment (VQA). The code and dataset will be released.

cs.CV

Layer Dependent Thermal Transport Properties of One- to Three-Layer Magnetic Fe:MoS2

Two-Dimensional (2D) transition metal dichalcogenides (TMDs) have been the subject of extensive attention thanks to their unique properties and atomically thin structure. Because of its unprecedented room-temperature magnetic properties, iron-doped MoS2 (Fe:MoS2) is considered the next-generation quantum and magnetic material. It is essential to understand Fe:MoS2's thermal behavior since temperature and thermal load/activation are crucial for their magnetic properties and the current nano and quantum devices have been severely limited by thermal management. In this work, Fe:MoS2 is synthesized by doping Fe atoms into MoS2 using the chemical vapor deposition (CVD) synthesis and a refined version of opto-thermal Raman technique is used to study the thermal transport properties of Fe:MoS2 in the forms of single (1L), bilayer (2L), and tri-layer (3L). In the Opto-thermal Raman technique, a laser is focused on the center of a thin film and used to measure the peak position of a Raman-active mode. The lateral thermal conductivity of 1-3L of Fe:MoS2 and the interfacial thermal conductance between Fe:MoS2 and the substrate were obtained by analyzing the temperature-dependent and power-dependent Raman measurement, laser power absorption coefficient, and laser spot sizes. We also characterized Fe:MoS2's thermal transport at high temperature, and calculated Fe:MoS2's thermal transport by density theory function. These findings will shed light on the thermal management and thermoelectric designs for Fe:MoS2 based nano and quantum electronic devices.

cond-mat.mtrl-sci

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi-center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and postprocessing (66%). The "typical" lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

cs.CV

Advanced Deep Networks for 3D Mitochondria Instance Segmentation

Mitochondria instance segmentation from electron microscopy (EM) images has seen notable progress since the introduction of deep learning methods. In this paper, we propose two advanced deep networks, named Res-UNet-R and Res-UNet-H, for 3D mitochondria instance segmentation from Rat and Human samples. Specifically, we design a simple yet effective anisotropic convolution block and deploy a multi-scale training strategy, which together boost the segmentation performance. Moreover, we enhance the generalizability of the trained models on the test set by adding a denoising operation as pre-processing. In the Large-scale 3D Mitochondria Instance Segmentation Challenge at ISBI 2021, our method ranks the 1st place. Code is available at https://github.com/Limingxing00/MitoEM2021-Challenge.

cs.CV

Recurrent Dynamic Embedding for Video Object Segmentation

Space-time memory (STM) based video object segmentation (VOS) networks usually keep increasing memory bank every several frames, which shows excellent performance. However, 1) the hardware cannot withstand the ever-increasing memory requirements as the video length increases. 2) Storing lots of information inevitably introduces lots of noise, which is not conducive to reading the most important information from the memory bank. In this paper, we propose a Recurrent Dynamic Embedding (RDE) to build a memory bank of constant size. Specifically, we explicitly generate and update RDE by the proposed Spatio-temporal Aggregation Module (SAM), which exploits the cue of historical information. To avoid error accumulation owing to the recurrent usage of SAM, we propose an unbiased guidance loss during the training stage, which makes SAM more robust in long videos. Moreover, the predicted masks in the memory bank are inaccurate due to the inaccurate network inference, which affects the segmentation of the query frame. To address this problem, we design a novel self-correction strategy so that the network can repair the embeddings of masks with different qualities in the memory bank. Extensive experiments show our method achieves the best tradeoff between performance and speed. Code is available at https://github.com/Limingxing00/RDE-VOS-CVPR2022.

cs.CV

Retinal Vessel Segmentation with Pixel-wise Adaptive Filters

Accurate retinal vessel segmentation is challenging because of the complex texture of retinal vessels and low imaging contrast. Previous methods generally refine segmentation results by cascading multiple deep networks, which are time-consuming and inefficient. In this paper, we propose two novel methods to address these challenges. First, we devise a light-weight module, named multi-scale residual similarity gathering (MRSG), to generate pixel-wise adaptive filters (PA-Filters). Different from cascading multiple deep networks, only one PA-Filter layer can improve the segmentation results. Second, we introduce a response cue erasing (RCE) strategy to enhance the segmentation accuracy. Experimental results on the DRIVE, CHASE_DB1, and STARE datasets demonstrate that our proposed method outperforms state-of-the-art methods while maintaining a compact structure. Code is available at https://github.com/Limingxing00/Retinal-Vessel-Segmentation-ISBI20222.

eess.IV

Investigation of photon emitters in Ce-implanted hexagonal boron nitride

Color centers in hexagonal boron nitride (hBN) are presently attracting broad interest as a novel platform for nanoscale sensing and quantum information processing. Unfortunately, their atomic structures remain largely elusive and only a small percentage of the emitters studied thus far has the properties required to serve as optically addressable spin qubits. Here, we use confocal fluorescence microscopy at variable temperature to study a new class of point defects produced via cerium ion implantation in thin hBN flakes. We find that, to a significant fraction, emitters show bright room-temperature emission, and good optical stability suggesting the formation of Ce-based point defects. Using density functional theory (DFT) we calculate the emission properties of candidate emitters, and single out the CeVB center - formed by an interlayer Ce atom adjacent to a boron vacancy - as one possible microscopic model. Our results suggest an intriguing route to defect engineering that simultaneously exploits the singular properties of rare-earth ions and the versatility of two-dimensional material hosts.

cond-mat.mes-hall

Spontaneous emission dynamics of $Eu^{ 3+}$ ions coupled to hyperbolic metamaterials

Sub-wavelength nanostructured systems with tunable electromagnetic properties, such as hyperbolic metamaterials (HMMs), provide a useful platform to tailor spontaneous emission processes. Here, we investigate a system comprising $Eu^{ 3+}(NO_{3})_{3}6H_{2}O$ nanocrystals on an HMM structure featuring a hexagonal array of Ag-nanowires in a porous $Al_{2}O_{3}$ matrix. The HMM-coupled $Eu^{ 3+}$ ions exhibit up to a 2.4-fold increase of their decay rate, accompanied by an enhancement of the emission rate of the $^{ 5}D_{0}\rightarrow$ $^{ 7}F_{2}$ transition. Using finite-difference time-domain modeling, we corroborate these observations with the increase in the photonic density of states seen by the $Eu^{ 3+}$ ions in the proximity of the HMM. Our results indicate HMMs can serve as a valuable tool to control the emission from weak transitions, and hence hint at a route towards more practical applications of rare-earth ions in nanoscale optoelectronics and quantum devices.

physics.optics

Coefficient of performance at maximum figure of merit and its bounds for low-dissipation Carnot-like refrigerators

The figure of merit for refrigerators performing finite-time Carnot-like cycles between two reservoirs at temperature $T_h$ and $T_c$ ($<T_h$) is optimized. It is found that the coefficient of performance at maximum figure of merit is bounded between 0 and $(\sqrt{9+8\varepsilon_c}-3)/2$ for the low-dissipation refrigerators, where $\varepsilon_c =T_c/(T_h-T_c)$ is the Carnot coefficient of performance for reversible refrigerators. These bounds can be reached for extremely asymmetric low-dissipation cases when the ratio between the dissipation constants of the processes in contact with the cold and hot reservoirs approaches to zero or infinity, respectively. The observed coefficients of performance for real refrigerators are located in the region between the lower and upper bounds, which is in good agreement with our theoretical estimation.

cond-mat.stat-mech