SearcharxivSearch

arXiv subjects

Mingxi Chen

Publications and source records attributed to Mingxi Chen.

9 recordsLinked to original sources

Cladding Layer Enhanced GHz Bulk Acoustic Wave Resonance in Sodium Niobate Thin Films on Silicon

Bulk acoustic wave resonators (BAWR) and bandpass filters operating at GHz frequency are the workhorse of (Vo-)LTE telecommunication and broadband internet. In line with the Singapore Green Plan 2030 for innovating environmentally friendly products, we fabricated lead-free BAWR with sodium niobate (NaNbO3) piezoelectric on silicon with a high electromechanical coupling factor up to 31.3% operating at ~4 GHz. We disclose our crucial strategy where the NaNbO3 layer is cladded between two thin layers of high band gap insulators, which satisfies two primary objectives, i.e. leakage current mitigation and crack avoidance. In addition, we also verified the efficacy of reducing lattice parameters of the cladding layers in promoting vertically distorted tetragonal phase NaNbO3 and producing stronger BAWR signals.

cond-mat.mtrl-sci

Blowouts of Nascent Wind Bubbles in Pulsar-Driven Supernovae

Formation of a rapidly spinning, strongly magnetized neutron star (NS) may occur in various classes of core-collapse events. If the NS injects an amount of energy comparable to the explosion energy of the accompanying supernova (SN) before the SN ejecta becomes transparent, the nascent NS wind bubble can overtake the outer ejecta and undergo a blowout driven by hydrodynamic instabilities. Based on multidimensional numerical studies, we construct a minimal semi-analytic framework to follow the post-blowout dynamics and radiative evolution, map the blowout conditions by scanning the ejecta and NS parameters, and compute survey-ready multi-band light curves. For stripped-envelope SNe with an ejecta mass of $M_\mathrm{ej} \sim 10\,M_\odot$ and an explosion energy of $E_\mathrm{sn} \sim 10^{51}\,\mathrm{erg}$, blowout occurs for NSs with magnetic field strengths of $B_{\mathrm{dip}} \gtrsim 10^{13}\,\mathrm{G}$ and spin periods of $P_\mathrm{NS} \lesssim \mathrm{a\ few}\,\mathrm{ms}$. Relatively weak-field cases with $B_\mathrm{dip} \lesssim 10^{14}\,\mathrm{G}$ produce luminous double-peaked UV/optical light curves, as observed in the superluminous SN LSQ14bdq, while stronger-field cases with $B_\mathrm{dip} \gtrsim 10^{14}\,\mathrm{G}$ result in hypernovae preceded by X-ray blowout precursors. We also examine weaker and lower-mass SN explosions representing ultra-stripped SNe and accretion- or merger-induced collapse events, in which blowout is more readily achieved over a broader range of NS parameters, producing fast X-ray transients with durations of $ 10^{2\mbox{--}4}\,\mathrm{s}$ and peak luminosities of $10^{42\mbox{--}48}\,\mathrm{erg\,s^{-1}}$. Our results encourage coordinated UV, optical, and X-ray observations which constrain the formation of the most energetic NSs in the universe.

astro-ph.HE

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat text-guided segmentation as one-shot grounding, where the model predicts pixel prompts in a single forward pass to drive an external segmentor, which limits verification, refocusing and refinement when initial localization is wrong. To address this limitation, we propose RSAgent, an agentic Multimodal Large Language Model (MLLM) which interleaves reasoning and action for segmentation via multi-turn tool invocations. RSAgent queries a segmentation toolbox, observes visual feedback, and revises its spatial hypothesis using historical observations to re-localize targets and iteratively refine masks. We further build a data pipeline to synthesize multi-turn reasoning segmentation trajectories, and train RSAgent with a two-stage framework: cold-start supervised fine-tuning followed by agentic reinforcement learning with fine-grained, task-specific rewards. Extensive experiments show that RSAgent achieves a zero-shot performance of 66.5% gIoU on ReasonSeg test, improving over Seg-Zero-7B by 9%, and reaches 81.5% cIoU on RefCOCOg, demonstrating state-of-the-art performance on both in-domain and out-of-domain benchmarks.

cs.CV

Levitated macroscopic rotors with 10 hours of free spin at room temperature

Low-dissipation rotors with large angular momentum are essential for precision sensing and probing macroscopic quantum phenomena. To date, low dissipation can only be achieved for micro-scale rotors. Here, we report a diamagnetically levitated millimeter-scale rotor exhibiting a measured dissipation rate as low as $3.85\,\mu\mathrm{Hz}$ at room temperature, corresponding to a free spinning duration exceeding 10 hours. The rotor is levitated stably over an axisymmetric permanent magnet trap, and can be driven up to 930 RPM using contactless electrostatic actuation in high vacuum. Leveraging its low damping rate and large angular momentum, we realize a precision gyroscope with a measured sensitivity of $6.5 \times 10^{-3}\ \mathrm{^\circ/s}$ and an estimated thermal-limited stability of $5.7 \times 10^{-7}\ \mathrm{^\circ/\sqrt{h}}$. These results establish diamagnetic levitation as a promising room-temperature platform for high-performance gyroscopes.

physics.app-ph

NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results

Cross-Domain Few-Shot Object Detection (CD-FSOD) poses significant challenges to existing object detection and few-shot detection models when applied across domains. In conjunction with NTIRE 2025, we organized the 1st CD-FSOD Challenge, aiming to advance the performance of current object detectors on entirely novel target domains with only limited labeled data. The challenge attracted 152 registered participants, received submissions from 42 teams, and concluded with 13 teams making valid final submissions. Participants approached the task from diverse perspectives, proposing novel models that achieved new state-of-the-art (SOTA) results under both open-source and closed-source settings. In this report, we present an overview of the 1st NTIRE 2025 CD-FSOD Challenge, highlighting the proposed solutions and summarizing the results submitted by the participants.

cs.CV

MedConv: Convolutions Beat Transformers on Long-Tailed Bone Density Prediction

Bone density prediction via CT scans to estimate T-scores is crucial, providing a more precise assessment of bone health compared to traditional methods like X-ray bone density tests, which lack spatial resolution and the ability to detect localized changes. However, CT-based prediction faces two major challenges: the high computational complexity of transformer-based architectures, which limits their deployment in portable and clinical settings, and the imbalanced, long-tailed distribution of real-world hospital data that skews predictions. To address these issues, we introduce MedConv, a convolutional model for bone density prediction that outperforms transformer models with lower computational demands. We also adapt Bal-CE loss and post-hoc logit adjustment to improve class balance. Extensive experiments on our AustinSpine dataset shows that our approach achieves up to 21% improvement in accuracy and 20% in ROC AUC over previous state-of-the-art methods.

cs.CV

ProjectedEx: Enhancing Generation in Explainable AI for Prostate Cancer

Prostate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy. Recent advancements in explainable AI and representation learning have significantly improved prostate cancer diagnosis by enabling automated and precise lesion classification. However, existing explainable AI methods, particularly those based on frameworks like generative adversarial networks (GANs), are predominantly developed for natural image generation, and their application to medical imaging often leads to suboptimal performance due to the unique characteristics and complexity of medical image. To address these challenges, our paper introduces three key contributions. First, we propose ProjectedEx, a generative framework that provides interpretable, multi-attribute explanations, effectively linking medical image features to classifier decisions. Second, we enhance the encoder module by incorporating feature pyramids, which enables multiscale feedback to refine the latent space and improves the quality of generated explanations. Additionally, we conduct comprehensive experiments on both the generator and classifier, demonstrating the clinical relevance and effectiveness of ProjectedEx in enhancing interpretability and supporting the adoption of AI in medical settings. Code will be released at https://github.com/Richardqiyi/ProjectedEx

eess.IV

Enhanced TM-Mode 3D Coupled Wave Theory for Photonic Crystal Surface-Emitting Terahertz Quantum Cascade Lasers

In this study, we propose and develop an enhanced three-dimensional coupled wave theory (3D CWT) to investigate the optical field behavior in photonic crystal surface-emitting terahertz quantum cascade lasers (THz-QCLs). By incorporating an effective permittivity enhancement (EP) model and a self-consistent iteration (SCI) method, we successfully address the numerical dispersion issues encountered in analytical methods when dealing with metallic waveguide structures. The results demonstrate that the EP and SCI-enhanced 3D TM mode CWT achieves computational accuracy comparable to traditional numerical simulation methods such as finite-difference time-domain (FDTD), while significantly reducing the required computational resources, including time and memory, to just tens of minutes. Moreover, this method provides a clear physical insight, revealing the reasons behind the current low extraction efficiency in surface-emitting THz-QCLs. Our study showcases the potential of the EP and SCI-enhanced 3D CWT as a powerful simulation tool in the research of photonic crystal surface-emitting lasers, offering a new theoretical foundation and optimization direction for future laser designs.

physics.optics

JointViT: Modeling Oxygen Saturation Levels with Joint Supervision on Long-Tailed OCTA

The oxygen saturation level in the blood (SaO2) is crucial for health, particularly in relation to sleep-related breathing disorders. However, continuous monitoring of SaO2 is time-consuming and highly variable depending on patients' conditions. Recently, optical coherence tomography angiography (OCTA) has shown promising development in rapidly and effectively screening eye-related lesions, offering the potential for diagnosing sleep-related disorders. To bridge this gap, our paper presents three key contributions. Firstly, we propose JointViT, a novel model based on the Vision Transformer architecture, incorporating a joint loss function for supervision. Secondly, we introduce a balancing augmentation technique during data preprocessing to improve the model's performance, particularly on the long-tail distribution within the OCTA dataset. Lastly, through comprehensive experiments on the OCTA dataset, our proposed method significantly outperforms other state-of-the-art methods, achieving improvements of up to 12.28% in overall accuracy. This advancement lays the groundwork for the future utilization of OCTA in diagnosing sleep-related disorders. See project website https://steve-zeyu-zhang.github.io/JointViT

cs.CV