SearcharxivSearch

arXiv subjects

Wenjing Jiang

Publications and source records attributed to Wenjing Jiang.

13 recordsLinked to original sources

Harmonic-Aware Transformer for Real-Time Catheter Localization in Interventional Procedures of Magnetic Particle Imaging

Magnetic particle imaging (MPI) enables real-time, radiation-free tracking of magnetic nanoparticle-coated instruments, making it highly suitable for interventional procedures. This study proposes a harmonic-aware transformer framework that directly predicts catheter tip positions from raw MPI voltage signals, eliminating the need for image reconstruction and reducing computational latency. The framework incorporates frequency-domain preprocessing to isolate the 2nd to 8th drive-field harmonics, enhancing the signal-to-noise ratio while preserving motion-relevant features. A transformer architecture with six encoder layers and eight attention heads is employed to learn spatio-temporal dependencies across the three receive axes (x, y, z) for accurate three-dimensional position estimation. The model is trained on simulated MPI signals and evaluated on real in vitro datasets under standard, bending, and heartbeat-like motion conditions. The proposed method achieves sub-millimeter localization accuracy, with a minimum L2 error of 0.103 +/- 0.092 mm and mean absolute errors (MAEs) of 0.039 +/- 0.046 mm, 0.054 +/- 0.049 mm, and 0.060 +/- 0.044 mm along the (x, y, z) axes, respectively, for the bending dataset. Across all datasets, the MAE ranges from 0.165 mm to 0.655 mm, demonstrating consistent performance. The optimized inference achieves a latency of 0.55 ms per frame and a throughput of approximately 1800 frames per second, confirming real-time capability. Compared with conventional MPI-guided approaches relying on image reconstruction, the proposed framework provides improved accuracy, reduced latency, and enhanced robustness under complex motion conditions. These results highlight the potential of harmonic-aware transformer models as efficient and scalable solutions for real-time catheter localization in interventional MPI.

physics.med-ph

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual search to introduce local evidence, but they typically do not explicitly distinguish perception from reasoning. In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as a Perceiver, and then answers the question as a Reasoner based on the annotated image and cropped regions. To better align training with this decoupled formulation, we further introduce Perception-Reasoning Alternating GRPO (PRA-GRPO), a role-aware reinforcement learning strategy that alternates between perception-focused and reasoning-focused updates using only final-answer supervision. Built on top of Qwen3-VL-Instruct-2B/4B/8B, P2R consistently improves performance across model scales. In particular, P2R-4B achieves 93.2% on V-Star, 81.9% on HR-Bench-4K, and 80.5% on HR-Bench-8K, substantially outperforming its corresponding backbone. Further experiments show that the benefits of P2R extend beyond high-resolution benchmarks to broader multimodal reasoning tasks. These results suggest that explicitly decoupling perception from reasoning provides an effective framework for fine-grained visual reasoning.

cs.CV

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of multimodal large language models purpose-built for content and AI safety, with both instruction-tuned and reasoning-oriented variants. Yuvion VL addresses this gap by treating safety as an inherently adversarial and multimodal problem and designing the entire pipeline around adversarial robustness. For data construction, we develop an automated pipeline integrating adversarial-aware data synthesis with multi-stage quality control, producing large-scale, high-quality multimodal samples augmented with domain knowledge and reasoning annotations. For training, we adopt a three-stage pipeline that includes continued pretraining for risk-concept cross-modal alignment, instruct post-training for production-grade safety tasks, and reasoning post-training for enhanced interpretability and performance in complex tasks. We further introduce Confuse-then-Contrast Fine-Tuning, a contrastive framework that mines model-specific confusions and constructs multi-image contrastive groups to enforce explicit discrimination of fine-grained visual-semantic elements, enabling the model to distinguish between visually similar cases with different safety implications in adversarial safety tasks. To support rigorous evaluation, we further introduce Yuvion VL RiskEval (YVRE), a collection of benchmarks covering diverse open and internal evaluations, with a focus on content and AI safety, adversarial robustness, and real-world capability requirements. Experiments show that Yuvion VL-32B achieves industry-leading safety performance, surpassing comparably sized open-source models and best closed-source commercial models, while maintaining comparable general capabilities.

cs.CV

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to evade model policies and safeguards. However, existing general-purpose model development largely overlook this adversarial nature, and often remain insufficient for realistic safety scenarios involving planning, tool use, and multi-step reasoning, causing measured safety performance to overestimate real deployment robustness. To address this gap, we present Yuvion LLM, a large language model built for adversarially robust content safety and broader AI safety. Yuvion LLM treats adversarial robustness and agentic capability as first-class objectives. Its pipeline combines adversarially aware data construction, knowledge-enhanced continued pretraining, and policy-grounded multi-task safety post-training, including risk-aware supervised fine-tuning and reinforcement learning-based policy optimization, together with safety-aware agentic reinforcement learning for tool use and multi-step reasoning in complex safety scenarios. We further introduce the Yuvion LLM RiskEval (YLRE), a collection of 93 benchmarks across four evaluation categories, covering diverse open and internal evaluations with a focus on safety, adversarial robustness, and real-world capability requirements. Across these evaluations, Yuvion LLM demonstrates clear advantages on safety-focused benchmarks and particularly strong robustness under adversarial conditions, while maintaining solid overall capability. Notably, Yuvion-8B outperforms most state-of-the-art baselines, including substantially larger models such as GPT-5.4 and Qwen3-MAX, on several safety tasks.

cs.CL

Cross-Axis Weighted Harmonic Method: A Frequency-Domain Approach for Enhanced Resolution in Magnetic Particle Imaging

Magnetic Particle Imaging (MPI) is a promising imaging modality that tracks magnetic nanoparticles (MNPs) to generate real time, high-resolution images. However, achieving an optimal balance between strong signal strength and sharp image clarity remains challenging. Higher drive field frequencies improve the signal-to-noise ratio (SNR), but also risk image blurring due to nanoparticle relaxation effects. To address this, we developed an end-to-end MPI simulation framework that models MNPs behavior, magnetic field dynamics, signal acquisition, and image reconstruction across a wide frequency range (20 to 85 kHz). Central to this framework is Cross-Axis Harmonic Analysis (CAHA), a novel, frequency-domain signal processing technique that adaptively extracts high-SNR harmonics from the x, y, and z directions for improved signal reconstruction. Using a simulated 3D vascular phantom, CAHA significantly enhanced image quality, achieving sub-millimeter resolution (0.8 mm FWHM at 85 kHz), strong noise suppression (nRMSE as low as 0.01), and structural fidelity (SSIM up to 0.94 at 55 kHz). The peak SNR reached 29.7 dB at 85 kHz. The signal processed with CAHA was also tested with other reconstruction methods; when combined with total variation regularization, CAHA achieved a pSNR of 37.91 dB. Evaluation on the Open MPI dataset further demonstrated up to 20% resolution improvement, confirming CAHA's robustness on real-world data. Although minor blurring was observed at the highest frequency due to relaxation, CAHA consistently maintained image clarity. By leveraging directional harmonic content rather than the full signal, CAHA sets a new benchmark for sharper, faster, and more robust MPI imaging.

physics.med-ph

Dynamic Facial Expressions Analysis Based Parkinson's Disease Auxiliary Diagnosis

Parkinson's disease (PD), a prevalent neurodegenerative disorder, significantly affects patients' daily functioning and social interactions. To facilitate a more efficient and accessible diagnostic approach for PD, we propose a dynamic facial expression analysis-based PD auxiliary diagnosis method. This method targets hypomimia, a characteristic clinical symptom of PD, by analyzing two manifestations: reduced facial expressivity and facial rigidity, thereby facilitating the diagnosis process. We develop a multimodal facial expression analysis network to extract expression intensity features during patients' performance of various facial expressions. This network leverages the CLIP architecture to integrate visual and textual features while preserving the temporal dynamics of facial expressions. Subsequently, the expression intensity features are processed and input into an LSTM-based classification network for PD diagnosis. Our method achieves an accuracy of 93.1%, outperforming other in-vitro PD diagnostic approaches. This technique offers a more convenient detection method for potential PD patients, improving their diagnostic experience.

cs.CV

High-Resolution Magnetic Particle Imaging System Matrix Recovery Using a Vision Transformer with Residual Feature Network

This study presents a hybrid deep learning framework, the Vision Transformer with Residual Feature Network (VRF-Net), for recovering high-resolution system matrices in Magnetic Particle Imaging (MPI). MPI resolution often suffers from downsampling and coil sensitivity variations. VRF-Net addresses these challenges by combining transformer-based global attention with residual convolutional refinement, enabling recovery of both large-scale structures and fine details. To reflect realistic MPI conditions, the system matrix is degraded using a dual-stage downsampling strategy. Training employed paired-image super-resolution on the public Open MPI dataset and a simulated dataset incorporating variable coil sensitivity profiles. For system matrix recovery on the Open MPI dataset, VRF-Net achieved nRMSE = 0.403, pSNR = 39.08 dB, and SSIM = 0.835 at 2x scaling, and maintained strong performance even at challenging scale 8x (pSNR = 31.06 dB, SSIM = 0.717). For the simulated dataset, VRF-Net achieved nRMSE = 4.44, pSNR = 28.52 dB, and SSIM = 0.771 at 2x scaling, with stable performance at higher scales. On average, it reduced nRMSE by 88.2%, increased pSNR by 44.7%, and improved SSIM by 34.3% over interpolation and CNN-based methods. In image reconstruction of Open MPI phantoms, VRF-Net further reduced reconstruction error to nRMSE = 1.79 at 2x scaling, while preserving structural fidelity (pSNR = 41.58 dB, SSIM = 0.960), outperforming existing methods. These findings demonstrate that VRF-Net enables sharper, artifact-free system matrix recovery and robust image reconstruction across multiple scales, offering a promising direction for future in vivo applications.

physics.med-ph

ML-based AIG Timing Prediction to Enhance Logic Optimization

As circuit designs become more intricate, obtaining accurate performance estimation in early stages, for effective design space exploration, becomes more time-consuming. Traditional logic optimization approaches often rely on proxy metrics to approximate post-mapping performance and area. However, these proxies do not always correlate well with actual post-mapping delay and area, resulting in suboptimal designs. To address this issue, we explore a ground-truth-based optimization flow that directly incorporates the exact post-mapping delay and area during optimization. While this approach improves design quality, it also significantly increases computational costs, particularly for large-scale designs. To overcome the runtime challenge, we apply machine learning models to predict post-mapping delay and area using the features extracted from AIGs. Our experimental results show that the model has high prediction accuracy with good generalization to unseen designs. Furthermore, the ML-enhanced logic optimization flow significantly reduces runtime while maintaining comparable performance and area outcomes.

cs.AR

IR-Aware ECO Timing Optimization Using Reinforcement Learning

Engineering change orders (ECOs) in late stages make minimal design fixes to recover from timing shifts due to excessive IR drops. This paper integrates IR-drop-aware timing analysis and ECO timing optimization using reinforcement learning (RL). The method operates after physical design and power grid synthesis, and rectifies IR-drop-induced timing degradation through gate sizing. It incorporates the Lagrangian relaxation (LR) technique into a novel RL framework, which trains a relational graph convolutional network (R-GCN) agent to sequentially size gates to fix timing violations. The R-GCN agent outperforms a classical LR-only algorithm: in an open 45nm technology, it (a) moves the Pareto front of the delay-power tradeoff curve to the left (b) saves runtime over the prior approaches by running fast inference using trained models, and (c) reduces the perturbation to placement by sizing fewer cells. The RL model is transferable across timing specifications and to unseen designs with fine tuning.

cs.AR

A Machine Learning Approach to Improving Timing Consistency between Global Route and Detailed Route

Due to the unavailability of routing information in design stages prior to detailed routing (DR), the tasks of timing prediction and optimization pose major challenges. Inaccurate timing prediction wastes design effort, hurts circuit performance, and may lead to design failure. This work focuses on timing prediction after clock tree synthesis and placement legalization, which is the earliest opportunity to time and optimize a "complete" netlist. The paper first documents that having "oracle knowledge" of the final post-DR parasitics enables post-global routing (GR) optimization to produce improved final timing outcomes. To bridge the gap between GR-based parasitic and timing estimation and post-DR results during post-GR optimization, machine learning (ML)-based models are proposed, including the use of features for macro blockages for accurate predictions for designs with macros. Based on a set of experimental evaluations, it is demonstrated that these models show higher accuracy than GR-based timing estimation. When used during post-GR optimization, the ML-based models show demonstrable improvements in post-DR circuit performance. The methodology is applied to two different tool flows - OpenROAD and a commercial tool flow - and results on 45nm bulk and 12nm FinFET enablements show improvements in post-DR slack metrics without increasing congestion. The models are demonstrated to be generalizable to designs generated under different clock period constraints and are robust to training data with small levels of noise.

cs.AR

Wafer-level substrate-free low-stress silicon nitride platform for THz metadevices and monolithically integrated narrowband metamaterial absorbers

The implementation of terahertz (THz) wafer-level metadevices is critical to advance the science for applications including (I) integrated focal plane array which can image for biology and (II) integrated narrowband absorbers for high spectral resolution THz spectroscopy. Substantial progress has been made in the development of THz metamaterials; however, a wafer-level low-stress THz metadevices platform remains a challenge. This paper experimentally demonstrates a substrate-free THz metadevices platform adopting engineered Si-rich and low-stress silicon nitride (SiNx) thin films, achieving an extensive THz transparency up to f = 2.5 THz. A new analytical model is first reported from the Lorentz model that can accurately predict spectral responses of metal insulator metal (MIM) metamaterial absorbers. The model is experimentally validated in the THz range and exploited for the first demonstration of a THz absorber, which exhibits performance approaching the predicted results. Our results show that the wafer-level SiNx platform will accelerate the development of large-scale, sophisticated substrate-free THz metadevices. The Lorentz model and its quadratic model will be a very practical method for designing THz metadevices.

physics.optics

Substrate-free THz focal plane metamaterial array with high absorption ratio

Microelectromechanical system (MEMS) focal plane array (FPA) with optical readout offers exciting opportunities for real-time terahertz (THz) imaging. However, conventional FPA suffers from a low THz absorption ratio, which further decreases the performance of THz imaging. Here, we present a simple and scalable approach for the realization of THz focal plane metamaterial array with a relatively high absorption ratio. The key idea is to combine the advantages of substrate-free structures with metamaterial. A 100 x100 THz FPA with a 150 x 150 μm pixel is designed, fabricated, and characterized. The dependence of the THz absorption ratio on the thickness of SiNx dielectric substrate film is investigated. The fabricated FPA exhibits a 90.6% resonant absorption at 1.36 THz, agreeing considerably with the theoretical simulation results. Our results imply that such a substrate-free THz focal plane metamaterial array enables the realization of THz imaging.

physics.app-ph

Dynamics of magnetic skyrmion clusters driven by spin-polarized current with a spatially varied polarization

Magnetic skyrmions are promising candidates for future information technology. Here, we present a micromagnetic study of isolated skyrmions and skyrmion clusters in ferromagnetic nanodisks driven by the spin-polarized current with spatially varied polarization. The current-driven skyrmion clusters can be either dynamic steady or static, depending on the spatially varied polarization profile. For the dynamic steady state, the skyrmion cluster moves in a circle in the nanodisk, while for the static state, the skyrmion cluster is static. The frequency of the circular motion of skyrmion is also studied. Furthermore, the dependence of the skyrmion cluster dynamics on the magnetic anisotropy and Dzyaloshinskii-Moriya interaction is investigated. Our results may provide a pathway to realize magnetic skyrmion cluster based devices.

cond-mat.mes-hall