SearcharxivSearch

arXiv subjects

Jiaxi Wang

Publications and source records attributed to Jiaxi Wang.

13 recordsLinked to original sources

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation

Autonomous endoscopic navigation can reduce clinicians' operational burden, yet robust control remains challenging due to tissue deformation, transient occlusions, and rapidly changing viewpoints. Existing learning-based policies typically predict actions from current observations without explicitly modeling future dynamics, limiting their robustness and reliability in safety-critical settings. World Action Models (WAMs) offer a promising alternative by coupling predictive visual dynamics with action generation, but extending them to robotic endoscopy remains challenging due to limited training data, restricted viewpoint diversity, deformable anatomy, and high inference latency. We present EndoWAM, which is, to our knowledge, the first WAM for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. Specifically, EndoWAM couples a lightweight diffusion transformer for future target-region prediction with a discrete action expert through a shared predictive representation. This design injects target-aware supervision into predictive dynamics modeling, improving robustness to visual degradation and viewpoint changes while enabling real-time control in a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset spanning three anatomically distinct procedures: ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms all baselines and alternative grounding strategies, while demonstrating strong zero-shot generalization to unseen viewpoints, environments, and targets. These results establish EndoWAM as a predictive, target-grounded framework for accurate, generalizable, and long-horizon navigation in visually constrained endoscopic environments.

cs.RO

Local magnetic correlations and light-sensitive centers in the Cr2AlC MAX phase

Cr2AlC MAX phase is synthesized by high-pressure solid-state annealing and investigated as a candidate platform for optically responsive magnetism. Structural characterization confirms the formation of the Cr2AlC phase, while magnetic and optical-magnetic properties are examined by superconducting quantum interference device (SQUID) magnetometry, electron spin resonance (ESR), and first principles calculations. SQUID magnetometry identifies Cr 2AlC as a weak, field-linear metallic paramagnet dominated by Pauli-like susceptibility of itinerant Cr-derived states. Its non-monotonic temperature dependence is described by an additional contribution from antiferromagnetically coupled Cr-Cr dimers, whereas the low-temperature Curie-like upturn originates from only a trace population of localized Cr centers. Under red-light illumination, SQUID magnetometry does not reveal an intrinsic macroscopic optomagnetic response. In contrast, ESR at 4 K shows a reversible light-induced reduction of a local magnetic signal, but the optically modified spin population corresponds only to several tens of ppm of the Cr sublattice. Ab initio Bethe-Salpeter equation (ai-BSE) calculations combined with the maximally localized Wannier function analysis suggest that optical excitation can redistribute spin polarization between neighboring Cr sites with the opposite local moments. The combined experiment-theory approach therefore establishes the hierarchy of magnetic contributions in Cr 2AlC and identifies the microscopic origin of its local optical sensitivity. This provides a reference for designing MAX phases and related MXenes in which defects, surface terminations or reduced dimensionality may enhance optically active magnetic states.

cond-mat.mtrl-sci

Hybrid Disclination Skin-topological Effects in Non-Hermitian Circuits

The bulk-disclination correspondence (BDC) is a fundamental concept in Hermitian systems that has been widely applied to predict disclination states. Recently, disclination states have also been observed and experimentally verified in non-Hermitian systems with C6 lattice symmetry, where gain and loss are introduced to induce non-Hermiticity. In this Letter, we propose a non-Hermitian two-dimensional (2D) Su-Schrieffer-Heeger (SSH) disclination model with skin-topological (ST) disclination states, and calculate its biorthogonal Zak phase. Together with the real-space disclination index, we predict the emergence of disclination states in a C4-symmetric non-Hermitian lattice and the corresponding fractional charge. We also generalize the symmetry indicator within the biorthogonal framework to predict the anomalous filling near the disclination core. Experimentally, the model is implemented on a nonreciprocal circuit platform, where we analyze the impedance matrix characterized by complex eigenfrequencies and directly observe the ST disclination states. Our work further extends the bulk-disclination correspondence to the non-Hermitian realm.

cond-mat.mtrl-sci

MIPS: a Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction

Polymers, composed of repeating structural units called monomers, are fundamental materials in daily life and industry. Accurate property prediction for polymers is essential for their design, development, and application. However, existing modeling approaches, which typically represent polymers by the constituent monomers, struggle to capture the whole properties of polymer, since the properties change during the polymerization process. In this study, we propose a Multimodal Infinite Polymer Sequence (MIPS) pre-training framework, which represents polymers as infinite sequences of monomers and integrates both topological and spatial information for comprehensive modeling. From the topological perspective, we generalize message passing mechanism (MPM) and graph attention mechanism (GAM) to infinite polymer sequences. For MPM, we demonstrate that applying MPM to infinite polymer sequences is equivalent to applying MPM on the induced star-linking graph of monomers. For GAM, we propose to further replace global graph attention with localized graph attention (LGA). Moreover, we show the robustness of the "star linking" strategy through Repeat and Shift Invariance Test (RSIT). Despite its robustness, "star linking" strategy exhibits limitations when monomer side chains contain ring structures, a common characteristic of polymers, as it fails the Weisfeiler-Lehman~(WL) test. To overcome this issue, we propose backbone embedding to enhance the capability of MPM and LGA on infinite polymer sequences. From the spatial perspective, we extract 3D descriptors of repeating monomers to capture spatial information. Finally, we design a cross-modal fusion mechanism to unify the topological and spatial information. Experimental validation across eight diverse polymer property prediction tasks reveals that MIPS achieves state-of-the-art performance.

cs.LG

Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data

The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios.

cs.CV

Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation

Visual grounding aims to align visual information of specific regions of images with corresponding natural language expressions. Current visual grounding methods leverage pre-trained visual and language backbones independently to obtain visual features and linguistic features. Although these two types of features are then fused through elaborately designed networks, the heterogeneity of the features renders them unsuitable for multi-modal reasoning. This problem arises from the domain gap between the single-modal pre-training backbones used in current visual grounding methods, which can hardly be bridged by the traditional end-to-end training method. To alleviate this, our work proposes an Empowering Pre-trained Model for Visual Grounding (EpmVG) framework, which distills a multimodal pre-trained model to guide the visual grounding task. EpmVG relies on a novel cross-modal distillation mechanism that can effectively introduce the consistency information of images and texts from the pre-trained model, reducing the domain gap in the backbone networks, and thereby improving the performance of the model in the visual grounding task. Extensive experiments have been conducted on five conventionally used datasets, and the results demonstrate that our method achieves better performance than state-of-the-art methods.

cs.CV

Towards Predicting Equilibrium Distributions for Molecular Systems with Deep Learning

Advances in deep learning have greatly improved structure prediction of molecules. However, many macroscopic observations that are important for real-world applications are not functions of a single molecular structure, but rather determined from the equilibrium distribution of structures. Traditional methods for obtaining these distributions, such as molecular dynamics simulation, are computationally expensive and often intractable. In this paper, we introduce a novel deep learning framework, called Distributional Graphormer (DiG), in an attempt to predict the equilibrium distribution of molecular systems. Inspired by the annealing process in thermodynamics, DiG employs deep neural networks to transform a simple distribution towards the equilibrium distribution, conditioned on a descriptor of a molecular system, such as a chemical graph or a protein sequence. This framework enables efficient generation of diverse conformations and provides estimations of state densities. We demonstrate the performance of DiG on several molecular tasks, including protein conformation sampling, ligand structure sampling, catalyst-adsorbate sampling, and property-guided structure generation. DiG presents a significant advancement in methodology for statistically understanding molecular systems, opening up new research opportunities in molecular science.

physics.chem-ph

Understanding the Failure of Batch Normalization for Transformers in NLP

Batch Normalization (BN) is a core and prevalent technique in accelerating the training of deep neural networks and improving the generalization on Computer Vision (CV) tasks. However, it fails to defend its position in Natural Language Processing (NLP), which is dominated by Layer Normalization (LN). In this paper, we are trying to answer why BN usually performs worse than LN in NLP tasks with Transformer models. We find that the inconsistency between training and inference of BN is the leading cause that results in the failure of BN in NLP. We define Training Inference Discrepancy (TID) to quantitatively measure this inconsistency and reveal that TID can indicate BN's performance, supported by extensive experiments, including image classification, neural machine translation, language modeling, sequence labeling, and text classification tasks. We find that BN can obtain much better test performance than LN when TID keeps small through training. To suppress the explosion of TID, we propose Regularized BN (RBN) that adds a simple regularization term to narrow the gap between batch statistics and population statistics of BN. RBN improves the performance of BN consistently and outperforms or is on par with LN on 17 out of 20 settings, involving ten datasets and two common variants of Transformer Our code is available at https://github.com/wjxts/RegularizedBN.

cs.CL

EVIboost for the Estimation of Extreme Value Index under Heterogeneous Extremes

Modeling heterogeneity on heavy-tailed distributions under a regression framework is challenging, and classical statistical methodologies usually place conditions on the distribution models to facilitate the learning procedure. However, these conditions are likely to overlook the complex dependence structure between the heaviness of tails and the covariates. Moreover, data sparsity on tail regions also makes the inference method less stable, leading to largely biased estimates for extreme-related quantities. This paper proposes a gradient boosting algorithm to estimate a functional extreme value index with heterogeneous extremes. Our proposed algorithm is a data-driven procedure that captures complex and dynamic structures in tail distributions. We also conduct extensive simulation studies to show the prediction accuracy of the proposed algorithm. In addition, we apply our method to a real-world data set to illustrate the state-dependent and time-varying properties of heavy-tail phenomena in the financial industry.

stat.ME

Bringing in the outliers: A sparse subspace clustering approach to learn a dictionary of mouse ultrasonic vocalizations

Mice vocalize in the ultrasonic range during social interactions. These vocalizations are used in neuroscience and clinical studies to tap into complex behaviors and states. The analysis of these ultrasonic vocalizations (USVs) has been traditionally a manual process, which is prone to errors and human bias, and is not scalable to large scale analysis. We propose a new method to automatically create a dictionary of USVs based on a two-step spectral clustering approach, where we split the set of USVs into inlier and outlier data sets. This approach is motivated by the known degrading performance of sparse subspace clustering with outliers. We apply spectral clustering to the inlier data set and later find the clusters for the outliers. We propose quantitative and qualitative performance measures to evaluate our method in this setting, where there is no ground truth. Our approach outperforms two baselines based on k-means and spectral clustering in all of the proposed performance measures, showing greater distances between clusters and more variability between clusters.

eess.AS

Ultrafast Intercavity Nonlinear Couplings between Polaritons

Realizing nonlinear coupling across space can enable new scientific and technological advances, including ultrafast operation and propagation of information in IR photonic circuitry, remote triggering or catalyzing of chemical reactions, and new platforms for quantum simulations with increased complexities. In this report, we show that ultrafast nonlinear couplings are achieved between polaritons residing in different cavities, in the mid-infrared (IR) regime, e.g. by pumping polaritons in one cavity, the polaritons in the adjacent cavity can be affected. By hybridizing photon and molecular vibrational modes, molecular vibrational polaritons are formed that have the combined characteristic of both photon delocalization and molecular nonlinearity. Thus, although photons have little nonlinear coupling cross-section, and molecular nonlinearity is localized, the dual photon/molecule character of polaritons allows photons to affect each other across different cavities, through coupling to the same molecules - a novel property that neither molecular nor cavity mode would possess alone.

physics.optics

Revealing Hidden Vibration Polariton Interactions by 2D IR Spectroscopy

We report the first experimental two-dimensional infrared (2D IR) spectra of novel molecular photonic excitations - vibrational-polaritons. The application of advanced 2D IR spectroscopy onto novel vibrational-polariton challenges and advances our understanding in both fields. From spectroscopy aspect, 2D IR spectra of polaritons differ drastically from free uncoupled molecules; from vibrational-polariton aspects, 2D IR uniquely resolves hybrid light-matter polariton excitations and unexpected dark states in a state-selective manner and revealed hidden interactions between them. Moreover, 2D IR signals highlight the role of vibrational anharmonicities in generating non-linear signals. To further advance our knowledge on 2D IR of vibrational polaritons, we develop a new quantum-mechanical model incorporating the effects of both nuclear and electrical anharmonicities on vibrational-polaritons and their 2D IR signals. This work reveals polariton physics that is difficult or impossible to probe with traditional linear spectroscopy and lays the foundation for investigating new non-linear optics and chemistry of molecular vibrational-polaritons.

quant-ph

Solving the "Magic Angle" Challenge in Determining Molecular Orientation at Interfaces

We introduce a novel method to determine the orientation heterogeneity (mean tilt angle and orientational distribution) of molecules at interfaces using heterodyne two-dimensional sum frequency generation spectroscopy. By doing so, we not only have solved the long-standing "magic angle" challenge, i.e. the measurement of molecular orientation by assuming a narrow orientational distribution results in ambiguities, but we also are able to determine the orientational distribution, which is otherwise difficult to measure. We applied our new method to a CO2 reduction catalyst/gold interface and found that the catalysts formed a monolayer with a mean tilt angle between the quasi-C3 symmetric axis of the catalysts and the surface normal of 53 deg, with 5 deg orientational distribution. Although applied to a specific system, this method is a general way to determine the orientation heterogeneity of an ensemble-averaged molecular interface, which can potentially be applied to a wide-range of energy material, catalytic and biological interfaces.

physics.chem-ph