SearcharxivSearch

arXiv subjects

Ruixing Liang

Publications and source records attributed to Ruixing Liang.

At least 19 recordsLinked to original sources

The Role of Mixed and Augmented Reality in Medical Visualization: Literature Review and A Context-Aware Taxonomy

The discovery and evolution of medical imaging technologies have enabled non-invasive visualization of internal anatomy that has be- come essential for supporting diagnosis, monitoring, and treatment. However, because medical imaging relies on complex physical processes and contrast mechanisms for image formation, imaging alone is not sufficient to enable humans to leverage the resulting in- formation fully. In addition, traditional methods to visualize the resulting information use two-dimensional displays to present three-dimensional anatomical structures. The introduction of Augmented and Mixed Reality (AR/MR) technologies offers an opportunity to provide valuable paradigms for medical imaging visualization, allowing users to observe, explore, and interact with anatomical infor- mation in more spatially intuitive ways. However, naive implementation without careful design considerations can lead to perceptual inconsistencies, potentially compromising utility and effectiveness. In this paper, we present a structured taxonomy of medical AR/MR visualization strategies aimed at providing clearer insight into how visualization design varies across clinical use cases. The taxonomy organizes techniques based on four core design components: image modality, data dimensionality, display technology, and clinical application. In addition, we introduce two critical dimensions that are often overlooked in the literature: visualization anchoring (the spatial relationship between virtual content and the physical world), and perceptual awareness (the use of visual cues to support spatial interpretation). Together, these components form a comprehensive taxonomy, offering a detailed framework for selecting appropriate visualization techniques in medical applications.

cs.GR

Extend Your Horizon: A Device-Agnostic Surgical Tool Tracking Framework with Multi-View Optimization for Augmented Reality

Surgical navigation provides real-time guidance by estimating the pose of patient anatomy and surgical instruments to visualize relevant intraoperative information. In conventional systems, instruments are typically tracked using fiducial markers and stationary optical tracking systems (OTS). Augmented reality (AR) has further enabled intuitive visualization and motivated tracking using sensors embedded in head-mounted displays (HMDs). However, most existing approaches rely on a clear line of sight, which is difficult to maintain in dynamic operating room environments due to frequent occlusions caused by equipment, surgical tools, and personnel. This work introduces a framework for tracking surgical instruments under occlusion by fusing multiple sensing modalities within a dynamic scene graph representation. The proposed approach integrates tracking systems with different accuracy levels and motion characteristics while estimating tracking reliability in real time. Experimental results demonstrate improved robustness and enhanced consistency of AR visualization in the presence of occlusions.

cs.HC

Quasiparticle interference in LiFeAs: Signature of inelastic tunneling through spin fluctuations

Quasiparticle interference (QPI) is a powerful tool to characterize the symmetry of the superconducting order parameter in unconventional superconductors, by mapping the spatial dependence of elastic tunneling of electrons between the tip of a scanning tunneling microscope and a sample. Here, we consider the influence of inelastic tunneling on quasi-particle interference, exemplarily for the iron-based superconductor LiFeAs. We clearly observe replica features in both experimental QPI maps and the dispersion extracted from QPI, which from comparison with theoretical model calculations can be attributed to inelastic tunneling. Analysis of the QPI dispersion shows that the inelastic mode that gives rise to these replica features exhibits a resonance between 8 and 10 meV. Comparison of the energy scale of the resonance energy estimated from QPI with inelastic neutron scattering indicates that the replica features arise from interaction with spin fluctuations.

cond-mat.supr-con

SurgiSR4K: A High-Resolution Endoscopic Video Dataset for Robotic-Assisted Minimally Invasive Procedures

High-resolution imaging is crucial for enhancing visual clarity and enabling precise computer-assisted guidance in minimally invasive surgery (MIS). Despite the increasing adoption of 4K endoscopic systems, there remains a significant gap in publicly available native 4K datasets tailored specifically for robotic-assisted MIS. We introduce SurgiSR4K, the first publicly accessible surgical imaging and video dataset captured at a native 4K resolution, representing realistic conditions of robotic-assisted procedures. SurgiSR4K comprises diverse visual scenarios including specular reflections, tool occlusions, bleeding, and soft tissue deformations, meticulously designed to reflect common challenges faced during laparoscopic and robotic surgeries. This dataset opens up possibilities for a broad range of computer vision tasks that might benefit from high resolution data, such as super resolution (SR), smoke removal, surgical instrument detection, 3D tissue reconstruction, monocular depth estimation, instance segmentation, novel view synthesis, and vision-language model (VLM) development. SurgiSR4K provides a robust foundation for advancing research in high-resolution surgical imaging and fosters the development of intelligent imaging technologies aimed at enhancing performance, safety, and usability in image-guided robotic surgeries.

eess.IV

CASPER: A Large Scale Spontaneous Speech Dataset

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.

cs.CL

Thermal Hall conductivity in the strongest cuprate superconductor: Estimate of the mean free path in the trilayer cuprate HgBa$_2$Ca$_2$Cu$_3$O$_{8 + \delta}$

The thermal Hall conductivity of the trilayer cuprate HgBa$_2$Ca$_2$Cu$_3$O$_{8+\delta}$ (Hg1223) - the superconductor with the highest critical temperature $T_c$ at ambient pressure - was measured at temperatures down to 2 K for three dopings in the underdoped regime ($p$ = 0.09, 0.10, 0.11). By combining a previously introduced simple model and prior theoretical results, we derive a formula for the inverse mean free path, $1 / \ell$, which allows us to estimate the mean free path of $d$-wave quasiparticles in Hg1223 below $T_c$. We find that $1 / \ell$ grows as $T^3$, in agreement with the theoretical expectation for a clean $d$-wave superconductor. Measurements were also conducted on the single layer mercury-based cuprate HgBa$_2$CuO$_{6+\delta}$ (Hg1201), revealing that the mean free path in this compound is roughly half that of its three-layered counterpart at the same doping ($p$ = 0.10). This observation is be attributed to the protective role of the outer planes in Hg1223, which results in a more pristine inner plane. We also report data in an ultraclean crystal of YBa$_2$Cu$_3$O$_y$ (YBCO) with full oxygen content $p$ = 0.18, believed to be the cleanest of any cuprate, and find that $\ell$ is not longer than in Hg1223.

cond-mat.supr-con

A novel open-source ultrasound dataset with deep learning benchmarks for spinal cord injury localization and anatomical segmentation

While deep learning has catalyzed breakthroughs across numerous domains, its broader adoption in clinical settings is inhibited by the costly and time-intensive nature of data acquisition and annotation. To further facilitate medical machine learning, we present an ultrasound dataset of 10,223 Brightness-mode (B-mode) images consisting of sagittal slices of porcine spinal cords (N=25) before and after a contusion injury. We additionally benchmark the performance metrics of several state-of-the-art object detection algorithms to localize the site of injury and semantic segmentation models to label the anatomy for comparison and creation of task-specific architectures. Finally, we evaluate the zero-shot generalization capabilities of the segmentation models on human ultrasound spinal cord images to determine whether training on our porcine dataset is sufficient for accurately interpreting human data. Our results show that the YOLOv8 detection model outperforms all evaluated models for injury localization, achieving a mean Average Precision (mAP50-95) score of 0.606. Segmentation metrics indicate that the DeepLabv3 segmentation model achieves the highest accuracy on unseen porcine anatomy, with a Mean Dice score of 0.587, while SAMed achieves the highest Mean Dice score generalizing to human anatomy (0.445). To the best of our knowledge, this is the largest annotated dataset of spinal cord ultrasound images made publicly available to researchers and medical professionals, as well as the first public report of object detection and segmentation architectures to assess anatomical markers in the spinal cord for methodology development and clinical applications.

eess.IV

SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

Surgical data science has seen rapid advancement with the excellent performance of end-to-end deep neural networks (DNNs). Despite their successes, DNNs have been proven susceptible to minor "corruptions," introducing a major concern for the translation of cutting-edge technology, especially in high-stakes scenarios. We introduce the SegSTRONG-C challenge dedicated to better understanding model deterioration under unforeseen but plausible non-adversarial "corruption" and the capabilities of contemporary methods that seek to improve it. Built on a dataset generated through counterfactual robotic replay, SegSTRONG-C provides paired clean and "corrupted" samples, enabling reproducible evaluation of model robustness. Participants are challenged to train tool segmentation algorithms on "uncorrupted" data and evaluate them on "corrupted" test domains for the binary robot tool segmentation task. Through comprehensive baseline experiments and participating submissions from widespread community engagement, SegSTRONG-C reveals key themes for model failure and identifies promising directions for improving robustness. The performance of challenge winners, achieving an average 0.9394 DSC and 0.9301 NSD across the unreleased test sets with "corruption" types: bleeding, smoke, and low brightness. This highlights how prior knowledge, customized training strategies, and architectural choice can be leveraged to improve robustness. In conclusion, the SegSTRONG-C challenge has identified practical approaches for enhancing model robustness. However, most approaches rely on conventional techniques that have known limitations. Looking ahead, we advocate for expanding intellectual diversity and creativity in non-adversarial robustness beyond data augmentation, calling for new paradigms that enhance universal robustness to unforeseen "corruptions" to facilitate richer applications in surgical data science.

cs.CV

Absence of Fermi surface reconstruction in pressure-driven overdoped YBCO

The evolution of the critical superconducting temperature and field, quantum oscillation frequencies and effective mass $m^{*}$ in underdoped YBa$_2$Cu$_3$O$_{7-\delta}$ (YBCO) crystals ($p$ = 0.11, with $p$ the hole concentration per Cu atom) points to a partial suppression of the charge orders with increasing pressure up to 7 GPa, mimicking doping. Application of pressures up to 25 GPa pushes the sample to the overdoped side of the superconducting dome. Contrary to other cuprates, or to doping studies on YBCO, the frequency of the quantum oscillations measured in that pressure range do not support the picture of a Fermi-surface reconstruction in the overdoped regime, but possibly point to the existence of a new charge order.

cond-mat.supr-con

Planar thermal Hall effect from phonons in cuprates

A surprising "planar" thermal Hall effect, whereby the field is parallel to the current, has recently been observed in a few magnetic insulators, and this has been attributed to exotic excitations such as Majorana fermions or chiral magnons. Here we investigate the possibility of a planar thermal Hall effect in three different cuprate materials, in which the conventional thermal Hall conductivity $\kappa_{\rm {xy}}$ (with an out-of-plane field perpendicular to the current) is dominated by either electrons or phonons. Our measurements show that the planar $\kappa_{\rm {xy}}$ from electrons in cuprates is zero, as expected from the absence of a Lorentz force in the planar configuration. By contrast, we observe a sizable planar $\kappa_{\rm {xy}}$ in those samples where the thermal Hall response is due to phonons, even though it should in principle be forbidden by the high crystal symmetry. Our findings call for a careful re-examination of the mechanisms responsible for the phonon thermal Hall effect in insulators.

cond-mat.str-el

Unidirectional brain-computer interface: Artificial neural network encoding natural images to fMRI response in the visual cortex

While significant advancements in artificial intelligence (AI) have catalyzed progress across various domains, its full potential in understanding visual perception remains underexplored. We propose an artificial neural network dubbed VISION, an acronym for "Visual Interface System for Imaging Output of Neural activity," to mimic the human brain and show how it can foster neuroscientific inquiries. Using visual and contextual inputs, this multimodal model predicts the brain's functional magnetic resonance imaging (fMRI) scan response to natural images. VISION successfully predicts human hemodynamic responses as fMRI voxel values to visual inputs with an accuracy exceeding state-of-the-art performance by 45%. We further probe the trained networks to reveal representational biases in different visual areas, generate experimentally testable hypotheses, and formulate an interpretable metric to associate these hypotheses with cortical functions. With both a model and evaluation metric, the cost and time burdens associated with designing and implementing functional analysis on the visual cortex could be reduced. Our work suggests that the evolution of computational models may shed light on our fundamental understanding of the visual cortex and provide a viable approach toward reliable brain-machine interfaces.

cs.CV

TAToo: Vision-based Joint Tracking of Anatomy and Tool for Skull-base Surgery

Purpose: Tracking the 3D motion of the surgical tool and the patient anatomy is a fundamental requirement for computer-assisted skull-base surgery. The estimated motion can be used both for intra-operative guidance and for downstream skill analysis. Recovering such motion solely from surgical videos is desirable, as it is compliant with current clinical workflows and instrumentation. Methods: We present Tracker of Anatomy and Tool (TAToo). TAToo jointly tracks the rigid 3D motion of patient skull and surgical drill from stereo microscopic videos. TAToo estimates motion via an iterative optimization process in an end-to-end differentiable form. For robust tracking performance, TAToo adopts a probabilistic formulation and enforces geometric constraints on the object level. Results: We validate TAToo on both simulation data, where ground truth motion is available, as well as on anthropomorphic phantom data, where optical tracking provides a strong baseline. We report sub-millimeter and millimeter inter-frame tracking accuracy for skull and drill, respectively, with rotation errors below 1{\deg}. We further illustrate how TAToo may be used in a surgical navigation setting. Conclusion: We present TAToo, which simultaneously tracks the surgical tool and the patient anatomy in skull-base surgery. TAToo directly predicts the motion from surgical videos, without the need of any markers. Our results show that the performance of TAToo compares favorably to competing approaches. Future work will include fine-tuning of our depth network to reach a 1 mm clinical accuracy goal desired for surgical applications in the skull base.

cs.CV

Twin-S: A Digital Twin for Skull-base Surgery

Purpose: Digital twins are virtual interactive models of the real world, exhibiting identical behavior and properties. In surgical applications, computational analysis from digital twins can be used, for example, to enhance situational awareness. Methods: We present a digital twin framework for skull-base surgeries, named Twin-S, which can be integrated within various image-guided interventions seamlessly. Twin-S combines high-precision optical tracking and real-time simulation. We rely on rigorous calibration routines to ensure that the digital twin representation precisely mimics all real-world processes. Twin-S models and tracks the critical components of skull-base surgery, including the surgical tool, patient anatomy, and surgical camera. Significantly, Twin-S updates and reflects real-world drilling of the anatomical model in frame rate. Results: We extensively evaluate the accuracy of Twin-S, which achieves an average 1.39 mm error during the drilling process. We further illustrate how segmentation masks derived from the continuously updated digital twin can augment the surgical microscope view in a mixed reality setting, where bone requiring ablation is highlighted to provide surgeons additional situational awareness. Conclusion: We present Twin-S, a digital twin environment for skull-base surgery. Twin-S tracks and updates the virtual model in real-time given measurements from modern tracking technologies. Future research on complementing optical tracking with higher-precision vision-based approaches may further increase the accuracy of Twin-S.

cs.HC

PQLM -- Multilingual Decentralized Portable Quantum Language Model for Privacy Protection

With careful manipulation, malicious agents can reverse engineer private information encoded in pre-trained language models. Security concerns motivate the development of quantum pre-training. In this work, we propose a highly Portable Quantum Language Model (PQLM) that can easily transmit information to downstream tasks on classical machines. The framework consists of a cloud PQLM built with random Variational Quantum Classifiers (VQC) and local models for downstream applications. We demonstrate the ad hoc portability of the quantum model by extracting only the word embeddings and effectively applying them to downstream tasks on classical machines. Our PQLM exhibits comparable performance to its classical counterpart on both intrinsic evaluation (loss, perplexity) and extrinsic evaluation (multilingual sentiment analysis accuracy) metrics. We also perform ablation studies on the factors affecting PQLM performance to analyze model stability. Our work establishes a theoretical foundation for a portable quantum pre-trained language model that could be trained on private data and made available for public use with privacy protection guarantees.

cs.LG

The Three-Dimensional Electronic Structure of LiFeAs: Strong-coupling Superconductivity and Topology in the Iron Pnictides

Amongst the iron-based superconductors, LiFeAs is unrivalled in the simplicity of its crystal structure and phase diagram. However, our understanding of this canonical compound suffers from conflict between mutually incompatible descriptions of the material's electronic structure, as derived from contradictory interpretations of the photoemission record. Here, we explore the challenge of interpretation in such experiments. By combining comprehensive photon energy- and polarization- dependent angle-resolved photoemission spectroscopy (ARPES) measurements with numerical simulations, we establish the providence of several contradictions in the present understanding of this and related materials. We identify a confluence of surface-related issues which have precluded unambiguous identification of both the number and dimensionality of the Fermi surface sheets. Ultimately, we arrive at a scenario which supports indications of topologically non-trivial states, while also being incompatible with superconductivity as a spin-fluctuation driven Fermi surface instability.

cond-mat.supr-con

A multi-component Fermi surface in the vortex state of an underdoped high-Tc superconductor

In order to understand the origin of superconductivity, it is crucial to ascertain the nature and origin of the primary carriers available to participate in pairing. Recent quantum oscillation experiments on high Tc cuprate superconductors have revealed the existence of a Fermi surface akin to normal metals, comprising fermionic carriers that undergo orbital quantization. However, the unexpectedly small size of the observed carrier pocket leaves open a variety of possibilities as to the existence or form of any underlying magnetic order, and its relation to d-wave superconductivity. Here we present quantum oscillations in the magnetisation (the de Haas-van Alphen or dHvA effect) observed in superconducting YBa2Cu3O6.51 that reveal more than one carrier pocket. In particular, we find evidence for the existence of a much larger pocket of heavier mass carriers playing a thermodynamically dominant role in this hole-doped superconductor. Importantly, characteristics of the multiple pockets within this more complete Fermi surface impose constraints on the wavevector of any underlying order and the location of the carriers in momentum space. These constraints enable us to construct a possible density-wave scenario with spiral or related modulated magnetic order, consistent with experimental observations.

cond-mat.supr-con

Orbital Symmetries of Charge Density Wave Order in YBa2Cu3O6+x

Charge density wave (CDW) order has been shown to compete and coexist with superconductivity in underdoped cuprates. Theoretical proposals for the CDW order include an unconventional $d$-symmetry form factor CDW, evidence for which has emerged from measurements, including resonant soft x-ray scattering (RSXS) in YBa$_2$Cu$_3$O$_{6+x}$ (YBCO). Here, we revisit RSXS measurements of the CDW symmetry in YBCO, using a variation in the measurement geometry to provide enhanced sensitivity to orbital symmetry. We show that the $(0\ 0.31\ L)$ CDW peak measured at the Cu $L$ edge is dominated by an $s$ form factor rather than a $d$ form factor as was reported previously. In addition, by measuring both $(0.31\ 0\ L)$ and $(0\ 0.31\ L)$ peaks, we identify a pronounced difference in the orbital symmetry of the CDW order along the $a$ and $b$ axes, with the CDW along the $a$ axis exhibiting orbital order in addition to charge order.

cond-mat.supr-con