SearcharxivSearch

arXiv subjects

Qing Xu

Publications and source records attributed to Qing Xu.

At least 19 recordsLinked to original sources

Vision Guided Target Conditioned Control for Autonomous Excavation

Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4\% for no-condition ACT, 63\% for non-paired mask-conditioned ACT, and 96\% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8\% of the pile versus 27.4\% and 15.7\% for the two baselines, with 91.0\% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.

cs.RO

A novel strategy for achieving a low-field lightweight permanent MRI magnet system with good magnetic field homogeneity and low eddy current

In low-field, lightweight, pole-pieceless permanent-magnet MRI systems built with sintered Nd-Fe-B or Sm-Co magnets, the rapid switching of gradient fields readily induces eddy currents in the sintered magnets, leading to image artifacts. To address this, we report for the first time a Sm-Fe-N permanent-magnet MRI system based on anisotropic Sm-Fe-N bonded magnets, whose high electrical resistivity reduces the eddy currents in the X, Y and Z directions to 0.093%, 0.172% and 2.38%, respectively, while a magnetic field inhomogeneity below 150 ppm is achieved at the boundary of a 220 mm diameter of spherical volume (DSV). Compared with sintered Nd-Fe-B and Sm-Co magnets, using Sm-Fe-N bonded magnets as the source of the static magnetic field not only suppresses eddy currents but also makes a closely tiled, densely packed magnetic-circuit layout feasible, providing a more uniform static magnetic field for the MRI system. Imaging results free of obvious geometric distortion and banding artifacts further indicate that the Sm-Fe-N magnet system delivers low eddy currents and high static magnetic field homogeneity.

physics.med-ph

High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions

Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most current MOBO approaches are limited to low-dimensional decision space due to its exponential sampling complexity. This paper presents decision variable interaction analysis-based MOBO, ViaMOBO, a generic framework for expensive multi-objective problems with high-dimensional decision space. The key idea of ViaMOBO is that it utilizes a variable interaction analysis model to determine whether the decision space can be completely or partially divided, and then performs local Bayesian optimization in the divided decision subspaces. Through the variable analysis model, it can be derived whether the objectives in black-box problems are separable, partially separable, or non-separable based on the potential independent or interdependent relationships among decision variables without any strong assumptions. We compare ViaMOBO with the state-of-the-art MOBO methods on both synthetic and real-world benchmarks. The experimental results demonstrate that ViaMOBO outperforms other related MOBO baselines in approximating the Pareto front of high-dimensional expensive multi-objective problems.

cs.LG

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation methods rely on deterministic cross-modal matching, which overlooks aleatoric uncertainty from ambiguous boundaries and epistemic uncertainty from limited training data, leading to fragile performance under domain shift. To address this issue, we propose DistMedVL, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty. Specifically, the PCM-Adapter comprises two sequential modules for progressive probabilistic alignment. We first devise a Mahalanobis Alignment Module (MAM) that models textual tokens as Gaussian distributions and computes patch-text compatibility via Mahalanobis distance, yielding variance-conditioned matching that downweights unreliable feature dimensions. Moreover, we devise a Distribution Flow Module (DFM) that estimates modality-wise confidence parameters and performs vision-guided refinement of textual distributions, accommodating distributional variation across imaging modalities. Extensive experiments across eight medical segmentation benchmarks demonstrate that DistMedVL outperforms state-of-the-art methods with only 6.3M trainable parameters, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.

cs.CV

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

Cell segmentation is critical for computational pathology and biomedical discovery. While recent Vision Foundation Models (VFMs) have demonstrated remarkable universal feature representations, unlocking their full potential for cellular imaging is currently bottlenecked by resource-intensive adaptation paradigms. Existing methods typically rely on fine-tuning heavy visual encoders, leading to extensive computational overhead and a dependency on large-scale annotations. To address this, we propose the EffiCell-Seg framework for highly efficient cell segmentation without re-training the visual encoder. Our core insight is that pretrained VFMs intrinsically encode complementary structural priors: global saliency for localizing potential cells, and local morphological patterns for delineating cellular structures. To harness these priors, we devise a Cell Structure Prompt Encoder (CSP-Encoder) that synthesizes semantic-aware saliency and principal morphological features from frozen VFM representations into explicit structural prior maps. Moreover, we propose a Synergistic Mask Decoder (SM-Decoder) that enforces contextual consistency by jointly predicting geometric distance fields and semantic maps via mutual cross-guidance. Extensive experiments demonstrate that EffiCell-Seg outperforms state-of-the-art methods across diverse cell imaging modalities while requiring only ~5M trainable parameters, over 130x fewer than fully fine-tuned VFM counterparts. The code is available at https://github.com/xq141839/EffiCell-Seg.

cs.CV

Carbon Layer Orientation and Closed-Pore Construction Achieving Ultra-Low Specific Surface Area Hard Carbon for High-Performance Na-ion Storage

Addressing the critical trade-off between initial Coulombic efficiency (ICE) and reversible capacity in hard carbon anodes for Na-ion batteries (NIBs), we introduce a novel coupling strategy that combines carbon layer orientation reconstruction with closed-pore construction to produce hard carbon with an ultra-low specific surface area. We demonstrate that the nanographite domains within the hard carbon precursor undergo entropy-driven orientation reconstruction through the synergistic regulation of heteroatom doping and medium-temperature carbonization. This process not only increases interlayer spacing and promotes structural disorder but also enables the formation of dense, closed pores and ultramicropores at domain boundaries via confined atomic migration, while simultaneously encapsulating surface open pores within internal closed ones. Due to this unique pore architecture, our hard carbon exhibits an ultra-low specific surface area of 1.89 m2 g-1 with a markedly higher proportion of closed pores. As a result, our hard carbon achieves a remarkable reversible capacity of 342.3 mAh g-1 at 20 mA g-1, with an exceptional ICE of 90.4% and a dominant plateau capacity of 262.3 mAh g-1 (76.6%) for NIBs. We believe this coupling strategy provides a new paradigm for the structural engineering of high-ICE anode materials in advanced NIBs.

cond-mat.mtrl-sci

DYNA-PRUNER: Input-Adaptive Data-Model Co-Pruning for Efficient and Scalable Spatio-Temporal Media Prediction

Spatio-temporal prediction supports radar/satellite nowcasting and city-scale traffic monitoring, but modern models are often too expensive for real-time deployment. This stems from a mismatch between dense computation and strong input-dependent redundancy (e.g., calm seas or clear skies). To enable automated, resource-aware architecture optimization in scalable media analysis, we propose Dyna-Pruner, an end-to-end framework for input-dependent co-pruning of data and model structure. A shared-importance synchronization mechanism generates coupled masks that prune redundant regions and their corresponding computational units (e.g., convolutional filters), yielding per-sample sparse sub-networks at inference time. Experiments on WeatherBench, SEVIR, and TaxiBJ show seamless integration with CNN, RNN, and Transformer backbones, reducing FLOPs by up to $70\%$ and achieving a $2.5\times$ speedup on NVIDIA Jetson AGX Orin with negligible accuracy loss ($<1\%$).

cs.CV

Cesarean Scar Defect Segmentation in Transvaginal Ultrasound Images: a Dataset and Benchmark

Cesarean Scar Defect (CSD) is one of the most prevalent complications following cesarean delivery. Transvaginal ultrasonography is widely used for primary CSD screening. Accurate determination of CSD outline and dimensions is crucial for treatment. However, CSDs are frequently overlooked by sonographers due to small size and irregular morphology, suboptimal image quality, and limited clinical awareness in resource-constrained settings. Despite artificial intelligence advances in medical imaging, no public dataset exists for transvaginal ultrasound CSD segmentation. To address this gap, we present a comprehensive CSD dataset comprising 1,111 images and 16 videos, yielding 501 positive samples with confirmed CSD and precise pixel-level manual annotations. Annotations are performed following standardized clinical guidelines through collaboration between experienced sonographers and trained PhD students. This work provides high-quality benchmark resources for advancing medical image segmentation algorithms and promoting clinical innovation. Ultimately, improved CSD diagnosis and subsequent treatment strategies can enhance the quality of life in women of reproductive age, representing significant value for both medical research and clinical practice.

cs.CV

Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective

This paper explores a scientific question in supervised fine-tuning (SFT): why SFT is broadly effective for small-scale deep neural networks, yet can produce inconsistent or even detrimental effects when applied to large language models (LLMs). Recent advances in interaction-based explanations suggest that interactions between words/tokens provide a faithful metric for quantifying the inference patterns encoded by LLMs. We find that the evolution of interactions during SFT can effectively explain the inconsistent effectiveness of SFT for LLMs. Specifically, we find that (1) SFT primarily removes noise-like interactions, while rarely acquiring reliable new interactions. (2) This denoising stage is extremely brief, after which continued fine-tuning tends to introduce overfitted interactions. We validate these findings across multiple LLMs and datasets. Our findings provide new insights into early stopping and offer practical guidance for LLM training.

cs.AI

Stochastic Momentum Tracking Push-Pull for Decentralized Optimization over Directed Graphs

Decentralized optimization over directed networks is frequently challenged by asymmetric communication and the inherent high variance of stochastic gradients, which collectively cause severe oscillations and hinder algorithmic convergence. To address these challenges, we propose the Stochastic Momentum Tracking Push-Pull (SMTPP) algorithm, which tracks the momentum term rather than raw stochastic gradients within the Push-Pull architecture. This design successfully decouples the variance reduction capacity from the algebraic connectivity of the graph.Although the inherent topology mismatch of directed graphs precludes exact convergence under persistent stochastic noise, SMTPP rigorously compresses this unavoidable steady-state error floor into a minimal neighborhood determined by network connectivity and gradient variance. Furthermore, SMTPP guarantees convergence on any strongly connected directed graph. Extensive experiments on non-convex logistic regression demonstrate that the algorithm is highly robust to network connectivity. By effectively dampening topology-induced oscillations, SMTPP achieves convergence rates and overall performance that closely match those of centralized baselines, regardless of whether the network is sparse or dense.

math.OC

Improved Convergence for Decentralized Stochastic Optimization with Biased Gradients

Decentralized stochastic optimization has emerged as a fundamental paradigm for large-scale machine learning. However, practical implementations often rely on biased gradient estimators arising from communication compression or inexact local oracles, which severely degrade convergence in the presence of data heterogeneity. To address the challenge, we propose Decentralized Momentum Tracking with Biased Gradients (Biased-DMT), a novel decentralized algorithm designed to operate reliably under biased gradient information. We establish a comprehensive convergence theory for Biased-DMT in nonconvex settings and show that it achieves linear speedup with respect to the number of agents. The theoretical analysis shows that Biased-DMT decouples the effects of network topology from data heterogeneity, enabling robust performance even in sparse communication networks. Notably, when the gradient oracle introduces only absolute bias, the proposed method eliminates the structural heterogeneity error and converges to the exact physical error floor. For the case of relative bias, we further characterize the convergence limit and show that the remaining error is an unavoidable physical consequence of locally injected noise. Extensive numerical experiments corroborate our theoretical analysis and demonstrate the practical effectiveness of Biased-DMT across a range of decentralized learning scenarios.

math.OC

XAttnRes: Cross-Stage Attention Residuals for Medical Image Segmentation

In the field of Large Language Models (LLMs), Attention Residuals have recently demonstrated that learned, selective aggregation over all preceding layer outputs can outperform fixed residual connections. We propose Cross-Stage Attention Residuals (XAttnRes), a mechanism that maintains a global feature history pool accumulating both encoder and decoder stage outputs. Through lightweight pseudo-query attention, each stage selectively aggregates from all preceding representations. To bridge the gap between the same-dimensional Transformer layers in LLMs and the multi-scale encoder-decoder stages in segmentation networks, XAttnRes introduces spatial alignment and channel projection steps that handle cross-resolution features with negligible overhead. When added to existing segmentation networks, XAttnRes consistently improves performance across four datasets and three imaging modalities. We further observe that XAttnRes alone, even without skip connections, achieves performance on par with the baseline, suggesting that learned aggregation can recover the inter-stage information flow traditionally provided by predetermined connections.

cs.CV

Electronic excitation of ultrafast collective amorphous-amorphous transitions in glassy phase-change material

The intrinsic nature of glass states and glass transitions remain a fundamental open question in condensed-matter physics and materials science. The key to solving the glass transition problem lies in achieving a complete understanding of the physics governing the structural relaxation. Nonetheless, directly probing dynamic atomic-scale structural changes in order to identify the precise local structural motifs and establish quantitative structure-property relationships remains an outstanding challenge. By combining femtosecond electron diffraction with time-dependent density-functional theory molecular dynamics simulations, we directly capture ultrafast amorphous-amorphous transitions indicated by collective bond stretching (0.2 ps) and angle bending (0.5-2 ps) in glassy phase-change material GeTe. The ultrafast bond stretching is accompanied by localized oscillation modes with the frequency of 3.10 THz, unambiguously signaling the local Peierls-like bonding structure and the flexibility of these polarized bonds. These ultrafast collective atomic motions, captured across timescales ranging from femtoseconds to picoseconds, directly reveals the structural origin of the boson peak and provide compelling evidence for many-body interactions in amorphous materials. Furthermore, the ultrafast amorphous-amorphous transitions induce a drastic insulator-metal transition, directly revealing both the underlying switching mechanism and the fundamental speed limit of the ovonic threshold switch. These insights establish a fundamental framework for rationally engineering relaxation pathways and phase-change/threshold switch in amorphous materials. Femtosecond electron diffraction provides a powerful novel approach to deciphering the structural complexity and functional mechanisms of amorphous materials by resolving collective atomic motions from random diffusion dynamics in the time domain.

cond-mat.mtrl-sci

From Optimizable to Interactable: Mixed Digital Twin-Empowered Testing of Vehicle-Infrastructure Cooperation Systems

Sufficient testing under corner cases is critical for the long-term operation of vehicle-infrastructure cooperation systems (VICS). However, existing corner-case generation methods are primarily AI-driven, and VICS testing under corner cases is typically limited to simulation. In this paper, we introduce an L5 ''Interactable'' level to the VICS digital twin (VICS-DT) taxonomy, extending beyond the conventional L4 ''Optimizable'' level. We further propose an L5-level VICS testing framework, IMPACT (Interactive Mixed-digital-twin Paradigm for Advanced Cooperative vehicle-infrastructure Testing). By enabling direct human interactions with VICS entities, IMPACT incorporates highly uncertain and unpredictable human behaviors into the testing loop, naturally generating high-quality corner cases that complement AI-based methods. Furthermore, the mixedDT-enabled ''Physical-Virtual Action Interaction'' facilitates safe VICS testing under corner cases, incorporating real-world environments and entities rather than purely in simulation. Finally, we implement IMPACT on the I-VIT (Interactive Vehicle-Infrastructure Testbed), and experiments demonstrate its effectiveness. The experimental videos are available at our project website: https://dongjh20.github.io/IMPACT.

cs.RO

Sub-angstrom many-body localization driven by phononic flat bands in real quantum materials

Defects, fluctuations, degenerate states and correlated interactions facilitate the emergence of exotic properties in condensed matter systems while also inducing atomic-scale local correlated structures that deviate from the average long-range order. Establishing the structure-property relationship from the perspective of these atomic-scale local correlated structures remains ambiguous and controversial due to the lack of direct methods for identifying such local correlated structures. In this work, based on the photoexcited ultrafast structural response, we propose a Bragg scattering phase breaking regime to identify sub-angstrom local correlated structures in quantum materials. With this regime, we unambiguously identify the many-body-interaction driven local correlated structures in the low temperature ground state of AgCrSe2, characterized by static off-center displacements of Ag atoms ranging from 0 to 0.5 angstrom. The competition between Ag-Ag Coulomb correlations and potential wells induced by CrSe2 layers, leading to phononic flat bands and driving the system into a many body localization (MBL) regime. As temperature rising, these static local correlated structures transform to a dynamic state where the thermal fluctuations overwhelm the multiple localized states. These distinctive local correlated structures constitute the first experimental observation of MBL with vortex-like topological characteristic in a real material system. Emergent vibrational modes arising from MBL have been confirmed and show excellent agreement with inelastic neutron scattering experiments. Our work not only offers a universal approach to characterize sub-angstrom local correlated structures across a wide range of quantum materials but also deepens our understanding of the fundamental mechanism behind exotic properties from the perspective of atomic-scale local correlated structures.

cond-mat.mtrl-sci

Multi-Source Human-in-the-Loop Digital Twin Testbed for Connected and Autonomous Vehicles in Mixed Traffic Flow

In the emerging mixed traffic environments, Connected and Autonomous Vehicles (CAVs) have to interact with surrounding human-driven vehicles (HDVs). This paper introduces MSH-MCCT (Multi-Source Human-in-the-Loop Mixed Cloud Control Testbed), a novel CAV testbed that captures complex interactions between various CAVs and HDVs. Utilizing the Mixed Digital Twin concept, which combines Mixed Reality with Digital Twin, MSH-MCCT integrates physical, virtual, and mixed platforms, along with multi-source control inputs. Bridged by the mixed platform, MSH-MCCT allows human drivers and CAV algorithms to operate both physical and virtual vehicles within multiple fields of view. Particularly, this testbed facilitates the coexistence and real-time interaction of physical and virtual CAVs \& HDVs, significantly enhancing the experimental flexibility and scalability. Experiments on vehicle platooning in mixed traffic showcase the potential of MSH-MCCT to conduct CAV testing with multi-source real human drivers in the loop through driving simulators of diverse fidelity. The videos for the experiments are available at our project website: https://dongjh20.github.io/MSH-MCCT.

cs.RO

A two-steps tensor eigenvector centrality for nodes and hyperedges in hypergraphs

Hypergraphs have been a powerful tool to represent higher-order interactions, where hyperedges can connect an arbitrary number of nodes. Quantifying the relative importance of nodes and hyperedges in hypergraphs is a fundamental problem in network analysis. In this paper, we propose a new tensor-based centrality measure for general hypergraphs. We use a third-order tensor to represent the relationship between nodes and hyperedges. The tensor's positive Perron vector is defined as the centrality vector of the hypergraph. The existence and uniqueness of this centrality vector are guaranteed by the Perron-Frobenius theorem for tensors. This new centrality measure captures a higher-order mutual reinforcement mechanism: a node's importance is determined by the importance of its incident hyperedges and the other nodes within these hyperedges; symmetrically, a hyperedge's importance is determined by the importance of its constituent nodes and the other hyperedges containing these nodes. We further provide a combinatorial interpretation by proving that the centrality vector represents the limit geometric capacity of two-steps expansion trees. We illustrate the centrality measure on real-world hypergraph datasets.

cs.SI

WaveRNet: Wavelet-Guided Frequency Learning for Multi-Source Domain-Generalized Retinal Vessel Segmentation

Domain-generalized retinal vessel segmentation is critical for automated ophthalmic diagnosis, yet faces significant challenges from domain shift induced by non-uniform illumination and varying contrast, compounded by the difficulty of preserving fine vessel structures. While the Segment Anything Model (SAM) exhibits remarkable zero-shot capabilities, existing SAM-based methods rely on simple adapter fine-tuning while overlooking frequency-domain information that encodes domain-invariant features, resulting in degraded generalization under illumination and contrast variations. Furthermore, SAM's direct upsampling inevitably loses fine vessel details. To address these limitations, we propose WaveRNet, a wavelet-guided frequency learning framework for robust multi-source domain-generalized retinal vessel segmentation. Specifically, we devise a Spectral-guided Domain Modulator (SDM) that integrates wavelet decomposition with learnable domain tokens, enabling the separation of illumination-robust low-frequency structures from high-frequency vessel boundaries while facilitating domain-specific feature generation. Furthermore, we introduce a Frequency-Adaptive Domain Fusion (FADF) module that performs intelligent test-time domain selection through wavelet-based frequency similarity and soft-weighted fusion. Finally, we present a Hierarchical Mask-Prompt Refiner (HMPR) that overcomes SAM's upsampling limitation through coarse-to-fine refinement with long-range dependency modeling. Extensive experiments under the Leave-One-Domain-Out protocol on four public retinal datasets demonstrate that WaveRNet achieves state-of-the-art generalization performance. The source code is available at https://github.com/Chanchan-Wang/WaveRNet.

cs.CV