SearcharxivSearch

arXiv subjects

Yi Yu

Publications and source records attributed to Yi Yu.

At least 37 records · Page 2Linked to original sources

Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation

Accurate segmentation of anatomical structures in medical images is essential for diagnosis and treatment planning. While recent interactive segmentation foundation models enhance generalization through large-scale multimodal pretraining, they still depend on precise prompts and can fail in underrepresented clinical contexts (e.g., small organs-at-risk). We present AtlasSegFM, an atlas-guided framework that customizes off-the-shelf foundation models to new clinical contexts with a single annotated example. AtlasSegFM 1) performs atlas-query registration to generate context-aware prompts, 2) refines the segmentation with a frozen foundation model, and 3) applies a lightweight adaptive fusion module to combine atlas priors with foundation-model inputs and predictions. Extensive experiments on six public and in-house datasets across radiotherapy and vascular scenarios show consistent gains, with the largest improvements on small and delicate structures. AtlasSegFM provides a lightweight, deployable solution for one-shot customization of segmentation foundation models in real-world clinical workflows.

cs.CV

Towards Voxel Spacing Consistency for Medical Image Segmentation

Volumetric medical image segmentation is essential for both preoperative diagnosis and intraoperative guidance. While recent years have witnessed rapid progress in segmentation architectures, comparatively little attention is paid to the physical voxel spacing of anatomical data. Indeed, volumetric image resampling is a ubiquitous preprocessing step before segmentation, yet its interaction with downstream segmentation has not been systematically exploited. In this work, we study the correlation between image resampling and segmentation, and propose Consispace, a semantic-aware resampling framework that achieves consistent voxel spacing in the axial direction while preserving anatomical and semantic consistency. Consispace introduces an ODE-based anatomical constraint to model inter-slice dynamics with a continuous interpolator, enabling faithful reconstruction under complex anatomical transitions beyond discrete interpolation. To further couple resampling with segmentation objectives, we leverage dense features from a pretrained vision model to build intra-slice semantic correlation maps and inject class-wise semantic consistency via feature reweighting during resampling. Both intra-slice and inter-slice constraints are integrated into an implicit neural network, supporting arbitrary-scale resampling. Extensive experiments on multiple datasets demonstrate that Consispace achieves superior reconstruction quality and perceptual fidelity, produces smoother inter-slice anatomy, and improves downstream segmentation performance when used as a preprocessing step.

cs.CV

Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence

Current embodied world models are primarily optimized for predictive objectives, limiting their ability to generalize under distribution shifts and reason systematically about unseen situations and hypothetical interventions. We argue that embodied intelligence should move beyond predictive world modeling toward self-evolving cognitive systems that continually construct and refine internal causal representations through interaction with the environment. To this end, we propose a self-evolving cognitive framework via causal world modeling for embodied scientific intelligence, which integrates three complementary components: causal world modeling, intervention-driven causal reasoning, and continual cognitive refinement. The proposed framework continuously revises and expands its internal causal world model through causal discovery, intervention-driven feedback, and counterfactual reasoning, supporting continual cognitive refinement and enabling cognition itself to evolve over time. Furthermore, we reinterpret embodied interaction not merely as a means of trajectory optimization, but as an epistemic process for causal hypothesis generation, intervention-driven experimentation, and continual knowledge acquisition. This work provides a conceptual and theoretical foundation for a transition from predictive intelligence toward epistemic intelligence, in which intelligence emerges through the continual construction, revision, and refinement of causal world models via interaction with the environment. Accordingly, an intervention-driven causal-epistemic benchmarking paradigm is suggested for evaluating self-evolving embodied scientific intelligence.

cs.AI

ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization

Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization techniques allow such models to imitate a specific subject, style, or motion pattern from only a few reference clips. This capability creates a new data-misuse risk: videos shared online can be collected and used for unauthorized T2V fine-tuning. Existing protective perturbations are mainly designed for image recognition or text-to-image personalization, and therefore focus on corrupting static appearance cues rather than the temporal denoising dynamics that make video personalization possible. To address this gap, we introduce ChronoLock, the first proactive protection framework that makes released videos difficult to exploit for unauthorized T2V personalization. ChronoLock targets the motion-learning process directly by optimizing bounded perturbations over temporal denoising trajectories. It first disrupts intra-chunk temporal adaptation with a diffusion objective that combines fitting error, frame-relative denoising relations, and adjacent-frame variation, and then enlarges inter-chunk boundary mismatch to weaken long-range motion continuity. Transformation-sampled updates further improve robustness to common preprocessing operations.Experiments on UCF Sports and HMDB51 with popular T2V backbones and personalization scheme show that ChronoLock effectively reduces motion imitation under automatic metrics and human evaluation.

cs.CV

Quasi-Monte Carlo finite element approximation for singularly perturbed convection-diffusion problems with random velocity

This paper studies the numerical approximation of a singularly perturbed convection-diffusion problem over a bounded polygonal domain in $\mathbb{R}^d$ ($d=2,3$), where the velocity field is modeled by a log-uniform random field, a setting typical in uncertainty quantification. We introduce a novel numerical framework for computing the expected value of the linear functionals of the solution. The approach combines a finite element discretization of the problem, a truncated Karhunen--Loève expansion to represent the stochastic velocity field, and a lattice-based quasi-Monte Carlo (QMC) method to estimate expectations over the parameter space. We provide a rigorous error analysis of the proposed scheme, establishing bounds on the mean squared error and demonstrating that the QMC method achieves a nearly linear optimal convergence rate, with a constant independent of the integration dimension. Furthermore, the convergence rate is shown to be independent of the singular perturbation parameter.

math.NA

Nanobeam Laser Cavities with High Quality-factor and Near-Unity Outcoupling Efficiency

Cavities with high quality (Q) factor and small mode-volume are crucial to realize high-performance nanolasers suitable for optical interconnects. In this work, we propose a novel one-dimensional photonic crystal nanobeam cavity design with fins for controlled electrical injection into the active region. An effective optimization algorithm based on first-order perturbation theory of quasinormal modes is implemented and shown to significantly enhance the cavity quality factor. The one-dimensional geometry of the cavity lends itself to unidirectional coupling of the resonant mode into the waveguide by introducing asymmetry of the mirror. The resulting design is shown to achieve high extraction efficiencies ($>90\%$) while maintaining a high Q-factor ($>10 \cdot 10^3$). Through an analysis of the cavity's decay channels, we find that the introduced asymmetry induces unexpected interactions between the cavity's decay channels. Passive InP cavities are fabricated and experimentally characterized, demonstrating record-high quality factors exceeding $170 \cdot 10^3$ for designs without fins and up to $70 \cdot 10^3$ for designs with fins, confirming the efficacy of the optimization method and quality of the fabrication process.

physics.optics

Ionization-Induced Electrostatic Hose Instability in Electron-Beam-Sustained Plasmas

We report the discovery of a previously unrecognized electrostatic hose instability in electron-beam-sustained plasmas, driven by the coupling between the electron beam centroid and the plasma generated via the beam-impact ionization. Unlike the conventional hose instability of relativistic beams propagating in underdense plasmas, this instability requires only ionization-capable electron beams readily produced by common emission processes and sheath acceleration, indicating broad relevance across various discharges. A linear theory is developed to predict the hosing frequency and growth rate, and particle-in-cell/Monte Carlo simulations confirm both the onset of instability and the theoretical predictions.

physics.plasm-ph

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. However, existing traffic-oriented multimodal benchmarks largely emphasize passive visual recognition or isolated video understanding, offering limited support for evaluating structure-aware traffic reasoning under controlled conditions. We introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built around 12 real-world intersections reconstructed into editable 3D traffic environments and complemented by surveillance footage from two countries, OmniTraffic supports both controlled and natural-condition evaluation. It defines a three-level task hierarchy spanning scene perception, multi-view and temporal reasoning, and decision support. Using structured traffic metadata, OmniTraffic generates synchronized multi-view VQA samples covering vehicle states, lane functions, view--BEV correspondence, temporal dynamics, and signal-phase analysis, resulting in 8M VQA samples and a 3K human-verified test set. Evaluation of eleven frontier MLLMs reveals a large human--model gap, with the most pronounced failures in topology-grounded and spatio-temporal reasoning tasks. Fine-tuning a lightweight MLLM on simulated OmniTraffic data further improves performance on real-world traffic scenes, demonstrating the value of simulation-generated supervision for traffic-specific multimodal reasoning. Beyond a fixed dataset, OmniTraffic provides an extensible pipeline with configurable intersections, camera views, traffic demands, signal phases, visual conditions, and rare events.

cs.CV

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored. We present a standardized real-world benchmark for evaluating representative VLA and imitation learning policies on the low-cost SO-101 robotic platform. The benchmark comprises four representative manipulation tasks together with unified evaluation protocols, enabling systematic comparison under embodiment uncertainty. Using real-world teleoperated demonstrations, we fine-tune and evaluate $π_{0.5}$, SmolVLA, Wall-X, and ACT directly on the physical platform. Beyond conventional task success rates, the benchmark incorporates a structured failure taxonomy, semantic- and execution-level failure decomposition, and recovery-aware evaluation metrics to characterize policy robustness. Experimental results show that stronger pretrained VLA policies generally outperform the imitation learning baseline, although performance remains highly task-dependent under low-cost robotic deployment conditions. Execution instability emerges as the dominant failure source, while recovery capability varies substantially across architectures. These results highlight the importance of failure and recovery analysis beyond binary task success and establish SO-101 as a practical benchmark for evaluating embodied AI systems under realistic low-cost robotic deployment conditions.

cs.RO

Online change point detection under heavy-tailedness and contamination

We study an online version of the robust mean change point detection problem under a dynamic Huber contamination model with arbitrary contamination distribution and inlier distribution possessing exponentially- or polynomially-decaying tails. This robustness framework is systematically studied for the first time in the change point literature. For univariate data, we characterise the detection delay by partitioning the parameter space into four regimes, in terms of the true change location, signal size and contamination level. Efficient detection procedures are accompanied by matching lower bounds, up to poly-logarithmic factors. For the multivariate setting, we devise an efficient robust mean testing procedure and apply this to the robust online change point problem. The theoretical analysis of the robust mean testing procedure is the first in dealing with both Huber contamination and heavy-tailedness, and is thus of independent interest. Extensive numerical experiments are conducted to support our theoretical findings.

math.ST

MemVerse: Multimodal Memory for Lifelong Learning Agents

Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember. Without reliable memory, agents catastrophically forget past experiences, struggle with long-horizon reasoning, and fail to operate coherently in multimodal or interactive environments. We introduce MemVerse, a model-agnostic, plug-and-play memory framework that bridges fast parametric recall with hierarchical retrieval-based memory, enabling scalable and adaptive multimodal intelligence. MemVerse maintains short-term memory for recent context while transforming raw multimodal experiences into structured long-term memories organized as hierarchical knowledge graphs. This design supports continual consolidation, adaptive forgetting, and bounded memory growth. To handle real-time demands, MemVerse introduces a periodic distillation mechanism that compresses essential knowledge from long-term memory into the parametric model, allowing fast, differentiable recall while preserving interpretability. Extensive experiments demonstrate that MemVerse significantly improves multimodal reasoning and continual learning efficiency, empowering agents to remember, adapt, and reason coherently across extended interactions.

cs.AI

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall short for the long tail of real-world care not covered by guidelines. Most medical large language models (LLMs), however, are trained to encode common, guideline-focused medical knowledge in their parameters. Current evaluations test models primarily on recalling and reasoning with this memorized content, often in multiple-choice settings. Given the fundamental importance of evidence-based reasoning in medicine, it is neither feasible nor reliable to depend on memorization in practice. To address this gap, we introduce OGCaReBench, a free-form retrieval-focused benchmark aimed at evaluating LLMs at answering clinical questions that require going beyond typical guidelines. Extracted from published medical case reports and validated by medical experts, OGCaReBench contains long-form clinical questions requiring free-text answers, providing a systematic framework for assessing open-ended medical reasoning in rare, case-based scenarios. Our experiments reveal that even the best-performing baseline (GPT-5.2) correctly answers only 56% of our benchmark with specialized models only reaching 42%. Augmenting models with retrieved medical articles improves this performance to up to 82% (using GPT-5.2) highlighting the importance of evidence-grounding for real-world medical reasoning tasks. This work thus establishes a foundation for benchmarking and advancing both general-purpose and medical LLMs to produce reliable answers in challenging clinical contexts.

cs.CL

Differentially private hypothesis testing in survival analysis

Survival analysis is widely used in applications involving sensitive individual-level data, yet differentially private hypothesis testing for right-censored data remains largely undeveloped. We initiate a finite-sample theory of private hypothesis testing in survival analysis applications. For Cox regression coefficients, we develop private partial-likelihood-ratio and score-type tests, including a private calibration procedure for the rejection threshold. For cumulative hazard functions, we propose a private distributed two-sample test. Across these problems, we prove differential privacy and finite-sample testing guarantees, as well as minimax lower bounds. Our results identify when privacy is statistically negligible, when it dominates the testing rate, and where optimal private rates for testing in semiparametric survival models remain open. This theoretical analysis is accompanied by numerical experiments on simulated data.

math.ST

STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation

SNNs promise energy-efficient and low-latency inference, but their performance still trails that of ANNs. ANN-to-SNN knowledge distillation helps narrow this gap, yet the original training data are often unavailable in practical deployment settings. Existing data-free knowledge distillation (DFKD) methods synthesize surrogate data by matching teacher-side priors, especially BN statistics, but these ANN-oriented constraints mainly regularize mean and variance and therefore remain under-constrained for SNN students whose responses depend on threshold-crossing dynamics. In this paper, we propose Spike Tail-Aware Relational Synthesis (STARS), a plug-and-play method for ANN-to-SNN DFKD that augments standard BN-guided synthesis with two complementary objectives: Relational Consistency Alignment, which preserves cross-sample relational consistency between teacher and student, and Tail-Aware Regularization, which regularizes threshold-relevant tail probabilities through soft exceedance over teacher-derived thresholds. Together, these objectives generate synthetic batches that remain teacher-valid while becoming more informative for SNN students. Experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet across multiple ANN-SNN pairs show that our method consistently improves conventional DFKD baselines and even surpasses several KD methods, with gains of up to 4.6\% on CIFAR-10 and 6.7\% on CIFAR-100, highlighting the importance of complementing BN matching with relational and tail-aware constraints in SNN-oriented DFKD.

cs.NE

Uncertainty-Aware Structured Data Extraction from Full CMR Reports via Distilled LLMs

Converting free-text cardiac magnetic resonance (CMR) reports into auditable structured data remains a bottleneck for cohort assembly, longitudinal curation, and clinical decision support. We present CMR-EXTR, a lightweight framework that converts free-text CMR reports into structured data and assigns per-field confidence for quality control. A teacher-student distillation pipeline enables fully offline inference while limiting manual annotation. Uncertainty integrates three complementary principles -- distribution plausibility, sampling stability, and cross-field consistency -- to triage human review. Experiments show that CMR-EXTR achieves 99.65% variable-level accuracy, demonstrating both reliable extraction and informative confidence scores. To our knowledge, this is the first CMR-specific extraction system with integrated confidence estimation. The code is available at https://github.com/yuyi1005/CMR-EXTR.

cs.CL

Online Learning for Autoregressive Multilayer Stochastic Block Models under Stationarity and Non-Stationarity

Dynamic multilayer networks arise in many applications where multiple types of relations among a common set of nodes evolve over time. Existing approaches often assume temporal independence, focus on single-layer networks or impose stationarity, limiting their applicability in practice. In this paper, we introduce a first-order autoregressive multilayer stochastic block model (AR(1)-MSBM), in which edge formation and dissolution probabilities between consecutive time points are determined by latent community memberships and shared across layers. Under stationarity, we propose an online estimation procedure based on recursive updates and tensor-based spectral refinement. We establish non-asymptotic estimation rates, prove their minimax optimality and derive guarantees for community recovery. We further consider a non-stationary setting that allows both abrupt changes and gradual shifts, and develop an adaptive windowed online algorithm that automatically adjusts to unknown structural changes. Under a quasi-stationary segmentation framework, we derive estimation and community recovery guarantees that match the stationary results when applied segmentwise. Our theoretical findings are supported by extensive numerical experiments, with code available online.

stat.ME

LAMOST J052016.79+345651.7: An EW-type Binary with Emission Line Spectra and Circumstellar Material

LAMOST J052016.79+345651.7 was identified as an EW-type eclipsing binary by Chen et al. when studying the periodic variable stars based on the ZTF telescope. An orbital period of 0.3507818 days has been reported. Using the ZTF g, r, i band light curves, we reproduced the orbital period and obtained a phase folded diagram. The multi-band apparent magnitudes from Pan-STARRS, 2MASS, and WISE, the color indices, and the infrared excess in the WISE w4 band all indicate that LAMOST J0520 contains cold circumstellar material. All 19 LRS from LAMOST exhibit prominent Halpha emission lines, along with clear N II and S II emission lines, indicating that LAMOST J0520 contains optically thin, warm ionized gas with extremely low electron density. From the 19 spectra, an effective temperature of 6200 \pm 400 K can be obtained. Therefore, the emission lines likely originate from shock-driven outflows and mass ejection, while the infrared excess likely comes from dust formed as the outflows cool. We performed light curve fitting and evolutionary simulation study on LAMOST J0520 using PHOEBE and MESA, respectively. The phase shift of the fitted model can be uniquely determined, but the inclination and the fillout_{factor} are degenerate. Considering the angular momentum loss process, the evolutionary simulation from detached binaries to contact binaries can reflect the main evolutionary process of LAMOST J0520. When the orbital period of the evolutionary model is 0.351 days, the other parameters are generally consistent with the observed characteristics of LAMOST J0520. Considering mass loss, LAMOST J052016.79+345651.7 will evolve into a blue straggler star in the future. Future LAMOST medium resolution time-domain spectra of LAMOST J0520 may offer an opportunity to reveal more detailed physical processes.

astro-ph.SR

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current guardrail models lack agentic risk awareness and transparency in risk diagnosis. To introduce an agentic guardrail that covers complex and numerous risky behaviors, we first propose a unified three-dimensional taxonomy that orthogonally categorizes agentic risks by their source (where), failure mode (how), and consequence (what). Guided by this structured and hierarchical taxonomy, we introduce a new fine-grained agentic safety benchmark (ATBench) and a Diagnostic Guardrail framework for agent safety and security (AgentDoG). AgentDoG provides fine-grained and contextual monitoring across agent trajectories. More Crucially, AgentDoG can diagnose the root causes of unsafe actions and seemingly safe but unreasonable actions, offering provenance and transparency beyond binary labels to facilitate effective agent alignment. AgentDoG variants are available in three sizes (4B, 7B, and 8B parameters) across Qwen and Llama model families. Extensive experimental results demonstrate that AgentDoG achieves state-of-the-art performance in agentic safety moderation in diverse and complex interactive scenarios. All models and datasets are openly released.

cs.AI