SearcharxivSearch

arXiv subjects

Zhihao Zhao

Publications and source records attributed to Zhihao Zhao.

At least 19 recordsLinked to original sources

Physical Kernel: Structured Visual Latents for Dark Manipulation

We study dark manipulation: after a brief lit Write encodes z0 = Enc(rgb), a policy pi(z) and open-loop dynamics f(z,a) complete contact-rich skills without further pixels (dark_f). On ManiSkill StackCube (n=160; seed packs 0/1000), dark_f attains 68.1% stacked on the five-rung chain (near_A -> grasped -> lifted -> on_B -> stacked), compared with 35.6% for per-step lit_reenc and 0% for freeze/encode_black. On a shared Write->HOLD protocol (n=40), occlusion and camera-aligned GT contact-neighbor masks drive lit lift from 43% to 0%, while dark_f holds 82.5%; shuffling actions inflates dynamics MSE by ~9.4x; write-time appearance shifts break encoding (night: 0% stacked), yet the same shifts during HOLD leave dark_f lift unchanged; Write length Tw is flat once the stop phase is reached, while earlier stops and write-time blur/JPEG sharply cut stacked. A dedicated pi_write reaches 35% vision-budget stacked (n=80); matched Dreamer-style/pixel nulls without privileged geom stay at 0%. Privileged state-RSSM MPC reaches ~35% stacked with 9D dark observations -- a stronger-observation null, not a matched visual baseline.

cs.RO

Automated 2D and 3D Segmentation of AMD and DME Lesions in OCT

Age-related macular degeneration (AMD) and diabetic macular edema (DME) are leading causes of vision loss, and optical coherence tomography (OCT) is the standard modality for detecting and monitoring the subtle lesions that drive treatment decisions. Most deep-learning segmentation work for OCT is validated only in-domain, leaving generalization to clinical data collected under different acquisition protocols largely untested. This work develops and systematically ablates four lesion-segmentation pipelines -- 2D and 3D variants for AMD and DME -- reaching Dice scores of 0.76 to 0.82 with strong volumetric and surface calibration (r vol, r surf greater than or equal to 0.97 across all four pipelines) on an in-domain validation set. The ablation process establishes a full-volume, calibration-aware adoption standard that catches mechanisms an ordinary slice-level evaluation would keep, and identifies ensemble composition as the most consistent driver of improvement. To test generalization, the models are evaluated on OLIVES, an external clinical cohort with no lesion-level ground truth, using a proxy-metric framework built around biomarker AUROC, central subfield thickness (CST) correlation, and longitudinal concordance. Predictions track clinical biomarkers outside the training distribution, though less strongly than in-domain -- evidence for, not validation of, automated lesion-burden tracking as a clinical tool.

cs.CV

On semi-stable integral models for Shimura varieties

We construct (potentially) semi-stable integral models for a class of Shimura varieties with maximal parahoric level at an odd prime $p$. The main input is an explicit construction of semi-stable equivariant modifications of the relevant canonical local models, where the underlying groups are Weil restrictions of unramified unitary similitude groups, symplectic similitude groups, or even orthogonal similitude groups. Via the local model diagram, these constructions give (potentially) semi-stable integral models of the corresponding Shimura varieties. In particular, the resulting models are regular and have reduced special fiber with normal crossings. As an application, we deduce the unipotence of the inertia action on nearby cycles and the $\ell$-adic cohomology of the geometric generic fibers.

math.NT

On $p$-adic integral moduli schemes and local models for PEL type D

We construct flat integral moduli schemes of PEL type D and the corresponding flat orthogonal Rapoport--Zink spaces with parahoric level structure over a $p$-adic integer ring. The construction relies on proving a conjecture of Pappas--Rapoport: for an even orthogonal similitude group over a complete discretely valued field of residue characteristic $p>2$, and for arbitrary parahoric level, the associated spin local model is flat, normal, Cohen--Macaulay, with reduced special fiber. In the course of the proof, we also show that in the quasi-split but non-split case, the Rapoport--Zink (naive) local model is topologically flat, verifying a conjecture of Pappas--Rapoport--Smithling. In the maximal parahoric case, we also describe the Schubert varieties in the special fiber in moduli-theoretic terms. Finally, for a maximal parahoric case we construct an explicit regular semi-stable model by blowing up the spin local model along the unique closed Schubert cell in its special fiber.

math.NT

KaLDeX: Kalman Filter based Linear Deformable Cross Attention for Retina Vessel Segmentation

Background and Objective: In the realm of ophthalmic imaging, accurate vascular segmentation is paramount for diagnosing and managing various eye diseases. Contemporary deep learning-based vascular segmentation models rival human accuracy but still face substantial challenges in accurately segmenting minuscule blood vessels in neural network applications. Due to the necessity of multiple downsampling operations in the CNN models, fine details from high-resolution images are inevitably lost. The objective of this study is to design a structure to capture the delicate and small blood vessels. Methods: To address these issues, we propose a novel network (KaLDeX) for vascular segmentation leveraging a Kalman filter based linear deformable cross attention (LDCA) module, integrated within a UNet++ framework. Our approach is based on two key components: Kalman filter (KF) based linear deformable convolution (LD) and cross-attention (CA) modules. The LD module is designed to adaptively adjust the focus on thin vessels that might be overlooked in standard convolution. The CA module improves the global understanding of vascular structures by aggregating the detailed features from the LD module with the high level features from the UNet++ architecture. Finally, we adopt a topological loss function based on persistent homology to constrain the topological continuity of the segmentation. Results: The proposed method is evaluated on retinal fundus image datasets (DRIVE, CHASE_BD1, and STARE) as well as the 3mm and 6mm of the OCTA-500 dataset, achieving an average accuracy (ACC) of 97.25%, 97.77%, 97.85%, 98.89%, and 98.21%, respectively. Conclusions: Empirical evidence shows that our method outperforms the current best models on different vessel segmentation datasets. Our source code is available at: https://github.com/AIEyeSystem/KalDeX.

eess.IV

The basic locus of regular ramified unitary Rapoport-Zink spaces at vertex-stabilizer level

We construct the Bruhat-Tits stratification of the reduced basic locus of regular ramified unitary Rapoport-Zink spaces of signature $(n\!-\!1,1)$ at vertex-stabilizer level. To study the Bruhat-Tits strata, we introduce strata models$\!-\!$simpler models that are étale-locally isomorphic to each stratum. They admit two complementary characterizations: (i) as strict transforms under the blow-up of the local model at its worst point, and (ii) via a partial moduli description given by explicit linear-algebraic conditions; from these we deduce smoothness, explicit dimension formulas and irreducibility of the Bruhat-Tits strata.

math.NT

The basic locus of unitary splitting Rapoport-Zink spaces with vertex stabilizer level

We construct the Bruhat-Tits stratification of the ramified unitary splitting Rapoport-Zink space, with the level being the stabilizer of a vertex lattice. To determine certain local properties of the Bruhat-Tits strata, we develop a theory of the strata splitting models. To study their global structure, we establish an explicit isomorphism between the Bruhat-Tits strata and certain (modified) Deligne-Lusztig varieties.

math.NT

Dynamic Structural Recovery Parameters Enhance Prediction of Visual Outcomes After Macular Hole Surgery

Purpose: To introduce novel dynamic structural parameters and evaluate their integration within a multimodal deep learning (DL) framework for predicting postoperative visual recovery in idiopathic full-thickness macular hole (iFTMH) patients. Methods: We utilized a publicly available longitudinal OCT dataset at five stages (preoperative, 2 weeks, 3 months, 6 months, and 12 months). A stage specific segmentation model delineated related structures, and an automated pipeline extracted quantitative, composite, qualitative, and dynamic features. Binary logistic regression models, constructed with and without dynamic parameters, assessed their incremental predictive value for best-corrected visual acuity (BCVA). A multimodal DL model combining clinical variables, OCT-derived features, and raw OCT images was developed and benchmarked against regression models. Results: The segmentation model achieved high accuracy across all timepoints (mean Dice > 0.89). Univariate and multivariate analyses identified base diameter, ellipsoid zone integrity, and macular hole area as significant BCVA predictors (P < 0.05). Incorporating dynamic recovery rates consistently improved logistic regression AUC, especially at the 3-month follow-up. The multimodal DL model outperformed logistic regression, yielding higher AUCs and overall accuracy at each stage. The difference is as high as 0.12, demonstrating the complementary value of raw image volume and dynamic parameters. Conclusions: Integrating dynamic parameters into the multimodal DL model significantly enhances the accuracy of predictions. This fully automated process therefore represents a promising clinical decision support tool for personalized postoperative management in macular hole surgery.

eess.IV

CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging

Recent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modality- and task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (\CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique "modality signature" for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model.

cs.CV

UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation

Significant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework \UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus image-based disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks.

cs.CV

Bifurcation formula for transition paths in stochastic dynamical systems by spectral flow

This paper investigates bifurcation phenomena and stability of most probable transition paths (MPTPs) in stochastic dynamical systems through a combined variational and spectral flow approach. Within the Onsager-Machlup framework, MPTPs are characterized as minimizers of an energy-dependent Lagrangian functional incorporating noise intensity. Existence criteria for such minimizers are established through critical value analysis and variational techniques. The main theoretical advancement is a spectral flow formula that detects bifurcation points and quantifies stability changes under noise perturbations. Specifically, the analysis reveals: (i) noise-sensitive MPTPs where variations in noise intensity destroy the minimizer property, and (ii) noise-robust MPTPs where stability is maintained despite finite noise fluctuations. These results establish a correspondence between Lagrangian bifurcations and stochastic phase transitions, providing a mathematical foundation for predicting noise-driven transition mechanisms in stochastic systems.

math.DS

Semi-stable and splitting models for unitary Shimura varieties over ramified places. I

We consider Shimura varieties associated to a unitary group of signature $(n-s,s)$ where $n$ is even. For these varieties, we construct smooth $p$-adic integral models for $s=1$ and regular $p$-adic integral models for $s=2$ and $s=3$ over odd primes $p$ which ramify in the imaginary quadratic field with level subgroup at $p$ given by the stabilizer of a $π$-modular lattice in the hermitian space. Our construction, which has an explicit moduli-theoretic description, is given by an explicit resolution of a corresponding local model.

math.NT

Semi-stable and splitting models for unitary Shimura varieties over ramified places. II

We consider Shimura varieties associated to a unitary group of signature $(n-1, 1)$. For these varieties, we construct $p$-adic integral models over odd primes $p$ which ramify in the imaginary quadratic field with level subgroup at $p$ given by the stabilizer of a vertex lattice in the hermitian space. Our models are given by a variation of the construction of the splitting models of Pappas-Rapoport and they have a simple moduli theoretic description. By an explicit calculation, we show that these splitting models are normal, flat, Cohen-Macaulay and with reduced special fiber. In fact, they have relatively simple singularities: we show that a single blow-up along a smooth codimension one subvariety of the special fiber produces a semi-stable model. This also implies the existence of semi-stable models of the corresponding Shimura varieties.

math.NT

WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving

Language models uncover unprecedented abilities in analyzing driving scenarios, owing to their limitless knowledge accumulated from text-based pre-training. Naturally, they should particularly excel in analyzing rule-based interactions, such as those triggered by traffic laws, which are well documented in texts. However, such interaction analysis remains underexplored due to the lack of dedicated language datasets that address it. Therefore, we propose Waymo Open Motion Dataset-Reasoning (WOMD-Reasoning), a comprehensive large-scale Q&As dataset built on WOMD focusing on describing and reasoning traffic rule-induced interactions in driving scenarios. WOMD-Reasoning also presents by far the largest multi-modal Q&A dataset, with 3 million Q&As on real-world driving scenarios, covering a wide range of driving topics from map descriptions and motion status descriptions to narratives and analyses of agents' interactions, behaviors, and intentions. To showcase the applications of WOMD-Reasoning, we design Motion-LLaVA, a motion-language model fine-tuned on WOMD-Reasoning. Quantitative and qualitative evaluations are performed on WOMD-Reasoning dataset as well as the outputs of Motion-LLaVA, supporting the data quality and wide applications of WOMD-Reasoning, in interaction predictions, traffic rule compliance plannings, etc. The dataset and its vision modal extension are available on https://waymo.com/open/download/. The codes & prompts to build it are available on https://github.com/yhli123/WOMD-Reasoning.

cs.RO

First-principles Investigation of Exceptional Coarsening-resistant V-Sc(Al2Cu)4 Nanoprecipitates in Al-Cu-Mg-Ag-Sc Alloys

Aluminum-copper-magnesium-sliver (Al-Cu-Mg-Ag) alloys are extensively utilized in aerospace industries due to the formation of Omega nano-plates.However, the rapid coarsening of these nano-plates above 475 K restricts their application at elevated temperatures.When introducing scandium (Sc) to these alloys, the service temperature of the resultant alloys can reach an unprecedented 675 K, attributed to the in situ formation of a coarsening-resistant V-Sc(Al2Cu)4 phase within the Omega nano-plates. However, the fundamental thermodynamic properties and mechanisms behind the remarkable coarsening resistance of V nano-plates remain unexplored.Here, we employ first-principles calculations to investigate the phase stability of V-Sc(Al2Cu)4 phase, the basic kinetic features of V phase formation within Omega nano-plates, and the origins of the extremely high thermal stability of V nano-plates. Our results indicate that V-Sc(Al2Cu)4 is meta-stable and thermodynamically tends to evolve into a stable ScAl7Cu5 phase. We also demonstrate that kinetic factors are mainly responsible for the temperature dependence of V phase formation. Notably, the formation of V-Sc(Al2Cu)4 within Omega nano-plates modifies the Kagome lattice in the shell layer of the Omega nano-plates, inhibiting further thickening of V nano-plates through the thickening pathway of Omega nano-plates. This interface transition leads to the exceptional coarsening resistance of the V nano-plates. Moreover, we also screened 14 promising element substitutions for Sc. These findings are anticipated to accelerate the development of high-performance Al alloys with superior heat resistance.

cond-mat.mtrl-sci

AgentAlign: Misalignment-Adapted Multi-Agent Perception for Resilient Inter-Agent Sensor Correlations

Cooperative perception has attracted wide attention given its capability to leverage shared information across connected automated vehicles (CAVs) and smart infrastructures to address sensing occlusion and range limitation issues. However, existing research overlooks the fragile multi-sensor correlations in multi-agent settings, as the heterogeneous agent sensor measurements are highly susceptible to environmental factors, leading to weakened inter-agent sensor interactions. The varying operational conditions and other real-world factors inevitably introduce multifactorial noise and consequentially lead to multi-sensor misalignment, making the deployment of multi-agent multi-modality perception particularly challenging in the real world. In this paper, we propose AgentAlign, a real-world heterogeneous agent cross-modality feature alignment framework, to effectively address these multi-modality misalignment issues. Our method introduces a cross-modality feature alignment space (CFAS) and heterogeneous agent feature alignment (HAFA) mechanism to harmonize multi-modality features across various agents dynamically. Additionally, we present a novel V2XSet-noise dataset that simulates realistic sensor imperfections under diverse environmental conditions, facilitating a systematic evaluation of our approach's robustness. Extensive experiments on the V2X-Real and V2XSet-Noise benchmarks demonstrate that our framework achieves state-of-the-art performance, underscoring its potential for real-world applications in cooperative autonomous driving. The controllable V2XSet-Noise dataset and generation pipeline will be released in the future.

cs.CV

Extrapolating Prospective Glaucoma Fundus Images through Diffusion Model in Irregular Longitudinal Sequences

The utilization of longitudinal datasets for glaucoma progression prediction offers a compelling approach to support early therapeutic interventions. Predominant methodologies in this domain have primarily focused on the direct prediction of glaucoma stage labels from longitudinal datasets. However, such methods may not adequately encapsulate the nuanced developmental trajectory of the disease. To enhance the diagnostic acumen of medical practitioners, we propose a novel diffusion-based model to predict prospective images by extrapolating from existing longitudinal fundus images of patients. The methodology delineated in this study distinctively leverages sequences of images as inputs. Subsequently, a time-aligned mask is employed to select a specific year for image generation. During the training phase, the time-aligned mask resolves the issue of irregular temporal intervals in longitudinal image sequence sampling. Additionally, we utilize a strategy of randomly masking a frame in the sequence to establish the ground truth. This methodology aids the network in continuously acquiring knowledge regarding the internal relationships among the sequences throughout the learning phase. Moreover, the introduction of textual labels is instrumental in categorizing images generated within the sequence. The empirical findings from the conducted experiments indicate that our proposed model not only effectively generates longitudinal data but also significantly improves the precision of downstream classification tasks.

cs.CV

Atomic-scale Nucleation and Growth Pathway of Complex Plate-like Precipitates in Aluminum Alloys

Aluminum alloys, the most widely utilized lightweight structural materials, predominantly depend on coherent complex-structured nano-plates to enhance their mechanical properties. Despite several decades of research, the atomic-scale nucleation and growth pathways for these complex-structured nano-plates remain elusive, as probing and simulating atomic events like solid nucleation is prohibitively challenging. Here, using theoretical calculations and focus on three representative complex-structured nano-plates in commercial Al alloys, we explicitly demonstrate their associated structural transitions follow an inter-layer-sliding+shuffling mode. Specifically, partial dislocations complete the inter-layer-sliding stage, while atomic shuffling occurs upon forming the unstable basic structural transformation unit of the nano-plates. By identifying these basic structural transformation units, we propose structural evolution pathways for these nano-plates within the Al matrix, which align well with experimental observations and enable the evaluation of critical nuclei. These findings provide long-sought mechanistic details into how coherent nano-plates nucleate and grow, facilitating the rational design of higher-performance Al alloys and other structural materials.

cond-mat.mtrl-sci