SearcharxivSearch

arXiv subjects

Rong Zhou

Publications and source records attributed to Rong Zhou.

At least 19 recordsLinked to original sources

Positive $\tau$-bi-Ricci curvature and Mean curvature flow with surgery in hyperbolic space

Let $n\geq3$ and $0\leq\tau\leq2$. We prove that every smooth closed connected immersed hypersurface in hyperbolic space whose induced metric has positive $\tau$-bi-Ricci curvature admits a mean curvature flow with surgery which has only finitely many surgery times and terminates. In the range compatible with cylindrical necks, the key ingredients are a preserved quantitative spectral pinching condition, which converts the intrinsic hypothesis into uniform two-convexity, and cylindrical and derivative estimates that remain valid across the hyperbolic standard neck replacement. At the endpoint $n=3$, $\tau=2$, positive $\tau$-bi-Ricci curvature is positive Ricci curvature and forces strict convexity, so the ordinary mean curvature flow converges to a round point. Consequently, the underlying manifold is diffeomorphic to a sphere or to a finite connected sum of copies of $\mathbb{S}^{n-1}\times\mathbb{S}^1$. If the initial hypersurface is embedded and bounds a compact domain, that domain is a one-handlebody, namely a ball with finitely many one-handles attached.

math.DG

Weinstock inequalities for outward-minimizing domains

We prove sharp Weinstock inequalities for the first nonzero Steklov eigenvalue of smooth outward-minimizing domains in Euclidean space and hyperbolic space. The method is based on the weak inverse mean curvature flow of Huisken--Ilmanen. The main new ingredient is an endpoint distributional monotonicity argument obtained from the calibrated weak formulation and the Gauss--Green formula for divergence-measure fields.

math.DG

Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges

Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auction outcomes: DSPs throttle participation under compute and budget constraints, reducing the effective use of limited bidding capacity. We present a competition-aware request dispatch framework that uses distributional bid prediction and probabilistic forwarding to decide whether each request should be sent to each DSP. The system adapts per-DSP thresholds over time through lightweight policy optimization to track non-stationary market conditions. We evaluate the framework through four sequential online experiments on a production platform serving over 20 billion daily requests. A full multi-DSP deployment reduces DSP request volume under the policy by 34.2% while increasing net revenue by 4.6% (p<0.001) in a recent 14-day window after an initial DSP adaptation period. Further analysis highlights strong heterogeneity across traffic segments and reveals that aggregate metrics can be misleading. Segment-level and per-DSP analyses suggest that the policy surfaces comparative advantages among DSPs, improving monetized outcomes without increasing overall request volume.

cs.AI

Free boundary flows by powers of the Gauss curvature in the unit ball

We study smooth compact strictly convex hypersurfaces in the unit ball that meet the support sphere orthogonally and evolve by the $\alpha$-Gauss curvature flow $\partial_tX=-K^\alpha\nu$, $\alpha>0$. We prove that the solution remains strictly convex, becomes extinct in finite time and contracts to a single point on the support sphere. If $\alpha >\frac{1}{n+2}$, we apply a Cayley-type conformal map that sends the extinction point to the origin of a Euclidean half-space and then normalize the enclosed half-space volume. The resulting normalized hypersurfaces converge smoothly to the unit hemisphere. The proof combines boundary identities for the spherical free boundary, a boundary-adapted Tso estimate, an almost-monotonicity formula for a half-space entropy, and uniform curvature estimates for the normalized flow.

math.DG

Higher regularity of the inverse anisotropic mean curvature flow

We prove an anisotropic analogue of the higher regularity theorem of Huisken and Ilmanen for inverse mean curvature flow. For an arbitrary smooth Minkowski norm, we first prove a Huisken--Ilmanen type Harnack estimate for smooth closed strictly star-shaped solutions. We then construct global smooth solutions starting from $C^1$ strictly star-shaped hypersurfaces with bounded nonnegative weak anisotropic mean curvature. Combining this construction with the asymptotic theory for weak inverse anisotropic mean curvature flow, we show that weak solutions starting from bounded smooth initial sets become smooth outside a compact set.

math.DG

Contraction of hypersurfaces with positive sectional curvature in hyperbolic space

We study contracting curvature flows of compact hypersurfaces with positive sectional curvature in hyperbolic space $\mathbb{H}^{n+1}$. The speed is assumed to be homogeneous of degree one in the principal curvatures and to satisfy certain conditions. This class of flows includes the $k$th mean curvature flow as a special case. We show that if the initial hypersurface has positive sectional curvature, then this property is preserved along the flow, and the evolving hypersurface contracts to a round point in finite time.

math.DG

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.

cs.AI

Adaptive Clinical-Aware Latent Diffusion for Multimodal Brain Image Generation and Missing Modality Imputation

Multimodal neuroimaging provides complementary insights for Alzheimer's disease diagnosis, yet clinical datasets frequently suffer from missing modalities. We propose ACADiff, a framework that synthesizes missing brain imaging modalities through adaptive clinical-aware diffusion. ACADiff learns mappings between incomplete multimodal observations and target modalities by progressively denoising latent representations while attending to available imaging data and clinical metadata. The framework employs adaptive fusion that dynamically reconfigures based on input availability, coupled with semantic clinical guidance via GPT-4o-encoded prompts. Three specialized generators enable bidirectional synthesis among sMRI, FDG-PET, and AV45-PET. Evaluated on ADNI subjects, ACADiff achieves superior generation quality and maintains robust diagnostic performance even under extreme 80\% missing scenarios, outperforming all existing baselines. To promote reproducibility, code is available at https://github.com/rongzhou7/ACADiff

cs.CV

Diffusion-Guided Pretraining for Brain Graph Foundation Models

With the growing interest in foundation models for brain signals, graph-based pretraining has emerged as a promising paradigm for learning transferable representations from connectome data. However, existing contrastive and masked autoencoder methods typically rely on naive random dropping or masking for augmentation, which is ill-suited for brain graphs and hypergraphs as it disrupts semantically meaningful connectivity patterns. Moreover, commonly used graph-level readout and reconstruction schemes fail to capture global structural information, limiting the robustness of learned representations. In this work, we propose a unified diffusion-based pretraining framework that addresses both limitations. First, diffusion is designed to guide structure-aware dropping and masking strategies, preserving brain graph semantics while maintaining effective pretraining diversity. Second, diffusion enables topology-aware graph-level readout and node-level global reconstruction by allowing graph embeddings and masked nodes to aggregate information from globally related regions. Extensive experiments across multiple neuroimaging datasets with over 25,000 subjects and 60,000 scans involving various mental disorders and brain atlases demonstrate consistent performance improvements.

cs.LG

Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models

Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.

cs.AI

Asymptotic behaviour of the weak inverse anisotropic mean curvature flow

We first establish a local gradient estimate for anisotropic $p$-harmonic functions. A key feature of our estimate is that the constant remains bounded as $p\to 1$; consequently, in the limit $p\to 1$, this estimate yields the local gradient estimate for weak solutions of the inverse anisotropic mean curvature flow (IAMCF). As an application, we show that the weak IAMCF is asymptotic to the expanding Wulff shape solution at the infinity, thereby extending the result of Huisken and Ilmanen in [8] to the anisotropic case.

math.DG

Spatio-Temporal Graph Deep Learning with Stochastic Differential Equations for Uncovering Alzheimer's Disease Progression

Identifying objective neuroimaging biomarkers to forecast Alzheimer's disease (AD) progression is crucial for timely intervention. However, this task remains challenging due to the complex dysfunctions in the spatio-temporal characteristics of underlying brain networks, which are often overlooked by existing methods. To address these limitations, we develop an interpretable spatio-temporal graph neural network framework to predict future AD progression, leveraging dual Stochastic Differential Equations (SDEs) to model the irregularly-sampled longitudinal functional magnetic resonance imaging (fMRI) data. We validate our approach on two independent cohorts, including the Open Access Series of Imaging Studies (OASIS-3) and the Alzheimer's Disease Neuroimaging Initiative (ADNI). Our framework effectively learns sparse regional and connective importance probabilities, enabling the identification of key brain circuit abnormalities associated with disease progression. Notably, we detect the parahippocampal cortex, prefrontal cortex, and parietal lobule as salient regions, with significant disruptions in the ventral attention, dorsal attention, and default mode networks. These abnormalities correlate strongly with longitudinal AD-related clinical symptoms. Moreover, our interpretability strategy reveals both established and novel neural systems-level and sex-specific biomarkers, offering new insights into the neurobiological mechanisms underlying AD progression. Our findings highlight the potential of spatio-temporal graph-based learning for early, individualized prediction of AD progression, even in the context of irregularly-sampled longitudinal imaging data.

cs.LG

SAMed-2: Selective Memory Enhanced Medical Segment Anything Model

Recent "segment anything" efforts show promise by learning from large-scale data, but adapting such models directly to medical images remains challenging due to the complexity of medical data, noisy annotations, and continual learning requirements across diverse modalities and anatomical structures. In this work, we propose SAMed-2, a new foundation model for medical image segmentation built upon the SAM-2 architecture. Specifically, we introduce a temporal adapter into the image encoder to capture image correlations and a confidence-driven memory mechanism to store high-certainty features for later retrieval. This memory-based strategy counters the pervasive noise in large-scale medical datasets and mitigates catastrophic forgetting when encountering new tasks or modalities. To train and evaluate SAMed-2, we curate MedBank-100k, a comprehensive dataset spanning seven imaging modalities and 21 medical segmentation tasks. Our experiments on both internal benchmarks and 10 external datasets demonstrate superior performance over state-of-the-art baselines in multi-task scenarios. The code is available at: https://github.com/ZhilingYan/Medical-SAM-Bench.

cs.CV

EfficientLLM: Efficiency in Large Language Models

Large Language Models (LLMs) have driven significant progress, yet their growing parameter counts and context windows incur prohibitive compute, energy, and monetary costs. We introduce EfficientLLM, a novel benchmark and the first comprehensive empirical study evaluating efficiency techniques for LLMs at scale. Conducted on a production-class cluster (48xGH200, 8xH200 GPUs), our study systematically explores three key axes: (1) architecture pretraining (efficient attention variants: MQA, GQA, MLA, NSA; sparse Mixture-of-Experts (MoE)), (2) fine-tuning (parameter-efficient methods: LoRA, RSLoRA, DoRA), and (3) inference (quantization methods: int4, float16). We define six fine-grained metrics (Memory Utilization, Compute Utilization, Latency, Throughput, Energy Consumption, Compression Rate) to capture hardware saturation, latency-throughput balance, and carbon cost. Evaluating over 100 model-technique pairs (0.5B-72B parameters), we derive three core insights: (i) Efficiency involves quantifiable trade-offs: no single method is universally optimal; e.g., MoE reduces FLOPs and improves accuracy but increases VRAM by 40%, while int4 quantization cuts memory/energy by up to 3.9x at a 3-5% accuracy drop. (ii) Optima are task- and scale-dependent: MQA offers optimal memory-latency trade-offs for constrained devices, MLA achieves lowest perplexity for quality-critical tasks, and RSLoRA surpasses LoRA efficiency only beyond 14B parameters. (iii) Techniques generalize across modalities: we extend evaluations to Large Vision Models (Stable Diffusion 3.5, Wan 2.1) and Vision-Language Models (Qwen2.5-VL), confirming effective transferability. By open-sourcing datasets, evaluation pipelines, and leaderboards, EfficientLLM provides essential guidance for researchers and engineers navigating the efficiency-performance landscape of next-generation foundation models.

cs.CL

ELITE: Embedding-Less retrieval with Iterative Text Exploration

Large Language Models (LLMs) have achieved impressive progress in natural language processing, but their limited ability to retain long-term context constrains performance on document-level or multi-turn tasks. Retrieval-Augmented Generation (RAG) mitigates this by retrieving relevant information from an external corpus. However, existing RAG systems often rely on embedding-based retrieval trained on corpus-level semantic similarity, which can lead to retrieving content that is semantically similar in form but misaligned with the question's true intent. Furthermore, recent RAG variants construct graph- or hierarchy-based structures to improve retrieval accuracy, resulting in significant computation and storage overhead. In this paper, we propose an embedding-free retrieval framework. Our method leverages the logical inferencing ability of LLMs in retrieval using iterative search space refinement guided by our novel importance measure and extend our retrieval results with logically related information without explicit graph construction. Experiments on long-context QA benchmarks, including NovelQA and Marathon, show that our approach outperforms strong baselines while reducing storage and runtime by over an order of magnitude.

cs.CL

On the basic locus of GSpin Shimura varieties with vertex stabilizer level

We study the basic locus of Shimura varieties associated to the group of spinor similitudes of a quadratic space over $\mathbb{Q}$ with level structure given by the stabilizer of a vertex lattice. We give a description of the underlying reduced scheme of the associated Rapoport--Zink space, generalizing results of Howard--Pappas [HP17] and Oki [Oki20b], in the case of self-dual, and almost self dual level structure.

math.NT

Strongly compatible systems associated to semistable abelian varieties

We prove a motivic refinement of a result of Weil, Deligne and Raynaud on the existence of strongly compatible systems associated to abelian varieties. More precisely, given an abelian variety $A$ over a number field $\mathrm{E}\subset \mathbb C$, we prove that after replacing $\mathbb E$ by a finite extension, the action of $\mathrm{Gal}(\overline{\mathrm E}/\mathrm E)$ on the $\ell$-adic cohomology $\mathrm H^1_{\mathrm{\acute{e}t}}(A_{\overline{\mathrm E}},\mathbb Q_\ell)$ gives rise to a strongly compatible system of $\ell$-adic representations valued in the Mumford--Tate group $\mathbf G$ of $A$. This involves an independence of $\ell$-statement for the Weil--Deligne representation associated to $A$ at places of semistable reduction, extending previous work of ours at places of good reduction.

math.NT

Two-stage deep learning framework for the restoration of incomplete-ring PET images

Positron Emission Tomography (PET) is an important molecular imaging tool widely used in medicine. Traditional PET systems rely on complete detector rings for full angular coverage and reliable data collection. However, incomplete-ring PET scanners have emerged due to hardware failures, cost constraints, or specific clinical needs. Standard reconstruction algorithms often suffer from performance degradation with these systems because of reduced data completeness and geometric inconsistencies. We present a two-stage deep-learning framework that, without incorporating any time-of-flight (TOF) information, restores high-quality images from data with about 50% missing coincidences - double the loss levels previously addressed by CNN-based methods. The pipeline operates in two stages: a projection-domain Attention U-Net first predicts the missing sections of the sinogram by leveraging spatial context from neighbouring slices, after which the completed data are reconstructed with OSEM algorithm and passed to a cascaded U-Net & warm-start diffusion model for image refinement. This module starts the reverse diffusion process from the U-Net coarse prediction rather than pure Gaussian noise. Using 613 simulated brain volumes from real scans (196 healthy brain samples, 217 Alzheimer's disease samples, and 200 Mild Cognitive Impairment samples), the result shows that our model successfully preserves most anatomical structures and tracer distribution features with PSNR of 38.18 to 38.59 dB and SSIM of 0.9904 to 0.9925. Our two-stage deep-learning framework effectively restores high-quality PET images from over 50% incomplete-ring data, achieving near-complete anatomical fidelity and robust performance without requiring TOF information.

cs.CV