Searcharxiv⌕ Search

arXiv subjects

Zhiyue Zhang

Publications and source records attributed to Zhiyue Zhang.

12 recordsLinked to original sources

FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry

We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data design. Stage 1 learns a high-fidelity, affine-invariant geometry model by initializing with DINOv3 and training on a curated 10.2M-sample multi-domain corpus with complementary local-detail supervision, yielding sharp boundaries and strong cross-domain generalization. Stage 2 moves beyond global scaling by introducing lightweight pixel-wise calibration fields for metric estimation: a scale field for spatially varying metric alignment and a ray-direction correction field that mitigates directional bias in point-map geometry, together producing metrically consistent 3D point maps. Beyond model design, we identify camera intrinsic coverage, especially focal length distribution mismatch between training and test data, as a key bottleneck for zero-shot metric generalization: performance drops sharply when test intrinsics fall outside the training distribution. To address this, we synthesize additional training data across diverse focal lengths using a Blender-based data engine, repairing under-covered focal regimes and improving robustness under intrinsic shift. Extensive zero-shot evaluations across seven benchmarks show that FoundationGeo significantly strengthens cross-domain robustness, staying near the top across diverse domains while avoiding the sharp cross-domain performance drops observed in other methods. This consistency translates into the best overall performance, surpassing heavier baselines by over 5.2% on average.

cs.CV↗

A Perturbation-Correction Method Based on Local Randomized Neural Networks for Quasi-Linear Interface Problems

For quasi-linear elliptic interface problems with discontinuous diffusion coefficients, randomized neural network approximations may exhibit stagnation because the associated objective functional is generally nonconvex. This paper proposes a Local Randomized Neural Network (LRaNN) perturbation-correction method, denoted by LRaNN-PC, to alleviate this stagnation. The method represents the solution using an LRaNN on each subdomain, coupled through the interface conditions in a domain-decomposed framework. It consists of a primary stage and a perturbation-correction (PC) stage. The primary stage computes the primary approximation by minimizing the original nonconvex objective functional. The PC stage constructs a residual-driven correction by performing a local expansion of the residual around the primary approximation and representing the correction in an independently generated randomized trial space. The correction coefficients are obtained by solving a least-squares residual-correction subproblem in this trial space. For the solution-dependent quasi-linear elliptic model under the stated sufficient assumptions, we derive a residual-controlled upper bound for the broken $H^1$ seminorm error. The bound involves the discrete residual, the quadrature error, the residual of the perturbation-correction subproblem, and the truncation remainder of the perturbation expansion. Numerical experiments further test irregular interfaces and high-contrast coefficients within this setting, and also examine gradient-dependent diffusivities and moving-interface extensions. In the tested benchmarks, LRaNN-PC reduces the relative $L^2$ error by up to $4$--$7$ orders of magnitude compared with the primary LRaNN stage.

math.NA↗

PCELM: Perturbation-Correction Extreme Learning Machine for the Stefan problem

For Stefan problems, characterized by moving boundaries and discontinuous coefficients due to phase changes, the inherent nonconvexity of the objective functional frequently causes optimization difficulty in randomized neural network approximations; to address this, we propose a Perturbation-Correction Extreme Learning Machine (PCELM) framework, built upon the extreme learning machine framework. This method first establishes a basic approximation during an initialization step by minimizing the original nonconvex residual, typically achieving only moderate accuracy, and then, in a subsequent correction step, determines a correction term by solving a subproblem based on a perturbation expansion around this basic approximation, thereby transforming it into a convex optimization problem for the output coefficients that ensures rapid convergence. We further provide a rigorous a convexity analysis, demonstrating that PCELM method solves a convex sub-problem. Numerical experiments on various Stefan problems, including multi-phase and multi-dimensional Stefan problems, confirm that the proposed PCELM method successfully overcomes optimization plateaus, with the correction step consistently delivering a significant improvement of 2-6 orders of magnitude in the relative L2 accuracy.

math.NA↗

Eligibility-Aware Evidence Synthesis: An Agentic Framework for Clinical Trial Meta-Analysis

Clinical evidence synthesis requires identifying relevant trials from large registries and aggregating results that account for population differences. While recent LLM-based approaches have automated components of systematic review, they do not support end-to-end evidence synthesis. Moreover, conventional meta-analysis weights studies by statistical precision without considering clinical compatibility reflected in eligibility criteria. We propose EligMeta, an agentic framework that integrates automated trial discovery with eligibility-aware meta-analysis, translating natural-language queries into reproducible trial selection and incorporating eligibility alignment into study weighting to produce cohort-specific pooled estimates. EligMeta employs a hybrid architecture separating LLM-based reasoning from deterministic execution: LLMs generate interpretable rules from natural-language queries and perform schema-constrained parsing of trial metadata, while all logical operations, weight computations, and statistical pooling are executed deterministically to ensure reproducibility. The framework structures eligibility criteria and computes similarity-based study weights reflecting population alignment between target and comparator trials. In a gastric cancer landscape analysis, EligMeta reduced 4,044 candidate trials to 39 clinically relevant studies through rule-based filtering, recovering all 13 guideline-cited trials. In an olaparib adverse events meta-analysis across four trials, eligibility-aware weighting shifted the pooled risk ratio from 2.18 (95% CI: 1.71-2.79) under conventional Mantel-Haenszel estimation to 1.97 (95% CI: 1.76-2.20), demonstrating quantifiable impact of incorporating eligibility alignment. EligMeta bridges automated trial discovery with eligibility-aware meta-analysis, providing a scalable and reproducible framework for evidence synthesis in precision medicine.

stat.ME↗

A Predictor Corrector Convex Splitting Method for Stefan Problems Based on Extreme Learning Machines

Solving Stefan problems via neural networks is inherently challenged by the nonlinear coupling between the solutions and the free boundary, which results in a non-convex optimization problem. To address this, this work proposes an Operator Splitting Method (OSM) based on Extreme Learning Machines (ELM) to decouple the geometric interface evolution from the physical field reconstruction. Within a predictor-corrector framework, the method splits the coupled system into an alternating sequence of two linear and convex subproblems: solving the diffusion equation on fixed subdomains and updating the interface geometry based on the Stefan condition. A key contribution is the formulation of both steps as linear least-squares problems; this transforms the computational strategy from a non-convex gradient-based optimization into a stable fixed-point iteration composed of alternating convex solvers. From a theoretical perspective, the relaxed iterative operator is shown to be locally contractive, and its fixed points are consistent with stationary points of the coupled residual functional. Benchmarks across 1D to 3D domains demonstrate the stability and high accuracy of the method, confirming that the proposed framework provides a highly accurate and efficient numerical solution for free boundary problems.

math.NA↗

LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing

Retrieval Augmented Generation (RAG) systems often struggle with domain-specific knowledge due to performance deterioration of pre-trained embeddings and prohibitive computational costs of large language model (LLM)-based retrievers. While fine-tuning data augmentation embedding models offers a promising direction, its effectiveness is limited by the need for high-quality training data and reliable chunking strategies that preserve contextual integrity. We propose LMAR (Language Model Augmented Retriever), a model-agnostic framework that addresses these challenges by combining LLM-guided data synthesis with contrastive embedding adaptation and efficient text clustering. LMAR consists of a two-stage pipeline: (1) Triplet sampling and synthetic data augmentation, where LLMs act as both labeler and validator to ensure high-fidelity supervision throughout the pipeline. Experimental results across multiple domain-specific benchmark datasets demonstrate that LMAR outperforms multiple baseline models, while maintaining moderate hardware requirements and low latency. Its model-agnostic nature further enables seamless integration with emerging RAG architectures and text embedding models, ensuring continual improvements without redesigning the pipeline. These results highlight LMAR as a practical and cost-effective solution for scalable domain-specific adaptation.

cs.IR↗

TransformerLSR: Attentive Joint Model of Longitudinal Data, Survival, and Recurrent Events with Concurrent Latent Structure

In applications such as biomedical studies, epidemiology, and social sciences, recurrent events often co-occur with longitudinal measurements and a terminal event, such as death. Therefore, jointly modeling longitudinal measurements, recurrent events, and survival data while accounting for their dependencies is critical. While joint models for the three components exist in statistical literature, many of these approaches are limited by heavy parametric assumptions and scalability issues. Recently, incorporating deep learning techniques into joint modeling has shown promising results. However, current methods only address joint modeling of longitudinal measurements at regularly-spaced observation times and survival events, neglecting recurrent events. In this paper, we develop TransformerLSR, a flexible transformer-based deep modeling and inference framework to jointly model all three components simultaneously. TransformerLSR integrates deep temporal point processes into the joint modeling framework, treating recurrent and terminal events as two competing processes dependent on past longitudinal measurements and recurrent event times. Additionally, TransformerLSR introduces a novel trajectory representation and model architecture to potentially incorporate a priori knowledge of known latent structures among concurrent longitudinal variables. We demonstrate the effectiveness and necessity of TransformerLSR through simulation studies and analyzing a real-world medical dataset on patients after kidney transplantation.

stat.ML↗

MoMA: Model-based Mirror Ascent for Offline Reinforcement Learning

Model-based offline reinforcement learning methods (RL) have achieved state-of-the-art performance in many decision-making problems thanks to their sample efficiency and generalizability. Despite these advancements, existing model-based offline RL approaches either focus on theoretical studies without developing practical algorithms or rely on a restricted parametric policy space, thus not fully leveraging the advantages of an unrestricted policy space inherent to model-based methods. To address this limitation, we develop MoMA, a model-based mirror ascent algorithm with general function approximations under partial coverage of offline data. MoMA distinguishes itself from existing literature by employing an unrestricted policy class. In each iteration, MoMA conservatively estimates the value function by a minimization procedure within a confidence set of transition models in the policy evaluation step, then updates the policy with general function approximations instead of commonly-used parametric policy classes in the policy improvement step. Under some mild assumptions, we establish theoretical guarantees of MoMA by proving an upper bound on the suboptimality of the returned policy. We also provide a practically implementable, approximate version of the algorithm. The effectiveness of MoMA is demonstrated via numerical studies.

cs.LG↗

A Neumann interface optimal control problem with elliptic PDE constraints and its discretization and numerical analysis

We study an optimal control problem governed by elliptic PDEs with interface, which the control acts on the interface. Due to the jump of the coefficient across the interface and the control acting on the interface, the regularity of solution of the control problem is limited on the whole domain, but smoother on subdomains. The control function with pointwise inequality constraints is served as the flux jump condition which we called Neumann interface control. We use a simple uniform mesh that is independent of the interface. The standard linear finite element method can not achieve optimal convergence when the uniform mesh is used. Therefore the state and adjoint state equations are discretized by piecewise linear immersed finite element method (IFEM). While the accuracy of the piecewise constant approximation of the optimal control on the interface is improved by a postprocessing step which possesses superconvergence properties; as well as the variational discretization concept for the optimal control is used to improve the error estimates. Optimal error estimates for the control, suboptimal error estimates for state and adjoint state are derived. Numerical examples with and without constraints are provided to illustrate the effectiveness of the proposed scheme and correctness of the theoretical analysis.

math.NA↗

Decoupling Numerical Method Based on Deep Neural Network for Nonlinear Degenerate Interface Problems

Interface problems depict many fundamental physical phenomena and widely apply in the engineering. However, it is challenging to develop efficient fully decoupled numerical methods for solving degenerate interface problems in which the coefficient of a PDE is discontinuous and greater than or equal to zero on the interface. The main motivation in this paper is to construct fully decoupled numerical methods for solving nonlinear degenerate interface problems with ``double singularities". An efficient fully decoupled numerical method is proposed for nonlinear degenerate interface problems. The scheme combines deep neural network on the singular subdomain with finite difference method on the regular subdomain. The key of the new approach is to split nonlinear degenerate partial differential equation with interface into two independent boundary value problems based on deep learning. The outstanding advantages of the proposed schemes are that not only the convergence order of the degenerate interface problems on whole domain is determined by the finite difference scheme on the regular subdomain, but also can calculate $\mathbf{VERY}$ $\mathbf{BIG}$ jump ratio(such as $10^{12}:1$ or $1:10^{12}$) for the interface problems including degenerate and non-degenerate cases. The expansion of the solutions does not contains any undetermined parameters in the numerical method. In this way, two independent nonlinear systems are constructed in other subdomains and can be computed in parallel. The flexibility, accuracy and efficiency of the methods are validated from various experiments in both 1D and 2D. Specially, not only our method is suitable for solving degenerate interface case, but also for non-degenerate interface case. Some application examples with complicated multi-connected and sharp edge interface examples including degenerate and nondegenerate cases are also presented.

math.NA↗

Viable embedded wormholes and energy conditions in $f(\mathcal{R},\mathcal{G})$ gravity

The current study explores the generalized embedded wormhole solutions in the background of $f(\mathcal{R},\mathcal{G})$ gravity, where $\mathcal{R}$ represents the Ricci scalar and $\mathcal{G}$ denotes the Gauss-Bonnet invariant. To investigate the necessary structures of the wormhole solutions we thoroughly analyzed the energy conditions under $f(\mathcal{R},\mathcal{G})$ gravity within the anisotropic source of matter. To meet this aim, we consider spherically symmetric geometry with the most generic gravity model of the gravity. A modified version of the field equations is calculated for two different embedded wormhole solutions. All the energy conditions are calculated and shown graphically with the regional ranges of the model parameter. Further, the invalid region of the energy conditions confirms the presence of exotic matter. Finally, we have concluding remarks.

gr-qc↗

Simple and Scalable Sparse k-means Clustering via Feature Ranking

Clustering, a fundamental activity in unsupervised learning, is notoriously difficult when the feature space is high-dimensional. Fortunately, in many realistic scenarios, only a handful of features are relevant in distinguishing clusters. This has motivated the development of sparse clustering techniques that typically rely on k-means within outer algorithms of high computational complexity. Current techniques also require careful tuning of shrinkage parameters, further limiting their scalability. In this paper, we propose a novel framework for sparse k-means clustering that is intuitive, simple to implement, and competitive with state-of-the-art algorithms. We show that our algorithm enjoys consistency and convergence guarantees. Our core method readily generalizes to several task-specific algorithms such as clustering on subsets of attributes and in partially observed data settings. We showcase these contributions thoroughly via simulated experiments and real data benchmarks, including a case study on protein expression in trisomic mice.

stat.ML↗