SearcharxivSearch

arXiv subjects

Peng Luo

Publications and source records attributed to Peng Luo.

At least 19 recordsLinked to original sources

Hessian degeneracy of the torsion function on smooth simply connected nonconvex planar domains

We construct a family of bounded, smooth, simply connected, nonconvex planar domains $\Omega_{a,\epsilon}$. Let $u_{a,\epsilon}$ be the corresponding torsion function satisfying \[ -\Delta u_{a,\epsilon}=1\quad\text{in }\Omega_{a,\epsilon},\qquad u_{a,\epsilon}=0\quad\text{on }\partial\Omega_{a,\epsilon}. \] There exists \(a^*\in(0,1)\) such that, for every \(a\in(a^*,1)\) and all sufficiently small \(\epsilon>0\), the origin is a unique global maximizer of $u_{a,\epsilon}$. Moreover, \[ \lim_{a\downarrow a^*}\lim_{\epsilon\to0} \lambda_{\max}\bigl(D^2u_{a,\epsilon}(0)\bigr)=0. \] In addition, the ratios $\text{diam}(\Omega_{a,\epsilon}) / \text{inrad}(\Omega_{a,\epsilon})$ are uniformly bounded. Hence the Hessian estimate proved by Steinerberger (J. Funct. Anal. 274, 1611--1630, 2018) for convex planar domains cannot be extended to smooth, simply connected, nonconvex planar domains. \vskip0.2cm Our domains are star-shaped, symmetric with respect to both coordinate axes and convex in the horizontal direction. When the limiting slit is sufficiently long, we further show that certain superlevel sets of the torsion function are not star-shaped. This strengthens the counterexample to the star-shapedness question raised by Gladiali and Grossi (Amer. J. Math. 144, 1221--1240, 2022) under stronger geometric assumptions.

math.AP

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard outcome-based methods in medicine often suffer from autoregressive credit assignment failure and gradient variance explosion. This leads to the "Right Answer, Wrong Reason" trap, where models inadvertently reinforce spurious correlations and dataset shortcuts rather than valid clinical deduction. In this work, we propose Causally-Aligned Reasoning Exploration (CARE), a theoretically grounded framework for intrinsic experience curation. CARE is built upon two rigorous conditions for high-quality training trajectories: Causal Sufficiency, which utilizes an agreement-based self-verification mechanism to mimic $do$-calculus interventions and effectively debias gradients; and Proximal Learnability, which employs dynamic entropy bounds to select experiences within the model's zone of proximal development for variance-bounded optimization. These rigorously filtered experiences are optimized via a dual-stream objective that combines on-policy group-relative exploration with difficulty-weighted experience replay. Extensive experiments on diverse medical multimodal and text-only benchmarks demonstrate that CARE consistently outperforms other strong competitors, substantially reducing correct-but-inconsistent reasoning and improving training stability.

cs.CL

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation

Large language models (LLMs) are increasingly deployed as planners and assistants in tasks with inherent spatial structure, such as navigation and route planning, yet they remain brittle in sequential spatial reasoning. We ask not merely whether LLMs fail at navigation but where in the spatial-cognition pipeline they get lost. We introduce a multi-scale diagnostic benchmark that decomposes maze navigation into three cognitive levels drawn from human spatial cognition: Fine (local passability), Meso (junction topology), and Macro (global goal direction). We evaluate three instruction-tuned chat LLMs (GPT-4o, DeepSeek-V3, Llama-3.3-70B) on 1,050 topology-annotated mazes spanning seven sizes (3x3 to 30x30) and three difficulty tiers. The benchmark is organized as three modules. (i) Input acquisition: among four input formats, structured coordinate text is the most navigable, far surpassing rendered images. (ii) Multi-scale representation: end-to-end one-shot navigation collapses to near zero by 10x10 for every model, yet the same models respond to isolated single-level probes (Fine, Meso, Macro) at 30-75% far beyond that size. A multi-hot first-error analysis localizes failures to Meso junction choices (59%) and Fine perception (39%), with global direction almost never at fault (1%). The barrier is therefore the cross-scale aggregation of individually available competences over a long sequential plan, not any single perceptual deficit. (iii) Hierarchical route planning: delegating per-step execution to a deterministic walker and querying the LLM only at junctions, with an explicit cell-type prompt, lifts GPT-4o success by up to 92 points at mid sizes, but the same scaling wall re-emerges by 30x30. We release the benchmark, mazes, and code as a reusable diagnostic instrument for spatial reasoning in LLMs, available at https://yuhanjiang415.github.io/lost-in-aggregation/.

physics.soc-ph

One Currency, Two Forward Prices: The Onshore-Offshore Renminbi Puzzle

Partially convertible economies face a market-design problem: trade integration, cross-border investment, and domestic balance-sheet exposure increase the demand for currency hedging before full financial integration is complete. China adopted a distinctive architecture for this problem by fostering a deliverable offshore Renminbi market (CNH) alongside the segmented onshore market (CNY), rather than relying only on non-deliverable forwards. This creates two venues for closely related claims on the same currency. Spot prices are tightly linked, yet CNY and CNH forwards display a persistent and economically large discrepancy. We study that discrepancy in a joint equilibrium model for spot and forward trading with transaction costs and segmented supply. In the benchmark case with common constant supply and deterministic costs, spot parity implies a forward differential with the wrong sign relative to the data. Random offshore stress, modeled as a jump in trading costs, overturns this benchmark while preserving tight spot parity. The model yields a semi-explicit representation in the CNY/CNH application and a calibration of the observed forward discrepancy in terms of the market-implied likelihood and severity of offshore liquidity stress.

q-fin.MF

A note on convergence rate for reflected BSDEs with quadratic generators by penalization method

In this paper, we study the convergence rate between reflected backward stochastic differential equations with quadratic generators and their penalized BSDEs. Using techniques of BMO martingales, we prove the convergence rate is at order $\frac{1}{2}$ as a function of the penalty parameter. Finally, the result is applied to study numerical approximation of reflected BSDEs with sub-quadratic generators by the Euler's polygonal line method.

math.PR

31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding

This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outlier-free low-bit quantization, a stacking-aware PNM architecture co-designed with blockwise vector quantization to reduce weight EMA overheads, and an adaptive parallel speculative decoding scheme with an out-of-order scheduler for high resource and bandwidth utilization. Our chip achieves 14.08-to-135.69token/s and 4.46-to-7.17x speedup over vanilla speculative decoding.

cs.AR

Doubly Reflected Backward SDEs Driven by $G$-Brownian Motion with Quadratic Generator

In this paper, we study the doubly reflected backward stochastic differential equations driven by $G$-Brownian motion ($G$-BSDEs for short) when the generator has quadratic growth in the $z$-component. Based on the theory of $G$-BMO martingale and $G$-Girsanov theorem, we establish the existence and uniqueness result when the upper obstacle is almost a generalized $G$-It\^{o}'s process. Moreover, the solution can be approximated monotonically by the solutions to a family of penalized reflected $G$-BSDEs with a lower obstacle, which plays an important role to establish the relation between doubly reflected $G$-BSDEs and fully nonlinear partial differential equations with double obstacles.

math.PR

Weighted Bayesian Conformal Prediction

Machine learning predictors are rarely deployed on data that match the data they were calibrated on. Conformal prediction and its risk-control generalization promise distribution-free guarantees at deployment, but only under exchangeability or a covariate shift whose likelihood ratio is known exactly -- and the estimated-ratio case, the one that occurs in practice, is explicitly open, out of reach of the Bayesian construction available under exchangeability. We reach it from a different starting point: pushing a Dirichlet process prior on the calibration distribution through the likelihood ratio yields the exact Bayesian posterior over the deployed risk. That posterior is exact given the weight function, which at deployment is itself estimated, so we further prove finite-sample upper and lower bounds on the risk the selected threshold actually incurs, in terms of the weight-estimation error and the conditional shift. On synthetic and real covariate shifts, in regression and classification, the realized failure rate then stays close to target where shift-blind selection fails, and what deployment receives is a posterior over the risk of every threshold rather than a single number.

cs.LG

Qualitative analysis on the critical points of the Kirchhoff-Routh function

In this paper, we study the number of critical points of the Kirchhoff-Routh function \begin{equation*} \mathcal{KR}_D(x,y)=\Lambda_1^2\mathcal{R}_D(x)+\Lambda_2^2\mathcal{R}_D(y)-2\Lambda_1\Lambda_2G_D(x,y), \end{equation*} where $D$ is a bounded domain in $\mathbb{R}^2$, $x,y\in D$, $\Lambda_1,\Lambda_2>0$, $\mathcal{R}_D$ is the Robin function, and $G_D$ is the Green function of the operator $-\Delta$ with $0$ Dirichlet boundary condition on $D$. This function arises from concentration phenomena in nonlinear elliptic problems and from the de-singularization problem for the steady Euler equation. For domains with a small hole, we establish not only the exact number and the location of the critical points of $\mathcal{KR}_D$, but also their nondegeneracy. We show that the location of the hole plays a crucial role. Finally in the context of elliptic problems, we establish the existence of multiple two-peak solutions.

math.AP

Maximum principle for robust utility optimization via Tsallis relative entropy

This paper investigates an optimal consumption-investment problem featuring recursive utility via Tsallis relative entropy. We establish a fundamental connection between this optimization problem and a quadratic backward stochastic differential equation (BSDE), demonstrating that the value function is the value process of the solution to this BSDE. Utilizing advanced BSDE techniques, we derive a novel stochastic maximum principle that provides necessary conditions for both the optimal consumption process and terminal wealth. Furthermore, we prove the existence of optimal strategy and analyze the coupled forward-backward system arising from the optimization problem.

q-fin.MF

GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models

Geospatial modeling provides critical solutions for pressing global challenges such as sustainability and climate change. Existing large language model (LLM)-based algorithm discovery frameworks, such as AlphaEvolve, excel at evolving generic code but lack the domain knowledge and multi-step reasoning required for complex geospatial problems. We introduce GeoEvolve, a multi-agent LLM framework that couples evolutionary search with geospatial domain knowledge to automatically design and refine geospatial algorithms. GeoEvolve operates in two nested loops: an inner loop leverages a code evolver to generate and mutate candidate solutions, while an outer agentic controller evaluates global elites and queries a GeoKnowRAG module -- a structured geospatial knowledge base that injects theoretical priors from geography. This knowledge-guided evolution steers the search toward theoretically meaningful and computationally efficient algorithms. We evaluate GeoEvolve on two fundamental and classical tasks: spatial interpolation (kriging) and spatial uncertainty quantification (geospatial conformal prediction). Across these benchmarks, GeoEvolve automatically improves and discovers new algorithms, incorporating geospatial theory on top of classical models. It reduces spatial interpolation error (RMSE) by 13-21% and enhances uncertainty estimation performance by 17\%. Ablation studies confirm that domain-guided retrieval is essential for stable, high-quality evolution. These results demonstrate that GeoEvolve provides a scalable path toward automated, knowledge-driven geospatial modeling, opening new opportunities for trustworthy and efficient AI-for-Science discovery.

cs.AI

Distributed-HISQ: A Distributed Quantum Control Architecture

The design of a scalable Quantum Control Architecture (QCA) faces two primary challenges. First, the continuous growth in qubit counts has rendered distributed QCA inevitable, yet the nondeterministic latencies inherent in feedback loops demand cycleaccurate synchronization across multiple controllers. Existing synchronization strategies -- whether lock-step or demand-driven -- introduce significant performance penalties. Second, existing quantum instruction set architectures are polarized, being either too abstract or too granular. This lack of a unifying design necessitates recurrent hardware customization for each new control requirement, which limits the system's reconfigurability and impedes the path toward a scalable and unified digital microarchitecture. Addressing these challenges, we propose Distributed-HISQ, featuring: (i) HISQ, A universal instruction set that redefines quantum control with a hardware-agnostic design. By decoupling from quantum operation semantics, HISQ provides a unified language for control sequences, enabling a single microarchitecture to support various control methods and enhancing system reconfigurability. (ii) BISP, a booking-based synchronization protocol that can potentially achieve zero-cycle synchronization overhead. The feasibility and adaptability of Distributed-HISQ are validated through its implementation on a commercial quantum control system targeting superconducting qubits. We performed a comprehensive evaluation using a customized quantum software stack. Our results show that BISP effectively synchronizes multiple control boards, leading to a 22.8% reduction in average program execution time and a $\sim$5$\times$ reduction in infidelity when compared to an existing lock-step synchronization scheme.

quant-ph

Feature-free regression kriging

Spatial interpolation is a crucial task in geography. As perhaps the most widely used interpolation methods, geostatistical models -- such as Ordinary Kriging (OK) -- assume spatial stationarity, which makes it difficult to capture the nonstationary characteristics of geographic variables. A common solution is trend surface modeling (e.g., Regression Kriging, RK), which relies on external explanatory variables to model the trend and then applies geostatistical interpolation to the residuals. However, this approach requires high-quality and readily available explanatory variables, which are often lacking in many spatial interpolation scenarios -- such as estimating heavy metal concentrations underground. This study proposes a Feature-Free Regression Kriging (FFRK) method, which automatically extracts geospatial features -- including local dependence, local heterogeneity, and geosimilarity -- to construct a regression-based trend surface without requiring external explanatory variables. We conducted experiments on the spatial distribution prediction of three heavy metals in a mining area in Australia. In comparison with 17 classical interpolation methods, the results indicate that FFRK, which does not incorporate any explanatory variables and relies solely on extracted geospatial features, consistently outperforms both conventional Kriging techniques and machine learning models that depend on explanatory variables. This approach effectively addresses spatial nonstationarity while reducing the cost of acquiring explanatory variables, improving both prediction accuracy and generalization ability. This finding suggests that an accurate characterization of geospatial features based on domain knowledge can significantly enhance spatial prediction performance -- potentially yielding greater improvements than merely adopting more advanced statistical models.

physics.soc-ph

Reframing Spatial Dependence as Geographic Feature Attribution

Spatial dependence, referring to the correlation between variable values observed at different geographic locations, is one of the most fundamental characteristics of spatial data. The presence of spatial dependence violates the classical statistical assumption of independent and identically distributed observations and implies a high degree of information redundancy within spatial datasets. However, this redundancy can also be interpreted as structured information, which has been widely leveraged in spatial modeling, prediction, and explanation tasks. With the rise of geospatial big data and the rapid advancement of deep learning and large models, effectively modeling and characterizing spatial dependence has become essential for enhancing the performance of spatial analysis and uncovering latent spatial processes. From a data-driven perspective, this study proposes a novel interpretation: spatial dependence can be understood as the contribution of geographic location -- specifically, latitude and longitude -- to the observed variation in target variables. To validate this hypothesis, we conduct simulation experiments in which data are generated based on known spatial processes. We train machine learning models to predict variable values using only coordinate information. Subsequently, XAI techniques are employed to quantify the contribution of spatial features. The resulting importance scores are then compared with local indicators of spatial association (LISA). Across a range of spatial process settings, we observe consistently high correlations (greater than 0.94) between coordinate-based contributions and LISA values. These findings offer a new data-driven perspective on spatial dependence, bridging traditional spatial statistical approaches with modern machine learning techniques.

physics.soc-ph

Morse index, topological degree and local uniqueness of multi-spikes solutions to the Lane-Emden problem in dimension two

We consider multi-spike positive solutions to the Lane-Emden problem in any bounded smooth planar domain and compute their Morse index, extending to the dimension $N=2$ classical theorems due to Bahri-Li-Rey (1995) and Rey (1999) when $N\geq 4$ and $N=3$, respectively. Furthermore, by deeply investigating their concentration behavior, we also derive the total topological degree. The Morse index and the degree counting formula yield a new local uniqueness result.

math.AP

Extended mean field games with terminal constraint via decoupling fields

We consider a class of extended mean field games with common noises, where there exists a strictly terminal constraint. We solve the problem by reducing it to an unconstrained control problem by adding a penalized term in the cost functional and then taking a limit. Using the stochastic maximum principle, we characterize the solution of the unconstrained control problem in terms of a conditional mean field forward-backward stochastic differential equation (FBSDE). We obtain the wellposedness results of the FBSDE and the monotonicity property of its decoupling field. Based on that, we solve the original constrained problem and characterize its solution in terms of a system of coupled conditional mean field FBSDE with a free backward part. In particular, we obtain the solvability of a new type of coupled conditional mean field FBSDEs.

math.OC

Mechanistic Fine-tuning for In-context Learning

In-context Learning (ICL) utilizes structured demonstration-query inputs to induce few-shot learning on Language Models (LMs), which are not originally pre-trained on ICL-style data. To bridge the gap between ICL and pre-training, some approaches fine-tune LMs on large ICL-style datasets by an end-to-end paradigm with massive computational costs. To reduce such costs, in this paper, we propose Attention Behavior Fine-Tuning (ABFT), utilizing the previous findings on the inner mechanism of ICL, building training objectives on the attention scores instead of the final outputs, to force the attention scores to focus on the correct label tokens presented in the context and mitigate attention scores from the wrong label tokens. Our experiments on 9 modern LMs and 8 datasets empirically find that ABFT outperforms in performance, robustness, unbiasedness, and efficiency, with only around 0.01% data cost compared to the previous methods. Moreover, our subsequent analysis finds that the end-to-end training objective contains the ABFT objective, suggesting the implicit bias of ICL-style data to the emergence of induction heads. Our work demonstrates the possibility of controlling specific module sequences within LMs to improve their behavior, opening up the future application of mechanistic interpretability.

cs.CL