SearcharxivSearch

arXiv subjects

Yifan Zhu

Publications and source records attributed to Yifan Zhu.

At least 19 recordsLinked to original sources

Diamond Agent: Agentic Control of Federated HPC Resources as a Service

Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving workflow context across independently administered clusters, moving large datasets between sites, reasoning about site-specific environments and scheduler policies, and exploiting live queue and resource states for efficient task scheduling. To this end, we design Diamond Agent, an agentic system that enables intelligent execution of HPC workflows across heterogeneous clusters with typed skills as the interface. Diamond Agent provides an agent-facing workspace and skills that unify cross-site resource discovery, resource specification, data movement, task execution, and result retrieval. A centralized Diamond Agent instance can operate multiple supercomputers without being deployed separately on each login node. Diamond Agent translates high-level agent actions into valid site-specific executions, moves data through Globus Transfer, and uses live system capability and queue information to select feasible placements. Its event-driven continuation mechanism decouples agent actions from long-running batch jobs: persistent services monitor remote execution and resume the agent only when a result or decision-relevant event is available. We experiment with 27 hours of telemetry and 19 matched multi-site submission rounds comprising 83 jobs across four production supercomputers. Compared with a fixed-site baseline, Diamond Agent reduces the median additional completion time relative to the fastest observed placement from 42 seconds to 4 seconds, a 10.5x reduction.

cs.DC

Finite-sample nonparametric mean tests: Leave-one-out duality and asymptotic optimality

We study finite-sample valid tests of the one-sided mean hypothesis $H_0:μ\leq 1$ against $H_1:μ>1$ for nonnegative random variables. To do so, we develop a leave-one-out dual certificate framework, where certain pointwise inequalities imply p-value validity under the conditional mean null $\mathbb{E}[X_i\mid\mathbf{X}_{-i}]\leq 1$, and which also gives conditions that allow combining dual certificates for p-values to show that their pointwise minimum is also a valid p-value. The framework proves finite-sample validity of Wang and Zhao's nonparametric likelihood-ratio statistic $T_{\mathrm{nplr}}$, yields a new p-value $T_{\mathrm{bin}+}$ extending the Clopper--Pearson binomial test to general nonnegative random variables, and shows that the pointwise minimum $\min\{T_{\mathrm{nplr}},T_{\mathrm{bin}+}\}$ is itself a valid and more powerful p-value. We establish sharp optimality results for such testing problems in two regimes: both $T_{\mathrm{nplr}}$ and $T_{\mathrm{bin}+}$ attain a universal detectability boundary for the null $H_0$ without moment or tail assumptions, and $T_{\mathrm{bin}+}$ attains a nonparametric power lower bound under $n^{-1/2}$-local alternatives to $H_0$. Efficient algorithms and numerical experiments demonstrate substantial finite-sample power gains over existing valid methods.

math.ST

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning latent preferences and evaluation constraints, where the resulting belief updates then guide later proposals. We formalise this setting as ProSE, a hidden-parameter sequential assistance problem, and instantiate it with a KL-regularised bounded-rational binary response model in which acceptance trades off value gain against a distance-dependent evaluability penalty. Analysing the planning consequence of this likelihood reveals that likely accepted proposals and informative probes need not coincide, which explains why planners that only pursue acceptance systematically underperform. We operationalise ProSE with \textsc{ProSE-Plan}, a depth-2 Bayes-adaptive planner that scores proposals by possible responses and response-induced posterior beliefs. In controlled graph simulations, \textsc{ProSE-Plan} improves over evaluability-unaware and myopic baselines when evaluation cost is the bottleneck, and a probe-commit ablation confirms that our approach selects informative proposals that simpler methods miss. Our results thus identify user evaluability as a planning-relevant dimension of AI assistance, complementary to generation quality and preference inference.

cs.AI

Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment

Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically optimize response-level preferences without explicitly modeling the latter. We operationalize Theory-of-Mind (ToM) inference as a control signal within Frictive Policy Optimization by extracting, at each referring expression, a four-part belief structure: the speaker's intended referent, the addressee's interpretation, and each participant's model of the other's belief. This makes friction mechanically computable from epistemic-state comparisons, capturing \emph{silent divergence}, where both participants proceed confidently while grounding to different referents. We evaluate the signal at two levels. At the representation level, ablating the second-order channel reduces misunderstanding recall from $65\%$ to $26\%$. At the policy level, reward-shaping (FAR) and trust-region (FTR) variants improve intervention F1 and warranted-context calibration over DPO, with Brier scores independently supporting the calibration gains. Across three training runs, FAR and FTR remain substantially more stable, whereas DPO varies widely and can degrade intervention competence already present in the base policy. Thus, ToM-grounded friction provides a trainable signal for context-sensitive intervention under referential belief divergence.

cs.CL

Assessing Parameter Redundancy in Transformers for Jet Tagging

Transformer-based jet taggers, such as the Particle Transformer (ParT) and the More-Interaction Particle Transformer (MIParT), achieve excellent discrimination by exploiting correlations among jet constituents, but often require more trainable parameters than earlier deep-learning taggers. In this paper, we investigate whether comparable discriminating power can be achieved with substantially fewer parameters. We introduce an hourglass structure that replaces the feed-forward networks (FFNs) in the attention blocks while leaving the particle-interaction attention unchanged. We also introduce a lightweight particle-embedding layer to replace the original dense embedding network. Applying both modifications to ParT and MIParT yields the hourglass (HG) variants ParT-HG and MIParT-HG, respectively. We evaluate both models on benchmark datasets for top tagging and quark-gluon discrimination. Both variants retain comparable tagging performance, including background rejection at fixed signal efficiencies, while using only approximately 48% and 39.7% of the parameters of their respective baselines. On the larger JetClass dataset, accuracy and AUC decrease by less than 1%, and background rejection also decreases for several signal classes. Overall, our approach provides an alternative way to reduce the parameter count of Transformer jet taggers while largely retaining their tagging performance.

hep-ph

Global Convergence of an SQP Method for Contact-Implicit Trajectory Optimization

Contact-Implicit Trajectory Optimization (CITO) is a powerful framework for planning motions of robots that interact with complex environments, but its convergence behavior remains difficult to characterize. Existing formulations either rely on off-the-shelf nonlinear programming solvers whose guarantees require constraint qualifications that are hard to verify for contact-rich systems, or require differentiable explicit dynamics maps that are difficult to formulate in the presence of impacts, changing contact modes, and geometric non-smoothness. This paper studies CITO for a broader class of constraint-rich dynamic systems formulated through implicit physics constraints. By exploiting maximal-coordinate structure, penalty relaxations of equality and inequality constraints, and finite-horizon Hamiltonian bounds, we prove global convergence in the numerical-optimization sense for a specific line-search Sequential Quadratic Programming (SQP) method with adaptive timestep refinement. Under the stated constraint-rich dynamic-system assumptions and fixed pre-horizon data satisfying the strict-interior predecessor condition, the method terminates finitely from any finite discretized optimized trajectory guess at an $ε$-feasible, force-parameterized stationary point---a unit-objective, penalty-induced force-parameterized Fritz--John certificate for smooth damping, weakening to a homogeneous force-parameterized Clarke--Fritz--John certificate whose objective multiplier may vanish under nonsmooth frictional contact.

math.OC

Multi-Branch Policy Optimization for Multimodal Large Language Models

Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially higher perceptual uncertainty than text-only settings, where the model must repeatedly re-examine visual information to verify intermediate interpretations, and different visual groundings can lead to divergent reasoning paths, making such uniform credit assignment particularly inadequate and causing relative advantages to progressively degenerate toward zero. To address these challenges, we propose Multi-Branch Policy Optimization (MBPO), a tree-based framework that constructs reasoning trees at vision-language decision boundaries, enabling sibling branches to explore diverse visual hypotheses and assigning segment-level credit through branch-relative advantages. We further introduce a temporal replay buffer to reuse informative segments while controlling policy staleness. Experiments on several multimodal reasoning benchmarks show that MBPO outperforms representative baselines, improving both learning signal quality and optimization efficiency. The code is publicly available at https://github.com/ShuaiLyu0110/MBPO.

cs.CV

Pressure-induced Superconductivity in Thermoelectric Semiconductor Mg3Sb2

The intrinsic electronic structures of narrow bandgap thermoelectric (TE) materials serve as a platform for the investigation of coupling effects of quasi-particles under high pressure, enabling the exploration of emerging electronic and phonon transport, superconductivity, and topological transition. Here, we report the discovery of pressure-induced superconductivity in the TE semiconductor Mg3Sb2. Upon the increased pressure, the metallization occurs at 8.7 GPa, followed by a superconducting transition concomitant with a carrier-type crossover from p- to n-type. This phenomenon arises from a pressure-induced structural phase transition from the semiconducting P-3m1 to the metallic C2/m-I phase. The superconducting critical temperature (Tc) exhibits a dome-shaped pressure dependence, peaking at 3.3 K at 12.6 GPa. Combined theoretical calculations, high-pressure Raman spectroscopy, and X-ray diffraction (XRD) measurements reveal an additional structural transition above 20 GPa, yielding a distinct C2/m-II phase. Our findings establish the high-pressure phase diagram of Mg3Sb2, elucidate its pressure-dependent electronic properties, and provide valuable insights for future investigations of TE materials under high pressure.

cond-mat.supr-con

SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption

Multimodal gradient balancing methods modulate encoder gradients with a shared scalar per modality, implicitly assuming that corruption is uniform across the training batch. In practice, corruption is sample-heterogeneous: within a single mini-batch, different samples may have different modalities corrupted. We prove that under this heterogeneous corruption model, any batch-level sample-agnostic linear estimator with a shared modulation parameter incurs an irreducible bias with respect to the clean-data gradient, and that sample-level all-or-nothing gating is the unique unbiased strategy within a natural distribution-free estimator class. Motivated by this result, we propose Sample-Adaptive Gradient Gating (SAGG), which makes a binary retain-or-discard decision per sample via an online feature-norm quality test and incorporates a truncation mechanism for variance control. We prove that SAGG-based SGD converges at the standard O(1/sqrt(T)) rate to stationary points of the clean loss without a corruption-dependent error floor, and derive a certified robustness radius for the independent-encoder architecture that connects per-modality Lipschitz constants to the classification margin. Experiments on Kinetics-Sounds and UCF-101 under Gaussian noise injection, partial modality missing, and natural contribution imbalance show that SAGG consistently outperforms ten existing methods, with the largest gains in high-corruption regimes where batch-level bias is most severe.

cs.LG

Deep-learning jet flavor tagging for precision hadronic Higgs measurements at future $e^+e^-$ Higgs factories

Precise measurements of Higgs decays into quarks and gluons are essential for probing the Yukawa couplings of the Higgs boson and testing the flavor structure of the Standard Model. We investigate the process $e^+e^- \to ZH$ at $\sqrt{s}=240~\mathrm{GeV}$ at a future $e^+e^-$ Higgs factory, taking the CEPC design as a benchmark. The analysis focuses on events with $Z\toν\barν$ and hadronic Higgs decays $H\to b\bar{b}$, $c\bar{c}$, $s\bar{s}$ and $gg$. Jet flavor is identified using state-of-the-art particle-level deep neural network taggers (ParticleNet, Particle Transformer and More-Interaction Particle Transformer), whose per-jet outputs are combined with global event observables in a two-stage analysis employing XGBoost classifiers to separate the four Higgs decay modes from the dominant two- and four-fermion Standard Model backgrounds. Assuming an integrated luminosity of $20~\mathrm{ab}^{-1}$, we obtain projected relative precision on $σ(ZH)\times\mathrm{Br}(H\to X)$ of 0.17% for $X=b\bar{b}$, 1.06% for $c\bar{c}$, 0.50% for $gg$ and 68% for $s\bar{s}$. Compared with the CEPC published results, the precisions for $H\to c\bar{c}$ and $H\to gg$ are improved by about 43% and 29%, respectively. For $H\to s\bar{s}$ we present a quantitative sensitivity estimation corresponding to a statistical significance of about $1.5σ$. These results highlight the potential of deep-learning-based jet flavor tagging for precision studies of Higgs decays at future $e^+e^-$ Higgs factories.

hep-ph

Elastic Trapped States at Dislocation Defects in Scaled Coupling and Hofstadter Models

Elastic topological dislocations provide a pathway for trapping elastic wave energy at internal defects, rather than being confined solely to external boundaries or corners, which are typically associated with topological insulators (TIs). However, two practical constraints persist. First, highly confined dislocation states based on conventional Su-Schrieffer-Heeger (SSH) dimerization usually require a large coupling contrast and a correspondingly enlarged bandgap, which may be challenging to realize. Second, some Hamiltonians with richer topological physics often contain complex hopping terms, synthetic gauge fields or nonlocal couplings, which substantially increase the geometric complexity of experimental samples. Here, dislocation-induced trapped states are demonstrated in both a scaled coupling (SC) model and a Hofstadter model (HM) within an elastic platform. In the SC model, the trapped mode is treated as a higher localized state in the continuum rather than an in-gap mode in the SSH model. Consequently, the SC-induced dislocation can trap an enhanced mode without the requirement of an enlarged bandgap. For the HM, Householder tridiagonalization is used to map the original tight-binding Hamiltonian with complex hopping terms onto a tridiagonal matrix with only positive-real-valued nearest-neighbour (NN) hopping terms. Truncation at a weak-hopping position preserves the topological phenomena and allows a dislocation defect to be constructed from the shortened aperiodic chain. The results establish a practical route for designing highly localized modes without relying solely on bandgap enlargement or complex couplings, which advance the topological physics of elastic wave systems and promise enhanced possibilities for elastic functional devices.

physics.app-ph

Chiral Landau levels induced by two in-plane pseudomagnetic fields in underwater acoustic metamaterials

The chiral zeroth Landau levels (LLs) constitute topologically protected bulk states that enable robust control of acoustic wave propagation. Given the central role of underwater acoustics in marine engineering, realizing such Landau-level physics in underwater acoustic systems is highly desirable. Nevertheless, existing studies have primarily been limited to airborne acoustic systems, and the implementation of chiral zeroth LLs in underwater acoustics remains a challenge due to the unavoidable fluid-solid interactions. In this study, we realize two kinds of chiral LLs in an open underwater spoof surface acoustic wave (SSAW) platform by introducing two perpendicular in-plane artificial pseudomagnetic fields (PMFs), oriented along the x and y directions, respectively, and reveal that scalar acoustic fields in water and vectorial elastic vibrations in solids can be jointly manipulated within a unified framework. Specifically, by strategically opening bandgaps at the Dirac points, position-dependent effective mass terms are introduced into the Dirac Hamiltonians, thereby synthesizing two in-plane PMFs. This results in the emergence of chiral LLs, which is confirmed both numerically and experimentally. The unidirectional propagation of the chiral LLs and their robustness against defects are also demonstrated. In addition, we achieve flexible manipulation of underwater ultrasonic energy carried by SSAWs, including beam splitting and arbitrary wave steering. Dual-band chiral LLs are also observed in small-scale underwater topological metamaterials. Our work provides a new route toward SSAW-based underwater ultrasonic control, opening opportunities for multiband underwater acoustic signal processing and detection, as well as underwater acoustic energy harvesting.

physics.app-ph

A GPU-Accelerated Framework for Multi-Attribute Range Filtered Approximate Nearest Neighbor Search

Range-filtered approximate nearest neighbor search (RFANNS) is increasingly critical for modern vector databases. However, existing solutions suffer from severe index inflation and construction overhead. Furthermore, they rely exclusively on CPUs for the heavy indexing and query processing, significantly restricting the throughput due to the limited memory bandwidth and parallelism. In this paper, we present Garfield, a GPU-accelerated framework for multi-attribute range filtered ANNS that overcomes these bottlenecks through designing a lightweight index structure and hardware-aware execution pipeline. Garfield introduces the GMG index, which partitions data into cells and builds local graph indexes. It guarantees linear storage and indexing overhead by adding a constant number of cross-cell edges. For queries, Garfield utilizes a cluster-guided ordering strategy that reorders query-relevant cells, enabling a highly efficient cell-by-cell traversal on the GPU that aggressively reuses candidates as entry points across cells. To handle datasets exceeding GPU memory, Garfield features a cell-oriented out-of-core pipeline. It dynamically schedules cells to minimize the number of active queries per batch and overlaps GPU computation with CPU-to-GPU index streaming. Extensive evaluations demonstrate that Garfield reduces index size by 4.4x, while delivering 119.8x higher throughput than state-of-the-art RFANNS methods.

cs.DB

GUI-AC: Enhancing Continual Learning in GUI Agents

Graphical User Interfaces (GUIs) serve as the dominant medium for human-computer interaction, yet building GUI agents that generalize across the vast diversity of real-world interface environments, with the same flexibility and robustness that humans naturally exhibit, remains unsolved. Notably, GUI data are inherently non-stationary: the continual emergence of previously unseen interface instances (e.g., novel domains and resolutions) induces persistent distribution shifts, significantly impeding the continual learning of existing GUI agents. Reinforcement fine-tuning (RFT) has attracted considerable attention as a promising approach. Nevertheless, RFT exhibits pronounced instability in its grounding capability, manifested as sharp reward discontinuities and high-variance oscillations. The imbalanced distribution of rollout outcomes introduces substantial noise into advantage estimation, leading to policy overconfidence. The fixed clipping bound suppresses the increase in policy probabilities needed to adapt to new distributions, leading to a collapse in exploration capacity. To address these challenges, we propose GUI-AC, a method that enhances the continual learning capability of GUI agents. GUI-AC introduces grounding certainty to support two core mechanisms: (i) Adaptive Advantage, which down-weights noisy advantage estimates to prevent policy overconfidence; and (ii) Dynamic Clipping, which relaxes the clipping bound to encourage exploration range. Extensive experiments show that these mechanisms jointly improve performance, enabling our method to surpass state-of-the-art baselines. Code is available anonymously at https://github.com/Can-Lin/GUI-AC.

cs.CV

Experimental Realization of Type-II Quadrupole Topological Insulator

The discovery of quadrupole topological insulators (QTIs) has spurred extensive research into higher-order topological phases. Recently proposed type-II QTIs exhibit unconventional topological behaviors with 1/2 edge polarization \operatorname{p}_x and zero edge polarization \operatorname{p}_y, due to the inequivalence between Wannier-band and edge-spectrum gap closures, yet their experimental realization remains challenging owing to the long-range and complex off-site hopping terms in their tight-binding model (TBM). Here, we circumvent this difficulty via an optimized Householder tridiagonalization (OHT) mapping that reduces the complex two-dimensional lattices to one-dimensional chains with only negative-real-valued nearest-neighbor hopping terms, greatly facilitating experimental sample fabrication. Using this strategy, we experimentally verify the type-II QTI phase, type-I QTI phase and trivial phase in elastic wave platforms via simple aperiodic plate-beam chain structures, where the plates reflect the on-site potential terms and beams correspond to the off-site hopping terms in the TBM. Our approach provides a versatile route for experimentally exploring more complex and richer topological phenomena based on TBM.

physics.app-ph

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

We prove that Thompson sampling exhibits $\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)})$ Bayesian regret in the linear-Gaussian bandit with a $\mathcal{N}(μ_0, Σ_0)$ prior distribution on the coefficients, where $d$ is the dimension, $T$ is the time horizon, $r$ is the maximum $\ell_2$ norm of the actions, and $σ^2$ is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ``burn-in'' term $d r \sqrt{\mathrm{Tr}(Σ_0)}$ decouples additively from the minimax (long run) regret $σd \sqrt{T}$. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ``elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.

cs.LG

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive access to databases. While recent approaches leveraging large-scale private LLMs such as GPT-4 have achieved state-of-the-art results, they face two critical challenges: the lack of openness and reproducibility, and the prohibitive computational cost of test-time scaling. To address these issues, we explore improving the model-level performance of small-scale public LLMs in NL2SQL under resource-constrained settings. Our exploratory experiments reveal the potential of task decomposition for enhancing NL2SQL performance, but also highlight the difficulty of enabling LLMs to decompose queries effectively. Motivated by these findings, we propose LearNAT, a novel framework designed to enhance decomposition capabilities of LLM. LearNAT introduces (1) a Decomposition Synthesis Procedure, which leverages AST-guided search with pruning strategies to generate verifiable and efficient decompositions, and (2) Margin-Aware Reinforcement Learning, which provides fine-grained preference optimization for multi-step reasoning beyond standard DPO. Extensive experiments on benchmark datasets demonstrate that LearNAT significantly improves the performance of small-scale LLMs, achieving results comparable to GPT-4 with only a 7B parameter model. These results validate the effectiveness of verifiable decomposition and fine-grained preference learning in advancing NL2SQL towards openness, transparency, and efficiency. Our code is publicly available at https://github.com/MrBlankness/LearNAT.

cs.CL

BrepLLM: Enabling Large Language Models to Understand Boundary Representations

Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain complex geometric and topological information. To this end, we propose BrepLLM, the first multimodal framework that enables LLMs to directly parse and reason over raw B-rep data. BrepLLM adopts a two-stage training pipeline: cross-modal alignment pre-training and two-stage LLM fine-tuning. In the first stage, we design an adaptive UV sampling strategy to convert B-reps into graph representations that integrate geometric and topological information. Subsequently, we construct a hierarchical BrepEncoder to extract features from geometric elements (faces and edges) and topology, generating a global token and a sequence of node tokens. Then, via contrastive learning, we conduct an initial alignment between this global token and the text embeddings of a frozen CLIP text encoder (ViT-L/14). In the second stage, we integrate the pre-trained BrepEncoder into the LLM and employ a two-stage progressive strategy to align the sequence of node tokens: (1) training an MLP-based semantic mapping network that utilizes the prior knowledge of a 2D-VLM to align the B-rep representation to the 2D visual semantic space; (2) utilizing LoRA for parameter-efficient fine-tuning of the Q-Former and the LLM backbone network to achieve the final 3D-language generation capability. Furthermore, we construct the Brep2Text dataset, which contains 269,444 B-rep and text question-answer pairs. Experiments demonstrate that BrepLLM achieves SOTA performance on 3D object classification and captioning tasks. The project page is available at https://user-deng.github.io/BrepLLM/.

cs.CV