SearcharxivSearch

arXiv subjects

Yiming Xu

Publications and source records attributed to Yiming Xu.

At least 19 recordsLinked to original sources

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present $\Phi$-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, $\Phi$-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.

cs.CL

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve the underlying 3D scene state or translate into executable actions. We introduce RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, covering 50 manipulation tasks across four regimes, with 5,000 episodes and 25,000 multi-view ground-truth videos. A defining feature of RoboPhys-3D is that generated and ground-truth videos are processed through the same 3D reconstruction pipeline, enabling reconstruction-induced error to be distinguished from generation-induced error. The RoboPhys-3D benchmark organizes 50 complementary metrics into 18 sub-dimensions across four levels: pixel-level fidelity, 3D geometry consistency, state-level understanding, and task-level completeness. We further introduce Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success. Among the four representative video world models, Cosmos 3 achieves the highest RoboPhyscore (0.6330, 92.7% of ground truth), while state- and execution-grounded metrics reveal substantial failures that perceptual and vision-language model-based judgments fail to capture. RoboPhyscore further exhibits strong agreement with human evaluation (Pearson r = 0.9761 and Spearman \r{ho} = 0.8962), demonstrating the importance of grounded, execution-aware evaluation for EWM capability.

cs.RO

Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models

LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show that this hidden convention is fragile: across 2 VLMs and 7 image and video benchmarks, the default layer is sub-optimal in 13 of 14 model-task pairs, and the best layer shifts with both task and visual backbone. Finding that layer by exhaustive layer-wise inference is prohibitively expensive, and no better fixed default exists. We therefore ask whether layer usefulness can instead be predicted from representation geometry. We study matrix-based entropy, introduced for unimodal layer analysis, which we compute over sample-level visual embeddings as Visual Dataset Entropy (VDE); and Gromov-Wasserstein (GW) distance, introduced for encoder-level VLM model selection, which we repurpose as a layer-wise visual--language alignment signal. Transferring these to LLaVA-based models is not obvious a priori: the vision tower is frozen while the multimodal projector is trained, so we profile both sides of the projector. We find that VDE transfers, and GW does not. Computed from 100 unlabeled task samples without downstream inference, pre-projector VDE tracks layer-wise accuracy and its top-ranked layers cover the oracle best layer on every task for the SigLIP-based LLaVA-Video, while giving region-level guidance for the CLIP-based Video-LLaVA. Post-projector profiles show that the projector reshapes visual geometry but does not erase the performance-relevant trend, leaving $\mathrm{VDE}_{\mathrm{pre}}$ the stronger signal. GW instead flattens after projection and is best read as an alignment diagnostic rather than a selector. VDE thus offers an interpretable, training-free policy that narrows the visual-layer search to a handful of candidates for limited downstream verification.

cs.CV

An Interactive, Automated 4D-STEM data acquisition and analysis routine for Scanning Electron Nanobeam Diffraction and Ptychography experiments

Modern transmission electron microscopes are versatile instruments which have become indispensable tools for understanding structure and chemical composition at the nano- and atomic scale. In the physical sciences these instruments are still largely manually controlled, requiring significant operator expertise, limiting throughput, and precluding statistical analysis of large datasets. Recent technical advances in both hardware and in control software now allow for the interaction with almost every functionality of the microscope through a programming interface. This enables better experimental design and data collection automation while also reducing operator collection bias and required expertise. In this study, we present an automated data collection routine with machine-driven decision-making to enable the collection of hundreds of 4D-STEM nanobeam diffraction and ptychography data from a large distribution of size-selectively deposited Pt nanoparticles. We present a semi-automated data analysis workflow to extract pertinent information from the large volumes of collected data. For the nanobeam diffraction data, reducing each dataset to its azimuthal variance profile and combining automated crystal orientation mapping with per-particle morphology descriptors reveals the orientation, shape and phase distributions across the ensemble, including a weak {110} texture. For the ptychographic data, an automated screening pipeline identifies on-zone-axis particles and enables atomic-resolution phase imaging and lattice-strain mapping of individual grains. Together these demonstrate how automation turns instrument throughput into statistically meaningful, atomic-scale microstructural information.

physics.ins-det

Balancing fractional Brownian motion

We study the discrepancy of balancing $n$ independent sample paths of fractional Brownian motion with Hurst exponent $H\in(0,1)$ on $[0,1]$, an infinite-dimensional analogue of balancing Gaussian vectors. We establish a phase transition at $H=1/2$: with high probability, the discrepancy is $\Omega(n^{1/2-H})$ and $\mathcal O(n^{1/2-H}(\log n)^{c(H)})$, where $c(H)=H+1/2$ if $H\geq 1/2$ and $c(H)=1/2$ otherwise. At the critical exponent $H=1/2$, we show that the discrepancy is $\Theta(1)$ with constant probability as $n\to\infty$. In this regime, we further characterize the geometry of the solution space by computing the expected number of local minima, establishing an overlap gap property near the existence threshold, and proving its absence at every diverging optimality threshold. We also give randomized polynomial-time algorithms that compute signings with discrepancy $\mathcal O(n^{1/2-H}\sqrt{\log n})$ for $H<1/2$, $\mathcal O((\log n)^{3/2})$ for $H=1/2$, and $\mathcal O(\sqrt{\log n})$ for $H>1/2$, with high probability. Our analysis combines a truncated balancing argument based on a wavelet representation of fractional Brownian motion with probabilistic methods.

math.PR

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered by three major challenges: (1) the scarcity of datasets for complex, logic-rich diagrams; (2) the absence of explicit layout priors, which leads to chaotic spatial arrangements; and (3) the lack of fine-grained visual feedback to validate rendered outputs and correct aesthetic defects. To address these challenges, at the data level, we introduce DocMeetSVG-100K, a large-scale SVG dataset tailored for document authoring and meeting review scenarios. At the model level, we propose GVR-Coder, a novel framework designed to generate high-quality logical diagrams from lengthy professional texts. Specifically, we adopt a curriculum-driven rejection sampling fine-tuning to progressively enhance the model's capability in modeling complex structures, while explicitly incorporating layout constraint knowledge during training. In addition, we introduce reinforcement learning from dual rendering feedback, a mechanism that provides implicit feedback through reward signals to jointly optimize structural complexity and visual aesthetics. Furthermore, we design a generate-verify-repair agent loop, which improves generation quality through explicit, fine-grained feedback and targeted refinement. Extensive experiments demonstrate that GVR-Coder outperforms competitive baselines and reliably produces logically coherent and visually appealing diagrams. Code and data are available at https://github.com/CurryaNa/GVR-Coder.

cs.LG

A Hierarchical Optimisation Framework for Integrated Electric-Hydrogen-Transport Systems

Integrated electric-hydrogen infrastructures are becoming increasingly important with the growing deployment of electric vehicles (EVs) and hydrogen vehicles (HVs) in transport systems. However, the strong coupling between vehicle scheduling and multi-energy dispatch introduces significant operational challenges. This paper models an integrated electric-hydrogen-transport system (EHTS) and proposes a hierarchical optimisation framework that couples vehicle scheduling and downstream energy dispatch through a sequential, demand-driven two-layer structure. In the vehicle scheduling layer, a solver-free greedy heuristic (SFGH) algorithm is developed to avoid repeated optimisation solving, enabling real-time EV charging and HV refuelling under non-preemptive service and within-interval sequential assignment. The resulting charging and refuelling demands are subsequently passed to the energy dispatch layer, where a deep reinforcement learning (DRL)-based approach is designed to optimise battery operation, hydrogen-tank operation, and PV generation allocation to minimise the overall operational cost of the EHTS while satisfying the scheduled transport demand. Representative case studies, together with comparative, ablation, and generalisation analyses, demonstrate the effectiveness and robustness of the proposed framework. Furthermore, the learned dispatch policy maintains strong performance across diverse transport-demand scenarios without retraining, demonstrating robust generalisation capability for practical deployment.

eess.SY

Convergence analysis of a family of Zermelo-type iterations for the Bradley--Terry model

Zermelo's algorithm is a classical method for computing the maximum likelihood estimator in the Bradley--Terry (BT) model, but its convergence can be slow in practice. To accelerate computation, Newman introduced a family of Zermelo-type fixed-point iterations parameterized by $\alpha$, with Zermelo's algorithm recovered at $\alpha=1$. Empirical evidence suggests that the choice $\alpha=0$ often converges substantially faster, making it a promising alternative, yet the mechanism underlying this acceleration remains elusive. This paper provides theoretical insight into this phenomenon through a systematic local convergence analysis. We derive closed-form expressions for local convergence factors under synchronous and asynchronous updates and analyze their dependence on $\alpha$ via spectral analysis of the associated Jacobian matrices. For synchronous updates, we show that the algorithm may fail to converge when $\alpha<1$, and its local convergence factor is quasi-convex in $\alpha$ under the population BT model. In contrast, asynchronous updates are always locally convergent, and their local convergence factor is provably monotonically increasing in $\alpha$ under the population BT model of consistently ordered bipartite comparison graphs, establishing the optimality of $\alpha=0$ in this setting. We further establish asymptotic approximation results for the population convergence factors under the BT model, justifying their practical relevance. Numerical experiments on synthetic and real-world datasets confirm the theory. Our analysis complements existing convergence results and shows that the acceleration of $\alpha=0$ arises not only from the parameter choice but, more importantly, from the use of asynchronous updates.

stat.ML

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention as a way to improve structured outputs without retraining model weights. An earlier project-specific approach, Orthogonal Residual Projection (ORP), exposed sensitive SwiGLU FFN sites and non-monotonic energy effects, but its direction-changing operation produced more regressions than repairs in a key diagnostic. We therefore propose Amplitude Gating (AG), which preserves pretrained FFN weight directions and modulates activation magnitudes during decoding. AG separates candidate generation, ranking, and a prospective acceptance/fallback decision. We also introduce Per-Sample Fix-Harm Evaluation (PFHE), a paired reporting protocol that complements native task metrics with fixes, harms, preserved-correct cases, and preserved-wrong cases. On the only cross-position union that passes source-alignment audit, an exploratory offline mixed selector raises the descriptive heterogeneous-scorer Qwen3.5-9B tool-route micro-average from 38.66% to 42.92% (+4.27 percentage points); two Hermes function-call endpoints improve by +7.64 and +7.62 points. The same-output PFHE-format view records 48 fixes, 26 harms, 294 preserved-correct cases, and 2,188 preserved-wrong cases over 2,556 units, with positive paired bootstrap intervals for native and strict effects. Protocol-separated Qwen3-8B and Qwen2.5-7B analyses retain oracle headroom but no positive train-selected fixed tool route. A grouped five-fold RF diagnostic suggests weak nonlinear ranking signal but forces intervention, lacks baseline fallback and paired uncertainty, and is not deployment evidence. The results support model- and task-specific selection with strict fallback, not a universal AG switch.

cs.CL

From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

Large Language Models (LLMs) are rapidly reshaping academic research across the natural sciences, social sciences, and humanities, yet the scientific community lacks a comprehensive, cross-disciplinary account of how these tools are being integrated, what they deliver, and where they fall short. This paper addresses that gap by mapping their current state and outlining an agenda for their responsible integration into scientific research. Our analysis reveals a consistent pattern: LLMs meaningfully accelerate research workflows -- from hypothesis generation and literature synthesis to data analysis and scientific writing -- while introducing serious challenges related to hallucination, reproducibility, dataset bias, and model opacity. Beyond technical limitations, we identify ten underexplored challenges, including the erosion of researcher autonomy, AI-driven confirmation bias, authorship ambiguity, and unequal access to these technologies -- systemic risks that demand interdisciplinary governance frameworks, robust validation standards, and expanded explainability research.

cs.DL

Phast: Simultaneous reconstruction of photoelectron count and time profiles from PMT waveforms via machine learning

Photomultiplier tubes (PMTs) are widely used in particle and nuclear physics experiments. The reconstruction of PMT waveforms is a fundamental task in these experiments, where accurate extraction of photoelectron (PE) multiplicities and time from the waveform is required for downstream event reconstruction and analysis. In realistic detector environments, PMT waveform reconstruction is complicated by electronic effects such as pileup, charge fluctuations, noise etc., which make precise recovery of physical observables challenging. To address these challenges, we present \phast{}, a machine-learning-based method that reconstructs PE count and time profile simultaneously. The model consists of a shared wave-transformer encoder followed by two dedicated branches: a counting branch for the total PE number prediction, and a time branch employing a count-conditioned query decoder with dynamic query activation. To study the reconstruction performance under controlled conditions, we construct several toy Monte Carlo PMT waveform datasets, including both uniform and mixed fast-slow double-temporal-components configurations. The proposed method demonstrates stable and accurate reconstruction performance across various waveform conditions, achieving high consistency in both PE counting and time reconstruction. These results indicate that architectures combining convolutional feature extraction with query-based transformer decoders provide an effective approach for complex PMT waveform reconstruction tasks.

hep-ex

Generalist Graph Anomaly Detection via Prototype-Based Distillation

Driven by the pressing demand for graph anomaly detection (GAD) in high-stakes domains, the generalist GAD paradigm, which trains a single detector transferable across new graphs, has recently gained growing attention. However, existing methods often rely on scarce and costly annotations for training and sometimes even require few-shot support at inference, which limits their robustness to diverse and unseen anomaly patterns. To address this limitation, we introduce ProMoS, the first unsupervised generalist GAD framework, which detects anomalies by modeling the abundant normality in unlabeled data. ProMoS adopts a knowledge-distillation paradigm to distill normality priors from a frozen self-supervised graph neural network (GNN) teacher to a mixture-of-students model with shared global and lightweight personalized branches, enabling efficient and expressive normality modeling without learning from scratch. We further propose prototype-guided soft-label distillation to align teacher and student in a shared prototype space, enhancing cross-graph generalizability. During inference, ProMoS performs zero-shot anomaly detection on unseen graphs via distillation bias and prototype geometric deviation. Extensive experiments show the effectiveness and efficiency of ProMoS, charting a practical path toward label-free, zero-shot generalist GAD.

cs.LG

TED: Related Party Transaction guided Tax Evasion Detection on Heterogeneous Graph

Tax evasion causes severe losses of government revenues and disturbs the economic order of fair competition. To help alleviate this problem, the latest tax evasion detection solutions utilize expert knowledge to extract features and then train classifiers to determine whether a company is suspected of tax evasion. However, existing solutions mainly focus on the statistical features of the company, but fail to exploit the rich interactive information in tax scenarios, which affect the detection performance. In this paper, we first model the tax scenario as a heterogeneous graph and study the tax evasion detection problem under the heterogeneous graph model. To improve the performance of tax evasion detection, a novel graph neural network model is proposed to extract the comprehensive information of heterogeneous graphs. Specifically, we use heterogeneous and complex related party transaction groups to filter low-level noise information. Moreover, a hierarchical attention mechanism is designed to capture the deeper structure and semantic information hidden in the related party transaction group. We apply our method to the real risk management system of the tax bureau, and evaluate it on two human-labeled real-world tax datasets. The results demonstrate that our method significantly outperforms the state-of-the-art in the tax evasion detection task.

cs.LG

Learning Dynamic Graph Representations through Timespan View Contrasts

The rich information underlying graphs has inspired further investigation of unsupervised graph representation. Existing studies mainly depend on node features and topological properties within static graphs to create self-supervised signals, neglecting the temporal components carried by real-world graph data, such as timestamps of edges. To overcome this limitation, this paper explores how to model temporal evolution on dynamic graphs elegantly. Specifically, we introduce a new inductive bias, namely temporal translation invariance, which illustrates the tendency of the identical node to keep similar labels across different timespans. Based on this assumption, we develop a dynamic graph representation framework CLDG that encourages the node to maintain locally consistent temporal translation invariance through contrastive learning on different timespans. Except for standard CLDG which only considers explicit topological links, our further proposed CLDG++ additionally employs graph diffusion to uncover global contextual correlations between nodes, and designs a multi-scale contrastive learning objective composed of local-local, local-global, and global-global contrasts to enhance representation capabilities. Interestingly, by measuring the consistency between different timespans to shape anomaly indicators, CLDG and CLDG++ are seamlessly integrated with the task of spotting anomalies on dynamic graphs, which has broad applications in many high-impact domains, such as finance, cybersecurity, and healthcare. Experiments demonstrate that CLDG and CLDG++ both exhibit desirable performance in downstream tasks including node classification and dynamic graph anomaly detection. Moreover, CLDG significantly reduces time and space complexity by implicitly exploiting temporal cues instead of complicated sequence models.

cs.LG

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stability. The performance of ADAS critically relies on accurate and reliable vehicle state estimation, particularly from vehicle dynamic sensors. Among these signals, wheel load is a key variable for chassis control and safety-critical functions, yet it remains difficult to estimate robustly due to complex suspension geometry, nonlinear dynamics, and measurement noise. To address this issue, we propose DBPnet, a Bayesian physics-informed neural network (PINN) with a physics-aware embedding module inspired by damper characteristics. First, this paper presents a suspension linkage-level modeling (SLLM) approach that constructs a nonlinear instantaneous dynamic model by explicitly considering the complex geometric structure of the suspension. Building upon SLLM, Bayesian inference is integrated into the PINN to effectively cope with noise and uncertainty in the vehicle chassis system, thereby improving the model's robustness. Then, a physics-informed loss function is employed to ensure consistency with fundamental physical principles, while the damper characteristics-inspired embedding module extracts temporal variation features of input signals and incorporates them into each layer of the PINN, ensuring that physical observations guide the neural network without being constrained by fixed physical models. Extensive evaluations on high-fidelity simulations and real-world experiments demonstrate that our DBPnet consistently achieves lower RMSE and MaxError than baseline methods. These results highlight the potential of our DBPnet to advance wheel load estimation and contribute to the development of more reliable ADAS actuator functions.

eess.SY

Solar phased arrays-based wireless power transfer for commercial airlines can reduce energy costs and carbon emissions in the United States

Decarbonizing aviation remains challenging because energy-dense jet fuels dominate beyond short-range operations, while batteries impose severe range and payload penalties. Here we evaluate a new infrastructure pathway in which utility-scale solar farms equipped with solar phased arrays wirelessly beam microwave power to hybrid-electric aircraft during cruise. Integrating 143,152 U.S. flight trajectories, 5,712 solar farms and wireless power transfer models, we quantify the spatial, temporal, and operational potential of this concept at continental scale. We find that benefits are highly concentrated in solar-rich, traffic-dense states and are dominated by short- and medium-range flights, accounting for nearly all delivered energy and cost savings. Schedule optimization and higher cruise altitudes further increase value by improving alignment between aircraft demand and beaming availability. Market penetration analysis reveals non-linear scaling between solar farm and flight adoption. These results show that wireless power beaming is best understood as a corridor-specific strategy complementing other aviation decarbonization pathways.

eess.SY

MKG-CARE: Case-Aware Reasoning with Multimodal Knowledge Graphs for Explainable Medical Image Diagnosis

Medical image diagnosis has achieved significant progress with deep learning, yet existing methods often rely on isolated visual evidence and lack the ability to effectively leverage similar cases and external knowledge. In clinical practice, diagnosis is typically supported by similar historical cases and their associated symptoms. To explicitly model this evidence-based diagnostic process, we propose MKG-CARE, a framework that performs case-aware reasoning using multimodal knowledge graphs for explainable medical image diagnosis. Specifically, we construct a case-aware multimodal knowledge graph as a structured diagnostic memory, where diseases, images, and symptoms are hierarchically organized. Given an input image, MKG-CARE adaptively retrieves similar cases from this memory and extracts their corresponding case-centered subgraphs. We further introduce a knowledge propagation and injection mechanism, where an image-centric Graph Attention Network aggregates heterogeneous semantics within the retrieved case subgraphs, followed by bidirectional cross-modal attention to align and inject the aggregated case knowledge into visual representations. To mitigate retrieval noise, we design a confidence-calibrated decision refinement scheme that estimates each retrieved case's reliability from prediction confidence and sample similarity, and reweights its contribution to the final prediction for interpretable case-level evidence attribution. Extensive experiments on multiple medical imaging datasets demonstrate consistent improvements over strong baselines, while ablation and qualitative analyses validate the effectiveness and interpretability of our method. The code is available at https://github.com/lyxuan1022/MKG-CARE.

cs.CV

RepoZero: Can LLMs Generate a Code Repository from Scratch?

Large Language Models (LLMs) have recently shown remarkable progress in code generation, yet their ability to construct complete software repositories from scratch remains poorly understood. A fundamental bottleneck is the lack of verifiable and scalable evaluation: existing benchmarks either focus on patch-based editing or rely on human or LLM-based judgments, which introduce bias and limit reproducibility. In this work, we present RepoZero, the first benchmark that enables fully automated, execution-based verification of repository-level generation from scratch. Our key idea is to reformulate generation as repository reproduction: given only API specifications, an agent must re-implement an entire repository such that its behavior matches the original implementation. This design allows for strict black-box validation via output equivalence, while naturally supporting large-scale construction by reusing existing open-source repositories. To further mitigate data leakage and shortcut solutions, we introduce cross-language constraints and a sandboxed evaluation protocol. Building on this benchmark, we propose an Agentic Code-Test Evolution (ACE) framework that performs iterative test generation and error-driven refinement, enabling effective test-time scaling for repository-level synthesis. Extensive experiments across multiple state-of-the-art LLMs and agent frameworks reveal that even the strongest LLM agents achieve only limited pass rates (30\% - 55\%), exposing a substantial gap between current capabilities and real-world software development requirements. Our results establish RepoZero as a challenging, scalable, and reliable testbed for end-to-end code generation, and highlight self-verification via test generation as a critical direction for advancing LLM-based coding agents.

cs.SE