SearcharxivSearch

arXiv subjects

Zhao Zhang

Publications and source records attributed to Zhao Zhang.

At least 19 recordsLinked to original sources

Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection

Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.

cs.CV

Diamond Agent: Agentic Control of Federated HPC Resources as a Service

Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving workflow context across independently administered clusters, moving large datasets between sites, reasoning about site-specific environments and scheduler policies, and exploiting live queue and resource states for efficient task scheduling. To this end, we design Diamond Agent, an agentic system that enables intelligent execution of HPC workflows across heterogeneous clusters with typed skills as the interface. Diamond Agent provides an agent-facing workspace and skills that unify cross-site resource discovery, resource specification, data movement, task execution, and result retrieval. A centralized Diamond Agent instance can operate multiple supercomputers without being deployed separately on each login node. Diamond Agent translates high-level agent actions into valid site-specific executions, moves data through Globus Transfer, and uses live system capability and queue information to select feasible placements. Its event-driven continuation mechanism decouples agent actions from long-running batch jobs: persistent services monitor remote execution and resume the agent only when a result or decision-relevant event is available. We experiment with 27 hours of telemetry and 19 matched multi-site submission rounds comprising 83 jobs across four production supercomputers. Compared with a fixed-site baseline, Diamond Agent reduces the median additional completion time relative to the fastest observed placement from 42 seconds to 4 seconds, a 10.5x reduction.

cs.DC

A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks

Physics-informed neural networks (PINNs) commonly evaluate the spatial derivatives appearing in partial differential equation residuals using automatic differentiation (AD), whose computational and memory costs can become substantial when multiple or high-order derivatives are required. We perform a controlled comparison of spatial AD and Fourier spectral differentiation in periodic physical-space PINNs. Within each paired experiment, the neural representation, temporal differentiation, optimizer, sampling procedure, and training schedule are held fixed, so that the two cases differ only in the spatial differentiation procedure. For the Fourier variant, network outputs are evaluated on a uniform periodic grid and transformed to Fourier space, where spatial derivatives are obtained through spectral multiplication and the same Fourier coefficients are reused across derivative orders. We compare the two procedures in standard PINNs for the Allen--Cahn and Korteweg--de Vries equations and in Causal PINNs for the Allen--Cahn, Korteweg--de Vries, and Kuramoto--Sivashinsky equations. Across these five equation--framework settings, Fourier differentiation yields mean paired end-to-end training speedups ranging from $2.90\times$ to $18.52\times$ and reduces peak allocated graphics processing unit (GPU) memory by $68.7\%$--$94.1\%$. The final relative $L_2$ errors remain of the same order, with neither differentiation procedure showing a consistent accuracy advantage. For the one-dimensional periodic benchmarks considered here, Fourier spectral differentiation therefore provides substantially lower training time and memory usage than spatial AD while retaining comparable solution error, at the cost of requiring a uniform structured spatial grid.

cs.LG

Local-to-Global Sentence-Level Graph Reranking for Scientific Synthesis

Retrieval-augmented scientific synthesis aims to answer complex research questions by integrating information from multiple papers into comprehensive and well-grounded responses. Since the generator can only synthesize the information selected and organized by the reranker, the quality of the generated synthesis depends critically on the reranked results. However, most rerankers operate at the passage level, which leaves key methodological, empirical, and comparative information buried in long and flat contexts, weakening the grounding of generated claims. Moreover, existing rerankers mainly rely on independent query-candidate scoring which overlooks complementary, contextual, and contrasting relations across scientific candidates, limiting information coverage and the comprehensiveness of the resulting synthesis. To address these limitations, we propose LoG-Reranker, a local-to-global sentence-level graph reranking framework for scientific synthesis. LoG-Reranker performs role-aware local scoring to identify fine-grained, query-relevant sentences and then models their relations on a sentence graph across the candidate set to globally refine sentence rankings. Top-ranked sentences and their connected neighbors are organized into a structured input context for generator to produce more grounded and comprehensive synthesis. Extensive experiments on scientific synthesis and reranking benchmarks show that LoG-Reranker consistently outperforms competitive rerankers, yielding more reliable rankings and improving the quality of generated synthesis.

cs.IR

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches, low-rank compression is particularly attractive because it represents every token in reduced dimensions. Previous low-rank methods typically derive fixed projection spaces from model weights, construct fixed spaces from calibration activations, or construct a shared basis over a broad cache region. Such representations may not capture detailed but important information. We partition each per-head KV cache into fixed-length logical pages and observe substantial low-rank structure within individual pages. Based on this observation, we propose PuzzleKV, a training- and calibration-free method that treats each completed page as an independent compression unit. PuzzleKV decomposes pages within each layer and KV head, computes attention directly over dense and factorized pages, and incrementally compresses newly eligible pages during autoregressive decoding. Experiments across models, context lengths, and benchmarks demonstrate the effectiveness of PuzzleKV under matched storage budgets. At approximately 60% of the original KV cache storage, PuzzleKV achieves more than 96% of Full KV performance across both evaluated models and all benchmark settings, with substantial gains over Global SVD on RULER and competitive performance on LongBench. To achieve a more aggressive compression ratio, PuzzleKV can be further combined with quantization while retaining more than 93% of Full KV performance using only 18.7% of the original storage.

cs.LG

Estimation of the cumulant $b_4$ and sextic self-coupling of the QCD axion at finite temperature and density

The sixth-order cumulant $b_4$ of the QCD topological charge distribution and the axion's sextic self-coupling at finite temperature $T$ and baryon chemical potential $\mu$ are calculated within the two-flavor quark-meson (QM) model without and with the Polyakov-loop dynamics. The $b_4$ in the vacuum falls within the range predicted by lattice SU(3) pure-gauge simulations, and at large $T$ and/or $\mu$, it tends to a constant predicted by the dilute instanton gas approximation. Both $b_4$ and the axion's sextic self-coupling diverge at the QCD critical point and flip sign across the phase transition. The convergence of the fourth- and sixth-order Taylor expansions of the free energy as a function of the $\theta$ parameter is investigated via comparison with the full potential at the mean-field level.

hep-ph

Compiler Framework for 3D Neutral-Atom Quantum Computers

Neutral-atom quantum computers can now arrange atoms in three-dimensional tweezer arrays, yet every existing compiler assumes a flat geometry. We present Piqasso, a compiler that exploits the vertical axis by stacking storage, entanglement, and readout into distinct layers. Its pipeline pairs an analytical placement respecting axial-clearance optics with a router that brings gate partners together via short vertical hops---bypassing in-plane crossing conflicts through out-of-plane detours---and a multi-AOD scheduler that parallelizes transport across focal planes. On 34 circuits, Piqasso reduces atom transport distance by 2.1$\times$ over a state-of-the-art planar compiler, yielding up to 7.3$\times$ faster execution, 2.2$\times$ higher movement fidelity, and 1.8$\times$ fewer serialized transport rounds, with all gains widening at scale.

quant-ph

Radially correlated partially coherent beams with a deterministic vortex structure

Partially coherent beams have attracted considerable attention due to their intrinsic resilience against complex environmental perturbations. However, the intrinsic wavefront fluctuations make it fundamentally challenging to preserve well-defined orbital angular momentum during propagation. In this work, we propose and experimentally demonstrate a class of radially correlated, partially coherent beams that carry deterministic vortex structures, generated via optical conformal mapping from Cartesian to log-polar coordinates. The resulting beams exhibit a ring-shaped coherence distribution, characterized by low coherence in the radial direction and high coherence in the azimuthal direction. This unique feature of such a beam supports a well-defined deterministic vortex phase, thereby enabling the beam to preserve its ring-shaped coherence distribution during propagation through a focusing system. Our results provide new insights into the design of new partially coherent beams and may facilitate the development of applications in optical encoding, free-space information transmission, and ultrafast light-matter interactions.

physics.optics

Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?

Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each node and all nodes across the traffic network. However, the underlying mechanism by which global information is modeled and extracted remains insufficiently investigated. Whether global information must be extracted by high-degree-of-freedom adaptive attention or can be captured by a simple global aggregation operator remains unclear. For this purpose, we design a controlled ablation framework that replaces only the spatial mixing module to test attention-based global interaction. Across six traffic benchmarks, standard spatial attention yields relative MAE changes of $-1.58\%$ to $+1.26\%$ compared with uniform full-range mixing, and we observe no consistent advantage for standard spatial attention, while uniform full-range mixing reduces node-scale spatial-mixing complexity from $O(N^2)$ to $O(N)$. We further propose a hypothesized model that decomposes spatial attention into a row-uniform global background and a non-uniform residual. The residual shows dataset-dependent effects. Overall, uniform full-range mixing provides a strong global spatial baseline, while the non-uniform attention residual is not consistently beneficial across datasets.

cs.AI

A cryogenic neutral-atom platform with full optical access and 2-hour trap lifetime

Neutral-atom quantum processors are rapidly scaling toward system sizes of more than ten thousand qubits, allowing for the realization of a new class of quantum computing algorithms and quantum simulation experiments. However, current neutral-atom platforms generally have to find a compromise between the optical accessibility and the storage time of atoms in optical potentials, limiting the available qubit numbers. Here we report on the operation of a novel, cryogenically enhanced, neutral-atom apparatus that overcomes these apparently conflicting requirements. We demonstrate vacuum-limited trapping lifetimes of up to two hours of single $^{88}\mathrm{Sr}$ atoms in an optical tweezer array while preserving full optical access and without the need for complex cryogenic enclosures. Our measurements show that exceptionally long single-atom lifetimes can be achieved with a relatively simple cryostat design. Our architecture can be straightforwardly ported to other atomic species and shows a viable path for scaling up to sorted arrays of tens of thousands of atoms.

physics.atom-ph

Absolute frequency measurement of the $^{176}$Lu$^+\,(^{3}\mathrm{D}_1)$ standard against the NRC-FCs2 fountain with $2.6\times10^{-16}$ uncertainty

We report an improved absolute frequency measurement of the $^{176}$Lu$^+\,(^{3}\mathrm{D}_1)$ optical frequency standard, evaluated via a remote link to the NRC-FCs2 caesium fountain primary frequency standard. Operating a single ion clock with 94.2% uptime over 10 days, and using an ambiguity-resolved precise point positioning (PPP-AR) link over the Global Positioning System (GPS), we determine an absolute frequency of $353\,638\,794\,073\,800.33(9)\,$Hz at a fractional uncertainty of $2.6 \times 10^{-16}$. This agrees with our previous result, which underpins the CIPM recommended frequency value, and reduces the uncertainty by a factor of 3.6.

physics.atom-ph

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the communication volume. Existing methods quantize gradients via linear or nonlinear mappings in Euclidean space, often degrading model performance because highly anisotropic gradients incur direction-dependent distortion. We present GIFT, a geometry-informed gradient scaling method that performs low-precision communication in geometry-aware coordinates. By transforming gradients into a near-isotropic space before quantization, GIFT makes low-precision representations substantially more faithful to their high-precision counterparts. GIFT only changes the coordinate system used for low-precision gradient communication and does not change the optimizer, training recipe, communication collective, or low-precision format. We also develop a simplified geometry-aware transformation algorithm with low-rank approximation and selective application to balance the computation overhead and communication reduction. We examine the empirical convergence of GIFT using Llama-300M and Llama-600M models. Our results show that GIFT reduces the end-to-end pretraining time of Llama-600M by 7.6% on 64 NVIDIA GH200 Superchips, while improving the downstream task preservation profile over direct Euclidean FP8 communication under the same optimizer and communication path.

cs.DC

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. Existing fault-tolerance mechanisms either impose non-trivial overhead during failure-free execution or suffer from prolonged recovery latency, particularly under scenarios where a small subset of compute nodes experience permanent failures. %The tradeoff between failure-free overhead and recovery latency forms a space forms a Pareto frontier We present PHOENIX to simultaneously address both optimization objectives. PHOENIX incorporates a fault-tolerance mechanism that restores LLM training via hot-swapping, namely by replacing failed nodes with spare nodes without terminating the complete job. The hot-swapping of PHOENIX is enabled by two ideas: First, it exploits an off-critical-path in-memory checkpointing mechanism for spatial redundancy. Second, it introduces a communicator reconstruction protocol that replaces failed nodes with spare nodes at runtime. PHOENIX efficiently overlaps the in-memory checkpointing with computation, thus introducing zero overhead during error-free execution. Upon permanent node failures, PHOENIX can rebuild memory states with minimal recomputation by leveraging in-memory checkpoints. We evaluate PHOENIX across scales (up to 512 NVIDIA A100 GPUs) and LLMs (up to 65B parameters), and observe zero checkpoint overhead with hot-swapping recovery completing in under 40 seconds. These results show that PHOENIX simultaneously achieves both zero-overhead error-free execution and extremely low recovery cost.

cs.LG

SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication

Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reconstruction overhead increase with model size and system scale. Existing communication-reduction methods either sparsify raw gradients, which can be unstable for modern Adam-style optimizers at high sparsity, or quantize communication, whose savings are fundamentally bounded by bit width and often incur additional runtime overhead. We present SCAPE, a communication-efficient distributed optimizer for LLM training that exploits the stability of AdamS's first-moment to enable aggressive sparsification without loss of LLM quality. Instead of constructing masks from raw gradients, SCAPE derives them from first-moment-based statistics, partitions mask generation across workers to align with optimizer sharding, and delays mask usage by one step so that mask synchronization can overlap with computation. SCAPE also reconstructs the quantities required for second-moment updates from a single synchronized sparse buffer, avoiding an additional collective. We implement SCAPE in Megatron-LM and evaluate its convergence by pre-training GPT-345M on OpenWebText and Llama-500M on SlimPajama-6B using 32 NVIDIA GH200 GPUs on TACC Vista. In both models, SCAPE preserves training stability, validation loss, and downstream task accuracy under 90\% and 99\% sparsity. For Llama-500M, SCAPE reduces end-to-end pre-training wall-clock time by up to 43.3\% while maintaining model quality comparable to dense AdamW and AdamS. For Llama-1.8B, SCAPE achieves up to 3.26$\times$ speedup per step compared to dense AdamS.

cs.LG

Improving Muon-Scattering Material Identification via Coarse Momentum Encoding and Unsupervised Domain Adaptation

Cosmic-ray muon scattering has shown considerable potential for detecting nuclear materials and other dense contraband, but practical deployment remains challenging. A major difficulty arises from the coupling between material properties and muon momentum, since the broad natural momentum distribution influences the scattering angle and prevents unambiguous material identification. In this work, we propose a Coarse Momentum-Aware Domain Adaptation (CMADA) method to enable precise identification of materials. Instead of relying on high-precision momentum measurements, the proposed framework adopts coarse momentum binning combined with unsupervised domain adaptation to learn transferable scattering representations. In addition, a precision review mode based on averaging repeated samplings was proposed to further enhances identification performance. The coarse momentum binning strategy improves same-domain identification accuracy from 62.15% without momentum information to 89.52% with 5-bin momentum information, and further to 93.37% (precision review mode). Furthermore, the proposed unsupervised domain adaptation framework improves the cross-domain identification accuracy from 71.71% for the source-only baseline to 89.00% without requiring target domain labels.

physics.ins-det

CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding

Code retrieval is becoming central to coding agents, but agentic coding requires more than matching a natural-language query to an isolated snippet. Given a user request, a coding agent needs to navigate a concrete repository state, locate relevant files and functions, gather supporting context, and filter similar in-repository distractors. Existing code retrieval benchmarks mainly evaluate docstring-to-function or snippet-level matching, thereby missing this requirement-driven repository search problem. To address this gap, we introduce CORE-Bench, a comprehensive benchmark for code retrieval in the era of agentic coding. CORE-Bench evaluates code retrieval ability at three levels: code understanding, issue-to-edit localization, and broader context retrieval. Built from curated code-search tasks and SWE-bench-series instances, CORE-Bench contains over 180K queries and 106K broader-context relevance labels. Experiments with representative embedding models show a sharp drop from traditional code search to code retrieval in agentic coding settings. Simple supervised fine-tuning of existing embedding models significantly improves performance in this setting, suggesting substantial room for further progress.

cs.IR

The Yang-Baxter Equation for the Chiral Potts Model and Integrable Parafermions

A new type of Yang-Baxter equation (YBE) for $R$-operators depending on three spectral parameters is constructed from the star-triangle relation for the chiral Potts model. As the $Z_N$ symmetric generalization to the Ising model, its Boltzmann weights are known to depend on two variables describing a curve with genus larger than one for $N>2$, except for the self-dual point corresponding to the Fateev-Zamolodchikov chain. Combined with the fact that quantum Hamiltonians of edge-type models such as the Ising model contain both nearest-neighbor interaction and onsite potential terms, this leads naturally to an additional spectral parameter in the associated $R$-operator. The construction extends the edge-vertex correspondence of solvable lattice models, and provides a bridge between the Bazhanov-Stroganov four-parameter $R$-matrix---realized as an intertwiner of cyclic representations of $U_q(\mathfrak{sl}_2)$ at a root of unity---and Shastry's two-parameter $R$-operator obtained from the decorated YBE.

math-ph

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $\pi_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.

cs.RO