Searcharxiv⌕ Search

arXiv subjects

Sheng Yang

Publications and source records attributed to Sheng Yang.

At least 37 records · Page 2Linked to original sources

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward methods typically rely on costly tree sampling. We view world knowledge as a latent world graph and each IS task as search within a latent task graph, where effective steps should make graph progress toward the answer node. Based on this prior, we propose Graph-Distance Contribution Reward (GDCR), a step-level process reward that scores newly-retrieved and newly-cited entities by their distance to the answer node in a training-time Entity-Relation (ER) graph. We further propose Step Advantage Policy Optimization (SAPO), which converts GDCR into step-level advantages and combines them with trajectory-level outcome advantages. Experiments on four challenging benchmarks validate the effectiveness of our method.

cs.AI↗

GSMap: 2D Gaussians for Online HD Mapping

Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-based approaches preserve topology but struggle with geometric fidelity, while rasterization-based approaches enable precise geometric supervision but produce unstructured outputs. To bridge this gap, we propose GSMap, a novel framework that unifies both paradigms via a learnable 2D Gaussian representation. Each map element is modeled as an ordered sequence of 2D Gaussians, whose centers correspond to the vertices of the vectorized polyline/polygon. This formulation enables simultaneous optimization through: (1) Differentiable rasterization that enforces pixel-level geometric constraints, and (2) Topology-aware vectorization that maintains structural regularity. Experiments on both nuScenes and Argoverse2 demonstrate that our Gaussian-based representation effectively unifies geometric and topological learning, achieving significant performance improvements and demonstrating strong compatibility with existing HD mapping architectures. Code will be available at https://github.com/peakpang/GSMap

cs.CV↗

Generalized Li-Haldane Correspondence in Critical Dirac-Fermion Systems

Topological phenomena in quantum critical systems have recently attracted growing attention, as they go beyond the traditional paradigms of condensed matter and statistical physics. However, a general framework for identifying such nontrivial phenomena, particularly in higher-dimensional systems, remains insufficiently explored. In this work, we propose a universal fingerprint for detecting nontrivial topology in critical free-fermion systems protected by global on-site symmetries. Specifically, we analytically establish an exact relation between the bulk entanglement spectrum and the boundary energy spectrum at topological criticality in arbitrary dimensions, demonstrating that the degeneracy of edge modes can be extracted from the bulk entanglement spectrum. These findings, further supported by numerical simulations of lattice models, provide a universal fingerprint for identifying nontrivial topology in critical free-fermion systems.

cond-mat.str-el↗

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability to perceive, interpret, and act over heterogeneous contexts such as images, videos, webpages, documents, GUIs. GLM-5V-Turbo is built around this objective: multimodal perception is integrated as a core component of reasoning, planning, tool use, and execution, rather than as an auxiliary interface to a language model. This report summarizes the main improvements behind GLM-5V-Turbo across model design, multimodal training, reinforcement learning, toolchain expansion, and integration with agent frameworks. These developments lead to strong performance in multimodal coding, visual tool use, and framework-based agentic tasks, while preserving competitive text-only coding capability. More importantly, our development process offers practical insights for building multimodal agents, highlighting the central role of multimodal perception, hierarchical optimization, and reliable end-to-end verification.

cs.CV↗

Second-Order Bilevel Optimization with Accelerated Convergence Rates

This paper studies second-order methods for nonconvex-strongly-convex bilevel optimization. We propose a novel fully second-order bilevel approximation method (FSBA) that achieves an iteration complexity of $\tilde{\mathcal{O}}(ε^{-1.5})$ for finding the $(ε, \mathcal{O}(\sqrtε))$ second-order stationary point of the hyper-objective function. Our results demonstrate that second-order methods can achieve an accelerated convergence rate than first-order methods in bilevel optimization. To address the heavy computational cost associated with the second-order oracle, we introduce a lazy variant of FSBA, called LFSBA, which reuses second-order information across several iterations. We prove that LFSBA exhibits better computational complexity than FSBA by a factor of $\sqrt{d}$, where $d$ is the dimension of the problem. We also apply a similar idea to nonconvex strongly-concave minimax optimization and propose the lazy minimax cubic-regularized Newton (LMCN) method with better computational complexity compared to existing second-order methods.

math.OC↗

Design and preliminary performance study of the broad-band spectrometer detector for POLAR-2

POLAR-2, the successor of the POLAR experiment aboard China's Tiangong-2 space lab, is set to be deployed on the China Space Station. The POLAR-2 mission aims to conducting high-precision polarization measurements of high-energy transients with a primary focus on Gamma-Ray Bursts (GRBs), following POLAR's pioneering accurate polarization measurements of GRB prompt emission. One of the key advancements in POLAR-2 is the inclusion of a dedicated Broad-band Spectrometer Detector (BSD) instrument, designed to provide precise measurements of GRB location and spectral parameters, which are critical inputs for accurate polarization analysis of POLAR-2's dedicated High-energy Polarimetry Detector (HPD), which is made of plastic scintillator bars array. BSD employs a coded-aperture mask imaging technique and pixelated GAGG scintillation crystals, offering a wide half-coded field of view of ~132° x 125° and an operational energy range of 10-1000 keV. Simulation results indicate that the instrument can achieve a localization accuracy of approximately 1.5° for faint GRBs similar to GRB 170817A, satisfying the core requirements of GRB polarimetry with HPD. BSD also has moderate capability for GRB polarimetry, particularly at several hundred keV energy. This paper outlines the preliminary design of BSD and presents an overall evaluation of its expected scientific performance, based on extensive Monte Carlo simulations and preliminary ground-based calibration tests.

astro-ph.IM↗

Anomalous Dynamical Scaling at Topological Quantum Criticality

We study the nonequilibrium driven dynamics at topologically nontrivial quantum critical points (QCPs), and find that topological edge modes at criticality give rise to anomalous dynamical scaling behavior. By analyzing the driven dynamics of bulk and boundary order parameters at topologically distinct QCPs in quantum spin chains, we demonstrate that, while the bulk dynamics remain indistinguishable and follow standard Kibble Zurek (KZ) scaling, the anomalous boundary dynamics are unique to topological criticality, obeying modified scaling relation beyond the traditional KZ framework. To elucidate the unified origin of this anomaly, we further study the dynamics of defect production at topologically distinct QCPs in free-fermion models and demonstrate similar anomalous scaling exclusive to topological criticality. These findings establish the existence of anomalous dynamical scaling arising from the interplay between topology and driven dynamics, challenging standard paradigms of quantum critical dynamics.

cond-mat.str-el↗

A search for successful and choked jets in nearby broad-lined Type Ic supernovae

The observational link between long gamma-ray bursts (GRBs) and broad-lined stripped-envelope core-collapse supernovae (SNe Ic-BL) is well established. Significant progress has been made in constraining what fraction of SNe Ic-BL may power high- or low-luminosity GRBs when viewed at small off-axis angles. However, the GRB-SN connection still lacks a complete understanding in the broader context of massive-star evolution and explosion physics. Models predict a continuum of outcomes for the fastest ejecta, from choked to ultra-relativistic jets, and observations from radio to X-rays are key to probing these scenarios across a range of viewing angles and velocities. Here, we present results from a coordinated radio-to-X-ray campaign targeting nearby (z<=0.1) SNe Ic-BL designed to explore this diversity. With eight new radio-monitored events and updated data for one previously observed SN, we further tighten constraints on the fraction of SNe Ic-BL as relativistic as SN 1998bw/GRB 980425. We identify SN 2024rjw as a new radio-loud event likely powered by strong interaction with circumstellar material (CSM), and add evidence supporting a similar interpretation for SN 2020jqm. We also establish new limits on the properties of radio-emitting ejecta with velocities consistent with cocoons from choked jets, highlighting SN 2022xxf as a promising cocoon-dominated candidate. These results refine our understanding of the continuum linking ordinary SNe Ic-BL, engine-driven explosions, and GRBs, and contribute to building a sample that will inform future multi-messenger searches for electromagnetic counterparts to high-energy neutrinos.

astro-ph.HE↗

GECAM discovery of a peculiar magnetar X-ray burst (MXB 221120) from SGR J1935+2154 associated with a fast radio burst

Fast radio bursts (FRBs) are enigmatic cosmic transients of millisecond duration observed in the radio band. The identification of FRB-associated magnetar X-ray bursts (MXBs) from galactic magnetar SGR J1935+2154 suggests that at least a fraction of FRBs can be produced from magnetar activity. However, the sample size of FRB-associated MXBs is still very small. Here we report a bright and peculiar FRB-associated MXB from SGR J1935+2154 detected by GECAM on November 20, 2022, dubbed MXB 221120. We find that both temporal and spectral properties of MXB 221120 exhibit distinctive features. Its light curve could be generally described by a single FRED function with superposition of several narrow pulses. Interestingly, we identify a possible QPO feature with center frequency of ~18 Hz in this MXB. The time-integrated spectrum is best fitted by a blackbody model with temperature (kT ) of 18.6 keV, rendering it the first thermal spectrum FRB-associated MXB from SGR J1935+2154. Compared to other MXBs with single emission episode, MXB 221120 has longer duration and higher blackbody temperature, making it an outlier in the burst sample. These results indicate that MXB 221120 may be produced by a special mechanism with extreme physical conditions.

astro-ph.HE↗

Comprehensive Measurement of Spectral Evolution in a GRB Flare: High Time-Resolution Insights into the "Double-Tracking" Phenomenon

The spectral evolution characteristics of the prompt emission in gamma-ray bursts (GRBs) have been extensively studied, but detailed investigations of spectral evolution in a GRB flare remain lacking. In this work, we present the first analysis of spectral parameter evolution in a GRB flare through high time-resolved spectral fitting of the Brightest Flare in GRB 221009A. We find that the $α$-Flux, $E_p$-Flux, and $E_p$-$α$ relationships during both the overall phase and the rise phase of flare can be well described by simple power-law model, showing positive correlations. Therefore, we conclude that Brightest Flare exhibits "Double-tracking" behavior. Since values of $α$ do not exceed the synchrotron "death line" (-2/3), we explain this phenomenon using a magnetic dissipation synchrotron radiation model. In the decay phase of flare, the $E_p$-Flux and $E_p$-$α$ correlations become notably flatter, with their power-law indices decreasing significantly compared to those in the rise phase. This may be due to the fact that the next flare begins to erupt before the Brightest Flare has completely ended, resulting in the combined effects of both two flares. Our study of spectral parameter relations of the Brightest Flare provides new insights into the radiation mechanisms of both GRB prompt emission and flares.

astro-ph.HE↗

A Telescope System for Charge and Position Measurement of High Energy Nuclei

A high-granularity telescope system with a large sensitive area and low material budget has been developed for high-energy heavy ion beam tests. The telescope consists of nine layers of silicon microstrip detectors (SSDs), whose performance was validated through a heavy ion beam test at the CERN SPS. A hybrid machine learning algorithm is proposed to address the challenges of nuclear charge measurement with SSDs. The system achieves a spatial resolution of $\mathcal{O}(1) \,$\SI{}{\micro\metre} and a charge resolution better than 0.16 charge units for nuclei from $Z = 1$ to $Z = 29$, with a sensitive area of $8 \times 8 \, \mathrm{cm}^2$. To the best of our knowledge, this represents the most precise charge and spatial resolution simultaneously achieved by a silicon telescope to date.

physics.ins-det↗

Beam Test Characterization of Silicon Microstrip Detector Flight-Model Ladders for the AMS-02 Upgrade

The AMS-02 experiment plans to install a new silicon microstrip tracker layer (Layer-0) on top of the existing detector, increasing the cosmic-ray acceptance by a factor of 3. Layer-0 employs a design in which multiple silicon microstrip detectors (SSDs) are connected in series to form long detector ladders. We present a detailed performance study of the flight-model ladders using a 350~GeV mixed hadron beam at the CERN SPS. The study focuses on the following aspects: (i) the performance of ladders with different numbers of SSDs, for which the intrinsic spatial resolution at normal incidence varies from $9.5~μ\mathrm{m}$ to $11.4~μ\mathrm{m}$ for ladders composed of 8 to 12 SSDs; (ii) the response consistency for particles impacting on the \emph{Head} and \emph{Tail} regions of the ladder; and (iii) the dependence of the detector performance on the particle incidence angle.

physics.ins-det↗

Performance of the Gamma-ray Transient Monitor at the IHEP Electron-Beam Facility

Gamma-Ray Transient Monitor (GTM) is an all-sky monitor onboard the Distant Retrograde Orbit-A (DRO-A) satellite, with the scientific objective of detecting gamma-ray bursts in the energy range of 20 keV to 1 MeV. GTM is equipped with five Gamma-Ray Transient Probes (GTPs), utilizing NaI(Tl) scintillators coupled with silicon photomultiplier (SiPM) arrays for signal readout. To test the performance of the GTP in detecting electrons, we used the IHEP Electron-Beam Facility (a continuous-energy-tunable, low-current, quasi-single-electron accelerator) for ground-based electron tests of the GTP. This paper provides a detailed description of the operating principles of the electron accelerator and presents the process and results of the GTP electron-beam tests. The test results show that the GTP has a dead time of less than 4 $μ$s for normal signals and approximately 70 $μ$s for overflow signals, consistent with the design specifications. The time-recording capability of the GTP was tested and found to be normal, with accurate recording of overflow events. The GTP's response to electrons in the 0.4-1.4 MeV range is also normal. Additionally, we used Geant4 to simulate the GTP's energy response and performed a comparative analysis of the simulation and experimental results. The performance tests and ground-based electron calibration validated the design of the GTP and enhanced the GTP's mass model, laying the foundation for payload development, in-orbit observation strategies, and scientific data analysis.

astro-ph.IM↗

GLM-OCR Technical Report

GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems.

cs.CL↗

A New Method for Identifying Contaminating Sources and Locating Target Sources through the Cross-Arm Features of Micro Pore Optics

The Pathfinder of the Type-A satellites in the Chasing All Transients Constellation Hunters (CATCH) space mission is equipped with Micro-Pore Optics (MPOs) and four single-pixel Silicon Drift Detectors (SDDs). Due to the lack of position resolution in an individual SDD, we propose a new method based on the cross-arms in the point spread function (PSF) of MPOs to enhance the satellite's capability in identifying contaminating sources and locating target sources. By placing one detector on each of the horizontal and vertical cross-arms on the focal plane, we can use the changes in the relative counts on the cross-arms detectors to deduce the location of the source. Simulated observations demonstrate that, for a target source with a flux of 1 Crab and an exposure time of 200 s, the cross-arms detectors can identify contaminating source with the same flux level at an off-axis angle larger than 8', and improve positioning accuracy to 6'. Furthermore, we extend the simulation study to CATCH Type-A, which plans to use an SDD array. In situations where sources exhibit the same flux of 1 Crab and the exposure time is merely 1 s, a 16x16 SDD array is capable of identifying contaminating source with an off-axis angle greater than 2.4' and can achieve a positioning precision of 1.8'.

astro-ph.IM↗

LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving

Accurate metric depth is critical for autonomous driving perception and simulation, yet current approaches struggle to achieve high metric accuracy, multi-view and temporal consistency, and cross-domain generalization. To address these challenges, we present DriveMVS, a novel multi-view stereo framework that reconciles these competing objectives through two key insights: (1) Sparse but metrically accurate LiDAR observations can serve as geometric prompts to anchor depth estimation in absolute scale, and (2) deep fusion of diverse cues is essential for resolving ambiguities and enhancing robustness, while a spatio-temporal decoder ensures consistency across frames. Built upon these principles, DriveMVS embeds the LiDAR prompt in two ways: as a hard geometric prior that anchors the cost volume, and as soft feature-wise guidance fused by a triple-cue combiner. Regarding temporal consistency, DriveMVS employs a spatio-temporal decoder that jointly leverages geometric cues from the MVS cost volume and temporal context from neighboring frames. Experiments show that DriveMVS achieves state-of-the-art performance on multiple benchmarks, excelling in metric accuracy, temporal stability, and zero-shot cross-domain transfer, demonstrating its practical value for scalable, reliable autonomous driving systems.

cs.CV↗

BEV-VLM: Trajectory Planning via Unified BEV Abstraction

This paper introduces BEV-VLM, a novel approach for trajectory planning in autonomous driving that leverages Vision-Language Models (VLMs) with Bird's-Eye View (BEV) feature maps as visual input. Unlike conventional trajectory planning approaches that rely solely on raw visual data (e.g., camera images), our method utilizes a highly compressed and informative BEV representation generated by fusing camera and LiDAR data, with subsequent alignment to High-Definition (HD) maps. This unified BEV-HD map format provides a geometrically consistent and semantically rich scene description, which enables VLMs to perform accurate and robust trajectory planning. Experimental results on the nuScenes dataset demonstrate that, compared with state-of-the-art vision-only methods, our approach achieves a 53.1% improvement in planning accuracy and realizes complete collision avoidance in evaluation scenarios. Our work highlights that VLMs can effectively interpret processed visual representations such as BEV features, expanding their applicability beyond raw image inputs for the task of trajectory planning.

cs.RO↗

Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving

In this work, we reconceptualize autonomous driving as a generalized language problem and formulate the trajectory planning task as next waypoint prediction. We introduce Max-V1, a novel framework for one-stage end-to-end autonomous driving, named in tribute to the renowned Dutch racing driver Max Verstappen. Our framework presents a single-pass generation paradigm that aligns with the inherent sequentiality of driving. This approach leverages the generative capacity of the Vision-Language Model (VLM) to enable end-to-end trajectory prediction directly from front-view camera input. The efficacy of this method is underpinned by a principled supervision strategy derived from statistical modeling. This provides a well-defined learning objective, which makes the framework highly amenable to mastering complex driving policies through imitation learning from large-scale expert demonstrations. Empirically, our method achieves state-of-the-art performance on the nuScenes dataset, delivering an overall improvement of over 30% compared to prior baselines. Furthermore, it exhibits superior generalization performance on cross-domain datasets acquired from diverse vehicles, demonstrating notable potential for cross-vehicle robustness and adaptability. With these empirical strengths, this work introduces a model that enables fundamental driving behaviors, laying the foundation for the development of more capable self-driving agents. Code will be available upon publication.

cs.CV↗