Searcharxiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 37 records · Page 2Linked to original sources

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' engagement dynamics. Existing methods often fail to capture these complex temporal patterns, leading to inaccurate long-term forecasts. This failure stems from two fundamental limitations: \ding{172} a superficial understanding of user feedback dynamics, which overlooks the mutually exciting and decaying nature of interactions such as likes, comments, and shares; and~\ding{173} retrieval mechanisms that rely solely on static content similarity, ignoring the crucial patterns of how a video's popularity evolves over time. To address these limitations, we propose \textbf{M$^3$TR}, a \textbf{T}emporal \textbf{R}etrieval enhanced \textbf{M}ulti-\textbf{M}odal framework that uniquely synergizes fine-grained temporal modeling with a novel temporal-aware retrieval process for \textbf{M}icro-video popularity prediction. At its core, M$^3$TR introduces a Mamba-Hawkes Process (MHP) module to explicitly model user feedback as a sequence of self-exciting events, capturing the intricate, long-range dependencies within user interactions (for \textbf{limitation} \ding{172}). This rich temporal representation then powers a temporal-aware retrieval engine that identifies historically relevant videos based on a combined similarity of both their multi-modal content (visual, audio, text) and their popularity trajectories (for \textbf{limitation} \ding{173}). By augmenting the target video's features with this retrieved knowledge, M$^3$TR achieves a comprehensive understanding of prediction. Extensive experiments on two real-world datasets demonstrate the superiority of our framework. M$^3$TR achieves state-of-the-art performance, outperforming previous methods by up to \textbf{19.3}\% in nMSE and showing significant gains in addressing long-term prediction challenges.

cs.MM↗

Focal-point scanning for dose delivery and optimization with focused laser-accelerated very-high-energy electron beams

Focused very-high-energy electron (VHEE) beams can produce localized dose enhancement at selected depths, but irradiation of a finite target requires coordinated control of multiple focal positions, incidence directions, and beam weights while limiting exposure of nearby organs at risk (OARs). We present Focal-Point Scanning (FPS), a dose delivery and optimization method developed for laser wakefield accelerator (LWFA)-driven VHEE beams. The method is based on a two-dipole focusing system that produces single-plane beam convergence and allows the focal position to be varied by changing the magnetic field strength. FPS distributes focal points throughout the planning target volume and determines focal-point-specific incidence sectors according to the geometry of nearby critical OARs. The method was evaluated using the AAPM TG119 C-shape benchmark and one previously treated lung radiotherapy case. At matched target coverage, FPS reduced the TG119 Core mean dose by approximately one half relative to parallel VHEE and intensity-modulated x-ray plans, approaching the single-field proton pencil-beam-scanning reference. In the lung case, FPS maintained target coverage comparable to the clinical volumetric modulated arc therapy reference while reducing the mean dose to every evaluated OAR; spinal-cord mean and maximum doses decreased by 93.2% and 87.2%, respectively. The evaluated OAR mean doses varied little across rms energy spreads of 0 to 10% and for a flat-top electron spectrum spanning 150 to 250 MeV. These results demonstrate that focal-point-specific angular selection can translate focused-beam physics into effective OAR sparing and support FPS as a planning strategy for broadband LWFA-VHEE radiotherapy.

physics.med-ph↗

ALMA visits the QSO MUSEUM: Connecting molecular gas and the cool circumgalactic medium around 37 z~3 quasars

Extended Ly$α$ emission is ubiquitous around quasars and traces the cool circumgalactic medium, providing insights into halo gas dynamics and active galactic nucleus feedback. However, its connection to the cold molecular gas of the host galaxies remains largely unexplored. We characterize the molecular gas reservoirs of quasars at cosmic noon and investigate their connection to extended Ly$α$ emission using ALMA CO(4-3) observations of 37 quasars at $z\sim3$ from the QSO MUSEUM survey, previously mapped in Ly$α$ with VLT/MUSE. We derive molecular gas masses and gas fractions, explore correlations with Ly$α$ nebula and quasar properties, and search for CO-emitting companions. We detect 21/37 quasars in CO(4-3), with gas masses of $M_\mathrm{gas}\approx(3-40) \times10^9\,\mathrm{M_\odot}$. Quasars with the most massive molecular gas reservoirs are associated with the centrally dimmest Ly$α$ nebulae, while those hosting the centrally brightest Ly$α$ nebulae are generally not detected in CO. This suggests that gas and dust in the hosts regulate Ly$α$ escape and consequently affect the emission from halo gas. We find evidence that lower-Eddington-ratio quasars harbor more massive gas reservoirs, while strongly accreting quasars ($λ_\mathrm{Edd} \gtrapprox 0.9$) likely deplete their gas; for example, through powerful quasar-driven outflows. Despite their higher molecular gas masses within the sample, CO-detected low-Eddington quasars exhibit low gas fractions with a median $M_\mathrm{gas}/M_* \sim 0.10$, below what is typically found for inactive star-forming galaxies. Six quasars are marginally resolved in CO, with effective radii up to $\sim 8\,\mathrm{kpc}$. In addition, we detect 14 high-fidelity companion galaxies, indicating overdense quasar environments with a quasar-galaxy cross-correlation length of $9.81^{+2.22}_{-2.05}\,h^{-1}\mathrm{cMpc}$.

astro-ph.GA↗

Harmonic Ranking for Edge-Weighted Oblivious Matching

We study edge-weighted oblivious bipartite matching. The weight of every potential edge is known, but its existence is revealed only when the edge is probed, and a successful probe between two free vertices must be accepted immediately. We give an explicit randomized algorithm with certified competitive ratio $0.698$, improving the previous best guarantee of $0.659$ (Huang, Sun, Wu, and Zhao, FOCS 2025). The result is computer-assisted and verified by a reproducible exact-integer computation. The same algorithm has a $0.698$-competitive online implementation for the vertex-weighted random-arrival model, improving the previous $0.696$ unweighted guarantee of Mahdian and Yan (STOC 2011) and the $0.686$ vertex-weighted guarantee of Peng and Tang (EC 2025). Our algorithm, Harmonic Ranking, is a role-symmetric generalization of \textsc{Ranking}. It assigns an independent random rank $x_z$ to each vertex and probes a potential edge $uv$ in decreasing order of \[ w_{uv}\frac{h(x_u)h(x_v)}{h(x_u)+h(x_v)}. \] This harmonic priority arises from a budget-balanced gain split and a mutual-proposal interpretation. The analysis lifts two cutoff curves into indicators, reducing the exponential-size factor-revealing problem to a polynomial-size directed minimum-cut instance. A maximum-flow computation with rounded-down integer capacities gives a rigorous certificate. Independently, we observe that the finite-grid unweighted relaxation of our factor-revealing program coincides exactly with a Mahdian--Yan program.

cs.DS↗

Breaking the 4-Approximation Barrier in Strategyproof Two-Facility Location

We study strategyproof mechanism design without transfers for the two-facility location problem in metric spaces. A mechanism selects two facility locations based on agents' reported locations; each agent incurs her distance to the nearer facility, and the objective is to minimize the expected social cost. A mechanism is strategyproof if no agent ever benefits from misreporting her location. The best approximation ratio achieved by a randomized strategyproof mechanism has been $4$, attained by the Proportional mechanism of Lu, Sun, Wang, and Zhu (EC 2010), and the best lower bound has been $1.045$, due to Lu, Wang, and Zhou (WINE 2009). Neither bound has moved since then, even on the line $\mathbb{R}$. We improve both bounds. Our main result is a randomized strategyproof mechanism with approximation ratio $11/3 \approx 3.667$ on every Ptolemaic metric space, a rich class containing all Euclidean spaces. The mechanism randomizes between the Proportional mechanism and a new mechanism that we call Global Pair. Global Pair draws an unordered pair of agents with probability proportional to their distance and opens facilities at their reported locations. Although Global Pair and Proportional each have approximation ratio $4$, the two mechanisms attain their worst-case approximation ratios on complementary instances. Randomizing between them balances these complementary weaknesses and breaks the $4$-approximation barrier. On the lower-bound side, we construct a new two-profile instance that yields a lower bound of $(1+\sqrt{2})/2 \approx 1.207$, improving upon the previous lower bound of $1.045$.

cs.GT↗

Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation

RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in gaming industry. However, generating high-quality RGBA animations for games remains challenging for two reasons. First, most existing RGBA video datasets are dominated by photorealistic content, with limited coverage of game assets. Second, the traditional generate-then-matte pipelines estimate alpha only after RGB synthesis, so semi-transparent regions are often blurred by background, resulting in unstable matting outputs. More recently, many methods have begun to model RGB and alpha jointly, but existing approaches are mostly text-conditioned, and still have unresolved issues in efficiency and quality. To address these challenges, we introduce GameAlpha-2.4K, a 2.4K-clip game-style RGBA video dataset built with matte-friendly synthesis, multi-hypothesis alpha recovery, and compositing-based quality gates. Using this dataset, we train a reference-conditioned RGBA video generator that jointly produces RGB frames and alpha mattes in a single pass. To improve efficiency, we propose a visibility router that identifies transparent tokens in an early stage and bypasses their later DiT updates, while x_0-lock guides them along the original flow-matching schedule toward self-predicted endpoints. Our model obtains lower FVD than traditional two-stage pipelines, and the visibility router skips 35% of token evaluations in the final two DiT denoising steps, providing a 1.2x backbone speedup with negligible quality degradation compared to dense inference.

cs.CV↗

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated objects receive equally strong visual attention in the model's mid-to-late layers, suggesting that the key issue may not be how much the model attends, but what it attends to and why. To this end, we decode the visual features of high-attention regions using Logit Lens, and observe that regions corresponding to real objects can be correctly decoded to the target object tokens, whereas those for hallucinated objects cannot. Building on this, we identify two hallucination mechanisms: (i) visual uncertainty, triggered by semantically similar or confusable regions; masking these regions eliminates the hallucination. (ii) contextual prior, triggered by strong co-occurrence priors; even when the initially attended region is masked, the hallucination persists and attention drifts to other regions. Based on these findings, we propose a simple yet effective training-free Detect-Mitigate framework comprising a Logit-Lens Consistency Check to detect hallucination and targeted remedies: High-Attention Regions Masking (HARM) for visual uncertainty hallucination, and Visual Evidence Enhanced Decoding (VEED) for contextual prior hallucination. Our approach achieves state-of-the-art results on multiple hallucination benchmarks. Code will be available.

cs.CV↗

FAST Ultra-Deep Survey: the baryonic Tully-Fisher relation in FUDS0 field

The Baryonic Tully-Fisher relation (BTFR) is one of the tightest scaling relations for disk galaxies in the local Universe, and therefore is an important tool for studying the fomation and evolution of galaxies. However, the evolution of the BTFR over cosmic time is poorly understood due to the limited sample of HI galaxies beyond the local Universe, limitations of optically-derived rotation curves, and selection effects. In this work, we explore the BTFR at redshifts up to $z=0.42$ from galaxies detected in the pilot FAST Ultra-Deep Survey (FUDS) field, FUDS0. As found in previous work, we identify two components in the plane of baryonic mass versus rotational velocity, $C_{\rm BTFR}$ (tight) and $C_{\rm Outlier}$ (dispersed). A Gaussian mixture model is employed to recover the BTFR, yielding the best fit parameters for the slope $k=3.32_{-0.11}^{+0.12}$, zero point $b=10.07_{-0.03}^{+0.03}$, and intrinsic scatter $σ_{\rm BTFR}=0.036_{-0.009}^{+0.010}$. A random forest classifier is used to investigate the origin of the outlier component. We find that low signal significance and inaccurate inclinations are the key factors that contribute to the outlier population, indicating that observational effects are the dominant origin. Evolutionary trends are examined in three different redshift bins. Both the slope and zero point show consistency within 1-$σ$ uncertainty in the two low redshift bins, indicating no significant evolution. The indirectly inferred BTFR parameters from the $C_{\rm Outlier}$ component in the highest redshift bin aligns with the conclusion. The ongoing full FUDS survey will provide a larger sample to enable more accurate constraints on BTFR evolution.

astro-ph.GA↗

Structured High-Angular-Momentum Coulomb Tensors from Real and Complex Solid-Harmonic Integral Engines: A Perspective

Electron-repulsion integrals describe the Coulomb interaction between charge distributions built from orbital basis functions. Most integral algorithms generate these quantities through Cartesian Gaussian functions, whose angular shapes are written as powers of $x$, $y$, and $z$, and then transform the result to spherical functions. This route is effective, but from $d$ shells onward the Cartesian representation contains more functions than the spherical space required by the calculation. Direct real or complex solid-harmonic engines work in that target space from the beginning. They therefore produce a smaller final Coulomb tensor while preserving the ordering, phase, and magnetic-quantum-number labels that describe its angular structure. Following this structure beyond integral evaluation reveals direct connections to the algorithms that use the tensor. Simple analytical counts quantify tensor size, angular blocks, radial Slater--Condon parameters, and pair-space work. These quantities guide low-rank factorization, local Hamiltonian construction, quantum simulation, and transformations to spinor or effective-model bases. In this way, solid-harmonic integral engines provide a direct bridge between efficient integral generation and structured many-electron computation.

physics.chem-ph↗

Random-Order Online Facility Location Beyond Uniform Opening Costs

We study online metric facility location in the random-order model with arbitrary positive opening costs. A finite set of candidate facilities and their costs is known in advance, while an adversary fixes a multiset of demand points that arrives in a uniformly random order. This setting includes both prescribed candidate sites and the classical finite full-space node-cost model. For a known horizon, we give a deterministic $4.2674$-competitive algorithm, improving the previous factor $33$ for nonuniform opening costs. At rank $t$, the algorithm uses the positive normalized rank $q_t=t/n$, chooses a candidate minimizing $d(x,y)+λ_t f_y$, where $λ_t=\min\{1,q_t/μ\}$, and opens it when the current connection distance covers this penalized objective. The analysis uses a monotone one-round charge and an upper-envelope decomposition to control later points and the first point of each optimal cluster. With unit opening costs, the rule reduces exactly to a cutoff on the distance improvement attainable from a nearest candidate. A supplementary appendix gives the sharper analysis of the closely related zero-start rank cutoff and obtains a ratio below $3.2805$. We also prove a $3-o(1)$ lower bound for arbitrary randomized online algorithms. The lower bound already holds with uniform costs on a prescribed candidate set and transfers, without loss, to the finite full-space model with nonuniform opening costs. Together with the recent competitive ratio below $2.42$ for full-space uniform costs, this yields a strict separation between the full-space uniform- and nonuniform-cost models.

cs.DS↗

Absolute charge calibration of DRZ phosphor screens for relativistic electron bunches

Laser-plasma accelerators have been the subject of extensive research in recent years. The electron beams they generate exhibit a broad energy spread. To conveniently characterize beams from laser wakefield acceleration (LWFA), electron spectrometers employing scintillating screens coupled with CCD cameras are typically used. In this work, we calibrate a series of DRZ phosphor screens and measure the spectra of the light they emit. The calibration was performed using the radio-frequency linear electron accelerator at Tsinghua University, which provided monoenergetic electron beams with peak energy of approximately 30 MeV.

physics.acc-ph↗

Unified Uncertainty Quantification Framework Bridging Noisy Quantum Backends Across Variational Quantum Algorithms and Quantum Signal Processing

We present an uncertainty quantification (UQ) framework for application level benchmarking and characterization of noisy quantum backends. The framework compares two workload classes under one statistical pipeline: noisy intermediate scale quantum (NISQ) variational quantum algorithms (VQAs) and Quantum Singular Value Transformation (QSVT) based Green's function reconstruction. For the VQA branch, we evaluate ten benchmark families spanning chemistry, optimization, simulation, compiling, linear solving, partial differential equations, metrology, error correction, tomography, and channel fidelity estimation. For the QSVT branch, we reconstruct orbital resolved Green's functions and spectral peaks from a block encoded real time propagator. The workflow combines Bayesian optimization, posterior distribution refinement, sensitivity analysis, robust parameter density estimation, backend ranking, noise correlation, and resource estimation analysis. Instead of reporting only one best parameter vector, the framework identifies robust parameter regions, residual gaps to ideal behavior, backend specific failure modes, and calibration sensitive uncertainty. The result is a common benchmark for variational and non-variational workloads that measures how reliably each backend reaches useful task level behavior.

cs.ET↗

The Stack Search Tests on FAST Data: Discovery of Six Faint Isolated Millisecond Pulsars in NGC 6517 and NGC 7078 (M15)

We report the discovery of six faint millisecond pulsars (MSPs) in the globular clusters NGC 6517 and NGC 7078 (M15) using the Five-hundred-meter Aperture Spherical radio Telescope (FAST). These discoveries were enabled by stacking power spectra from multiple observations, a method that effectively boosts the signal-to-noise ratio of faint sources. In NGC 6517, we identified four new MSPs (NGC 6517S-V) with spin periods ranging from 3.68 to 6.02 ms and dispersion measures (DMs) between 182.45 and 182.85 pc cm^-3. In M15, two additional MSPs (M15M and M15N) were discovered, with spin periods of 4.83 and 9.28 ms, and DMs of 67.89 and 66.65 pc cm^-3, respectively. A phase-coherent timing solution has been obtained for M15M; however, sparse detection rates currently preclude phase-connected solutions for the remaining five pulsars. Current timing parameters suggest all six MSPs are isolated, which is consistent with the expected pulsar populations in core-collapsed globular clusters. Notably, pulsars M15N, NGC 6517U, and NGC 6517V eluded detection by standard frequency-domain searches (e.g., PRESTO-based) and the Fast Folding Algorithm, demonstrating that the stack search technique significantly enhances detection sensitivity to inherently faint pulsar signals.

astro-ph.HE↗

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank from the recent trajectory and deciding whether to inject a memory-grounded reminder or remain silent. The module is plug-and-play with frontier action agents and existing agent harnesses. Across Terminal-Bench 2.0 and $τ^2$-Bench, it improves pass@1 for both weaker and stronger action agents, with gains of +8.3 pp on Terminal-Bench and +6.8 pp on $τ^2$-Bench. Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval. As an early step toward open-weight memory policies, we train Qwen3.5-27B on SETA using SFT and GRPO, improving validation reward and achieving partial transfer to Terminal-Bench.

cs.AI↗

Measurement-Access Risk Frontiers for Autonomous Scientific Control

Rapidly scaling autonomous science is limited not only by algorithms, compute or data volume, but by which physical records a platform exposes before action. We formulate physically accessible decision-making (PADM) and a measurement-access risk frontier: the Bayes-optimal target risk minimized over records realizable under cost, bandwidth, latency, disturbance, memory and actuation constraints. The frontier gives a no-free-autonomy limit: automation cannot collapse decision uncertainty by computation alone; an optimal controller cannot remove target components absent from its record, and closing that gap requires expanded access, auditing, tolerated disturbance, slower operation or restricted deployment. In monitored feedback, displacement-only control remains exposed to a hidden switching force, whereas a finite-bandwidth cue recovers part of the missing projection before action. A chemistry-aware candidate-ranking audit with a 1000-target stress panel, Gaussian sensing, hidden-regime decisions and cost-aware/thermodynamic channel selection provide reproducible checks. PADM identifies target-specific audit value and residual oracle gaps before deployment.

math-ph↗

CauScale: Neural Causal Discovery at Scale

Causal discovery is essential for advancing data-driven fields such as scientific AI and data analysis, yet existing approaches face significant time- and space-efficiency bottlenecks when scaling to large graphs. To address this challenge, we present CauScale, a neural architecture designed for efficient causal discovery that scales inference to graphs with up to 1000 nodes. CauScale improves time efficiency via a reduction unit that compresses data embeddings and improves space efficiency by adopting tied attention weights to avoid maintaining axis-specific attention maps. To keep high causal discovery accuracy, CauScale adopts a two-stream design: a data stream extracts relational evidence from high-dimensional observations, while a graph stream integrates statistical graph priors and preserves key structural signals. CauScale successfully scales to 500-node graphs during training, where prior work fails due to space limitations. Across testing data with varying graph scales and causal mechanisms, CauScale achieves 99.6% mAP on in-distribution data and 84.4% on out-of-distribution data, while delivering 4-13,000 times inference speedups over prior methods. Our project page is at https://github.com/OpenCausaLab/CauScale.

cs.LG↗

Coupled atmospHere Interior modeL Intercomparison (CHILI). I. Evolutionary Modelling -- Primordial Magma Oceans of Earth and Venus

Earth and Venus represent two evolutionary outcomes arising from initially molten 'magma ocean' periods, followed by lifetimes of chemical and geophysical divergence. Their physics is common to all rocky planets and is accessible to simulations that adopt coupled interior-atmosphere modelling approaches. Our understanding of planet histories and interpretation of current states is dependent on this modelling, yet existing codes vary in their approximations. Here, we present the first results from the Coupled atmospHere Interior modeL Intercomparison (CHILI) project; benchmarking planetary evolution codes in the context of Earth and Venus to identify key model sensitivities. Our 'nominal' Earth models predict magma ocean solidification timescales within 4 Myr of thermal evolution, and are consistent with empirical constraints on Earth's early history. Venus scenarios exhibit more diverse behaviours where prolonged magma ocean stages can be conditionally sustained for 50 Myr. Cooling timescales correlate with initial hydrogen and carbon budgets, but model-specific treatments of volatile partitioning and vertical energy transport introduce substantial inter-model variance. Different parametrisations of mantle geodynamics, convection, melting curves, rheological properties, and radiative transfer give rise to divergent evolutionary behaviours. Discrepancies in atmospheres generated by magma ocean outgassing underscore these differences, although C-H-O compositions with surface pressures exceeding 100 bar are favoured. This intercomparison identifies critical sensitivities in volatile partitioning, escape processes, mantle viscosity, and melting. Validating these treatments is essential for enabling deep insight into the early histories of the Solar System's terrestrial planets, and for drawing meaningful interpretations from ongoing observational exoplanet campaigns.

astro-ph.EP↗

High-Fidelity 4D Hand-Object Capture via Multi-View Spatiotemporal Tracking and Physics-Aware Gaussians

The growing demand for high-fidelity 4D hand-object interaction (HOI) data in embodied AI and spatial computing is currently bottlenecked by the reliance on pre-scanned object templates and physical markers. While recent methods have demonstrated promising results in reconstructing 4D hand-object interaction from videos, they are highly sensitive to initial estimates of hand and object poses. Yet, estimating these poses from images is challenging, in particular under severe occlusion which is inherent in hand-object interaction scenarios. We propose a novel system for the robust and accurate reconstruction of hands and objects from synchronized and calibrated multi-view videos without requiring any templates or markers. Our system consists of two main components with key innovations: (1) a multi-view feed-forward transformer model that aggregates cross-view geometry and temporal cues to provide a reliable, metric-consistent initialization for both poses and dense object geometry, and (2) a hand-object physics-aware Gaussian-based optimization framework to refine the initial estimates, integrating tetrahedral constraints, collision refinement, and appearance decomposition to produce physically plausible and visually accurate reconstruction. Validated on public benchmarks and an extensive internal dataset, our pipeline achieves highly robust, artifact-free reconstruction, providing an efficient foundation for automated 4D asset generation. Our project page are available at https://zyshen021.github.io/HOSTPG/.

cs.CV↗