SearcharxivSearch

arXiv subjects

Bo Zhang

Publications and source records attributed to Bo Zhang.

At least 19 recordsLinked to original sources

ALP-mediated inelastic dark matter and the LUX-ZEPLIN high-recoil candidate event LZ230616

The LUX-ZEPLIN (LZ) Collaboration has reported a high-energy candidate event LZ230616 with a reconstructed nuclear recoil energy $E_R=248\pm23_{\rm stat}\pm23_{\rm sys}~{\rm keV}$. We investigate a possible interpretation in terms of inelastic scattering between two Majorana dark matter states mediated by an axionlike particle coupled to gluons. The positive mass splitting suppresses low-energy recoils, while the momentum dependence of the interaction reshapes the high-energy spectrum. We treat the dark-sector and gluonic couplings independently and retain the momentum-dependent nucleon form factors and xenon nuclear responses. Using an approximate single-event likelihood, we find that, for $m_a=0.3~{\rm GeV}$, a narrow spectrum near the candidate energy arises at $m_\chi\simeq0.35~{\rm TeV}$ and $\delta\simeq330~{\rm keV}$, although this configuration requires a large coupling product and is highly sensitive to the Galactic halo speed cutoff. Our analysis establishes the kinematic and coupling requirements for subsequent tests using the thermal relic abundance and laboratory constraints on the mediator.

hep-ph

Rapid and high-sensitive NV-based microwave field imaging via digital lock-in amplification for on-chip microstrip diagnostics

High-resolution, high-sensitivity microwave (MW) magnetic field imaging is indispensable for non-destructive integrated circuit (IC) testing, radio-frequency device characterization, and spintronic research. Yet, the practical utility of these techniques is severely constrained by the pervasive challenge of isolating weak magnetic signatures from intense optical and electronic noise, which fundamentally limits both acquisition speed and detection sensitivity. Here, we overcome this barrier by introducing a wide-field imaging scheme based on an ensemble of diamond nitrogen-vacancy (NV) centers, synergistically combined with digital lock-in amplification (DLA). By exploiting digital demodulation, the DLA precisely extracts the MW-field response at a specific modulation frequency from background noise (e.g., laser intensity fluctuations), dramatically improving the signal-to-noise ratio (SNR). Consequently, our system attains a magnetic field sensitivity of 126 nT/$\sqrt(Hz)$. Critically, the unprecedented SNR permits a pixel dwell time of under one millisecond, allowing full-field images to be acquired within seconds-more than an order of magnitude faster than state-of-the-art NV-based wide-field techniques. This combination of speed, sensitivity, and micron-scale spatial resolution (1.6 $\mu$m) paves the way for quasi-real-time, non-invasive diagnostics of dynamic MW devices and integrated circuits.

physics.optics

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

cs.RO

MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon from a gradient-centric perspective. Our analysis shows that Long CoT induces larger gradient magnitudes and more concentrated update directions than Short CoT, with this effect becoming more pronounced as student model capacity increases. These findings suggest that effective Long CoT distillation requires balancing the reasoning information density of reasoning trajectories with their distributional alignment to the student model. Motivated by this insight, we propose \textbf{M}odel \textbf{I}nterporlation \textbf{Distillation} (\textbf{MI-Distillation}), a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation. To select suitable trajectories from this spectrum, we further introduce \textbf{Seq}uential \textbf{L}earnable \textbf{S}urprisal \textbf{S}core (\textbf{SeqLSS}), which favors reasoning paths that are both informative and learnable for the student. Extensive experiments on reasoning benchmarks show that MI-Distillation consistently improves small model CoT distillation over strong Long CoT baselines.

cs.CL

Towards Faithful and Efficient Semantic Communication: An Ontological Approach

In this paper, an ontology-driven semantic communication (ODSC) framework is proposed for multi-view visual question answering (VQA) tasks. In the considered framework, multiple transmitters observe a scene, extract the semantic information (SI) with vision-language models (VLMs), and transmit the scene graphs to a receiver. Due to the completeness, heterogeneity, and uninterpretability of the VLMs, the extracted scene graphs are redundant, ambiguous, and inconsistent. To solve these problems, the transmitters and the receiver share an ontology-based knowledge base that predefines synonyms, inference rules, and consistency constraints. For each transmitter, the proposed ODSC framework removes the partial scene graph that can be inferred based on the inference rules. For the receiver, the proposed framework aligns the SI of different views based on the synonyms and detects the inconsistency among the views based on the constraints. A metric of multi-view VQA accuracy (MVA) is defined to evaluate the proposed framework. Simulation results show that, compared with transmitting the complete scene graphs, the proposed framework reduces the data size of the SI by up to 87.1% while improving the answering accuracy by 4.5%. Moreover, the proposed framework yields up to a 16.0% improvement in terms of the MVA compared with the SI filtering approaches.

cs.IT

Searching for Solar-Basin Axionlike-Particle Decay with XMM-Newton Blank-Sky Observations

Axion-like particles (ALPs) bound in the solar gravitational field form the so-called ALP solar-basin. Since the two-photon decay of non-relativistic particles is approximately isotropic, this population can be searched for using observations in the anti-solar direction. In this work, we propose a search strategy for narrow decay-line signals from the ALP solar basin using \textit{XMM-Newton} blank-sky observations (XMM-BSOs) stacked spectra data taken in directions opposite to the Sun. By jointly fitting the signal and background model, we obtain limits on $g_{a\gamma\gamma}^{95}$ in the mass range $m_a=1.4\text{--}16~{\rm keV}$, with typical sensitivities of $g_{a\gamma\gamma}\sim10^{-10}\text{--}10^{-11}~{\rm GeV}^{-1}$. We have implemented the first anti-solar search for the solar basin, demonstrating that this strategy can exploit the stacked exposure of a large number of X-ray observations and provide a scalable analysis framework for future searches.

hep-ph

Giant Surface-driven Nonlinear Hall Effect in BiTeCl at Room Temperature

The nonlinear Hall effect (NLHE) provides a pathway to generate a Hall response in time-reversal-symmetric yet inversion-symmetry-broken systems. NLHE can rectify an alternating current into a transverse direct voltage, making it attractive for radio-frequency rectification, energy harvesting, and terahertz detection, applications for which device miniaturization remains a central pursuit. In this context, the inherent inversion symmetry breaking at surfaces is particularly appealing: because symmetry is necessarily broken at the surface of any crystal, irrespective of whether its bulk is centrosymmetric, surface-driven nonlinear responses lift the stringent constraint on bulk symmetry and open a route toward compact device architectures. Here we report the observation of a giant, surface-driven second-order nonlinear Hall effect in the Rashba-type polar semiconductor BiTeCl at room temperature. The determined second-order nonlinear Hall susceptibility at 300 K reaches 1.68 $\mu$mV$^{-1}$, which is 80 times larger than that of the best previously reported surface-dominated systems. We attribute this giant response to the synergistic interplay between BiTeCl's polar crystal structure and its rich surface states: the polar stacking renders the top and bottom surfaces inequivalent, so that the nonlinear response originates from a single surface without compensation from the other. Symmetry and scaling analyses suggest that both skew-scattering and side-jump mechanisms contribute to the observed effect. Our findings not only identify BiTeCl as a promising platform for future applications utilizing the NLHE, but also establish the asymmetry between the opposite surfaces of a polar crystal as a general design principle for discovering surface-driven materials with larger nonlinear Hall responses.

cond-mat.mtrl-sci

FASHI DR2: A Catalog of 132 Low-Redshift HI 21 cm Absorption Systems

We present an untargeted survey of 21 cm HI absorption systems based on the second data release of the FAST All Sky HI survey (FASHI DR2), covering approximately 19,500 deg$^{2}$ at $z\lesssim0.09$. A total of 132 HI absorbers are identified, including approximately 60 new discoveries, forming one of the largest homogeneous samples of low-redshift HI absorbers assembled to date. The sample extends to continuum flux densities as low as 2.6 mJy, substantially below the limits of previous flux-limited surveys. The absorber population is dominated by narrow systems ($W_{50}<100$ km s$^{-1}$), while broad absorbers ($W_{50}>200$ km s$^{-1}$) account for 13.6% of the sample. Most absorbers are optically thin, with a median optical depth of $\tau_{\rm HI}\approx0.14$. The velocity-offset distribution is broadly symmetric about the systemic velocities of the host galaxies. The associated absorbers are preferentially found in massive, actively star-forming galaxies. We find tentative evidence for a weak anti-correlation between HI column density and stellar mass, although the relation exhibits substantial scatter. These results provide the first statistical characterization of the low-redshift HI absorber population based on the FASHI DR2 sample and establish a valuable benchmark for future HI absorption surveys with next-generation radio facilities.

astro-ph.GA

Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis

While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks. To address this, we introduce (i) Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal-video-text tri-modal alignment. Based on this dataset, we present (ii) Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long-term medical MLLMs. Our project is available at https://github.com/ZJU4HealthCare/Holtercare-Bench.

cs.LG

SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains. We present SAGA, a generative action embedding model that encodes multi-surface user interaction sequences across a Financial Service organization's ecosystems, from checkout, peer-to-peer (P2P) transactions, in-app engagement, email to account actions, into a unified user representation for downstream recommendation tasks. Central to SAGA is a per-field tokenization schema that decomposes each action event into multiple field-level tokens (e.g. product, interaction, surface), enabling field-level attention and per-field training objectives that fused single-token approaches cannot support. Through an offline ablation study on loss formulation, tokenization granularity and training data scope, we isolate the contribution of each design choice. A downstream model integrated with SAGA-generated user embeddings delivers the strongest overall click and conversion lift across diverse downstream touchpoints, compared to all ablated and alternative architectures.

cs.LG

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based methods usually generate rubrics independently for each query. In specialized domains, such unconstrained rubrics may omit critical requirements and vary across samples, hindering the diagnosis and targeted repair of persistent capability deficiencies. We propose APTER (Adaptive Post-Training with Expert-Grounded Rubrics), a framework that integrates structured domain knowledge into fine-grained evaluation, optimization, and diagnosis for specialized complex reasoning. First, expert-grounded rubric construction starts from an expert criteria framework built by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query-level rubrics linked to their source criteria, turning reusable expert criteria into executable query-level supervision without reference answers. Second, adaptive post-training uses rubric verdicts as both optimization and criterion-level diagnostic signals. Aggregating low-scoring verdicts by criterion ID reveals persistent deficiencies and triggers targeted supervised fine-tuning updates during reinforcement learning. Experiments on mathematical reasoning and medical question answering show consistent gains across both domains. Across three model generations, APTER improves the mathematics and medical averages over the corresponding base models by up to 15.86 and 8.04 points, respectively. Code and rubric datasets are available at https://github.com/AntDT-APTER/APTER.

cs.AI

OH Line Detections in Southern Galaxies of the IRAS Revised Bright Galaxy Sample

We present a systematic study of OH main-line emission and absorption in 186 southern galaxies from the IRAS Revised Bright Galaxy Sample, using archival MeerKAT snapshot data. OH features are detected in 38 galaxies, including eight with OH maser emission (three new) and 30 showing OH absorption, mostly unreported previously. Four absorption systems exhibit weak OH emission superposed on strong absorption. OH-emitting regions are generally more compact than the associated radio continuum. Most absorption profiles are well fit by two Gaussian components (1667 and 1665 MHz), with an average integrated line ratio of $\sim$1.5. LIRGs show an OH emission detection rate of ~13\%, versus significantly lower rates in non-LIRGs. For sources with radio continuum flux densities >20 mJy, OH absorption detection rates reach ~36\% (LIRGs) and ~27\% (non-LIRGs), while no OH absorption features were detected among sources with lower radio continuum flux densities. This suggests that sufficient background continuum is likely an important factor for the detection of OH absorption. Detected OH emitters follow the empirical $L_{\rm OH}$--$L_{\rm FIR}$ relation, consistent with far-infrared pumping, while non-detections show upper limits below the relation. No significant differences are found between OH absorbers and non-detections in infrared luminosity or radio continuum compactness. Stacked spectra of non-detections reveal no significant OH features, suggesting that sensitivity and orientation alone do not fully explain the absence of absorption. In contrast, mid-infrared colors (e.g., W2--W3) and q_TIR differ between the two populations. OH absorption galaxies occupy an intermediate regime in L_HCN/L_CO between OH megamasers and non-detections, implying that OH absorption detectability is linked to dense molecular gas conditions, with extreme star formation potentially suppressing its occurrence.

astro-ph.GA

Towards Physics-Faithful Generation of Scientific Diagrams

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

cs.CV

A Helium-shell Burning Blue Horizontal Branch Star Produced from Common Envelope Evolution

Observationally, blue horizontal branch (BHB) stars are defined as hot stars occupying a characteristic region between the extreme blue horizontal branch and RR Lyrae variables in the Hertzsprung-Russell diagram. Most of them are interpreted as stripped core-helium-burning stars, but the role of binary interaction in their formation remains unclear. Here, we report the discovery of a metal-rich BHB star in a 0.82628-day binary system (Feige 64) comprising a $0.35\pm0.03\,M_{\odot}$ BHB star and a likely $1.26\pm0.17\,M_{\odot}$ white dwarf (WD). The BHB star has an effective temperature of $15{,}524\pm310$ K and a luminosity of $39.7\pm4.1\,L_{\odot}$. Stellar evolution modelling indicates that it is a helium-shell-burning star produced through the common-envelope channel, retaining a hydrogen-rich envelope that is more massive than previously thought for low-mass stars. This finding provides direct evidence for binary interaction in the formation of BHB stars, offering a fresh perspective on interpreting this emerging population.

astro-ph.SR

PE-CSNet: An equivariant network architecture with learnable patch-based sparse representation

Compressive sensing (CS) enables accurate signal reconstruction from sparse measurements and is widely applied in medical imaging, remote sensing, and image compression. However, designing an effective, task-specific sparse transform and the corresponding optimization procedure for high-quality CS remains challenging. This process typically requires expert domain knowledge and laborious parameter tuning. To address this issue, we present a Patch-based Equivariant deep unrolling architecture, termed PE-CSNet, for accurate CS recovery. While traditional CS methods generally use predefined patch-based transform sparsity, we generalize this idea by incorporating learnable transform sparsity that adapts to the specific CS task through an optimization-driven process. Specifically, we first establish a generalized patch-based CS model, which we solve via a block coordinate descent (BCD) algorithm. The BCD solver is then unrolled into a deep neural network, where all parameters of both the CS model and solver are learned through end-to-end training. To improve data efficiency, we introduce a stochastic equivariant training strategy that exploits the patch-wise structure of the network, enabling PE-CSNet to learn effectively even from limited data. We further provide a simpler, parameter-shared version of PE-CSNet and briefly discuss its convergence as an iterative solver. For practical applications, the network uses stage-specific (non-shared) parameters to enhance its expressive power and thereby improve its performance. On the tasks of CS magnetic resonance imaging (CS-MRI) and CS coded diffraction patterns (CS-CDP), PE-CSNet achieves state-of-the-art accuracy with fast computational speed, outperforming traditional methods and existing deep unrolling methods.

cs.CV

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM's intrinsic evidence and failing to accommodate the non-uniform spatiotemporal information density. In this paper, we propose a fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence. Our method efficiently probes visual evidence via sparse prefilling as a structured prior to guide distribution-aware dynamic sampling. Specifically, we efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart. Conditioned on three complementary attention components derived from this prior, we design a lightweight selector that not only precisely locates query-relevant timestamps but also adaptively adjusts the local sampling rate and spatial resolution. To enable evidence-conditioned spatiotemporal sampling, we formulate the selector as a stochastic policy and optimize it via GRPO under a joint accuracy--efficiency reward. By rewarding correct predictions under lower visual cost through group-relative comparisons, our method encourages the policy to allocate computation dynamically according to the information density of each video. Across three long video understanding benchmarks, EviSelect achieves superior performance compared to existing methods while reducing selected visual tokens by about 50\% and achieving a 3.9x end-to-end speedup.

cs.CV

Improving the efficiency of infectious disease prevention trials using negative control outcome event times

Baseline covariate adjustment can enhance the efficiency of randomized trials by improving precision of treatment effect estimates. However, the precision gain depends on how strongly the baseline covariates are prognostic for the primary outcome. In randomized trials of infectious disease prevention interventions (e.g., vaccines or passively administered antibodies), an individual's exposure to the pathogen is a leading prognostic factor but is rarely measurable at baseline. Hence, conventional covariate adjustment offers limited precision gain in prevention trials. We propose adjusting for a negative control outcome (NCO) event time, which is causally unaffected by the intervention but shares overlapping exposure mechanisms with the primary outcome. We formalize assumptions under which adjustment for the NCO event time is valid, and show that right-censoring of the NCO event time further complicates adjustment. We derive the efficient influence function for the treatment-arm-specific survivor function of the primary outcome when both the primary outcome and the NCO event time are right-censored, and use it to construct a cross-fitted, one-step estimator that is multiply robust to nuisance misspecification and asymptotically efficient when the nuisances are estimated accurately. In numerical experiments, our estimator compares comparably to benchmarks when the NCO event time is uninformative, and gains precision as the NCO event time is more prognostic for the primary outcome. We apply our method to HVTN 704/HPTN 085, a randomized, double-blinded trial of VRC01, a broadly neutralizing antibody against HIV-1. Adjusting for the time to a bacterial sexually transmitted infection --- a negative control outcome for HIV-1 acquisition --- reduced the estimated variance of the prevention efficacy estimate by approximately 27%, compared to roughly 2.5% for baseline covariate adjustment.

stat.ME

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a single large language model, yet these approaches often expose them to only a limited range of perspectives and directions. We present AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration. Heterogeneous agents asynchronously discuss scientific questions in a forum-style environment, while researchers can submit questions, browse and organize candidate ideas, engage agents in follow-up interactions, and optionally generate post-hoc summary reports. We evaluate AgentPanel in terms of idea quality, exploration breadth, interaction effectiveness, candidate-selection efficiency, and practical utility. Offline experiments show that AgentPanel outperforms a centralized multi-agent debate baseline. A human study with 20 participants further shows that users value AgentPanel for perspective diversity and exploration support. In experience-based comparisons with commonly used LLM tools, 65\% of participants favored AgentPanel for both breadth of research directions and overall suitability for early-stage exploration. The platform is publicly available at https://agentpanel.cc/.

cs.AI