SearcharxivSearch

arXiv subjects

Junjie Zhao

Publications and source records attributed to Junjie Zhao.

At least 19 recordsLinked to original sources

Improved proper motion and gravity tests with PSR J1913+1102

PSR J1913+1102 is a highly asymmetric double neutron star system and an excellent laboratory for testing scalar-tensor gravity theories, as well as a potential progenitor analogue of GW170817 that will merge in 470 Myr. We present an updated timing analysis combining 13 years of historical Arecibo observations and new FAST measurements, using two approaches to model dispersion-measure variations. The new timing solution provides precise measurements of four post-Keplerian parameters and improves the system mass estimates. Assuming general relativity and modelling the DM variation with a Gaussian process, we obtain a three-fold improvement in the total mass, m_{tot}=2.88948(20) M_\odot, and nearly four-fold improvements in the pulsar and companion masses, m_p=1.599(8) M_\odot and m_c=1.290(8) M_\odot, giving the mass ratio, q=0.807(8). We also measure an improved proper motion, \mu=7.71(25) mas yr^{-1}, enabling a more accurate correction of the observed orbital-period derivative. Combined with the improved orbital-decay measurement, this yields an intrinsic orbital-period derivative \dot{P}_b^{intr}=-4.60(6)\times10^{-13} s s^{-1}, five times more precise than the previous value and fully consistent with the general-relativistic prediction for gravitational-wave damping. The improved masses and precise \dot{P}*b^{intr} place stringent constraints on dipolar gravitational-wave emission and the spontaneous-scalarisation window around 1.6 M*\odot. The refined proper motion and mass measurements also provide tighter constraints on the final helium-star mass immediately prior to its core collapse and formation of the second NS in a supernova, as well as on the magnitude and direction of the associated natal kick of the DNS system.

astro-ph.HE

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

cs.AI

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, dense matter, gravitation, and cosmology. The GW sources detected to date are from mergers of compact object binaries, which possess the potential to generate detectable electromagnetic (EM) counterparts. Searching for associations between GW signals and their EM counterparts is an essential step toward enabling subsequent multi-messenger studies. In the era of next-generation GW and EM detectors, the rapid increase in the number of events brings not only unprecedented scientific opportunities, but also substantial challenges to the existing data analysis paradigm. To help address these challenges, we develop GW-Eyes, an agentic framework powered by large language models (LLMs). For the first time, GW-Eyes integrates domain-specific tools and autonomously performs counterpart association tasks between GW and candidate EM events. It supports natural language interaction to assist human experts with auxiliary tasks such as catalog management, skymap visualization, and rapid verification. Our framework leverages the complex decision-making capabilities of LLMs and their traceable reasoning processes, offering a new perspective to the multi-messenger astronomy.

astro-ph.IM

Null-Space Flow Matching for MIMO Channel Estimation in Latency-Constrained Systems

Accurate yet low-latency channel state information (CSI) acquisition is essential for multiple-input multiple-output (MIMO) communication systems. While advanced deep generative models, such as score-based and diffusion models, enable high-fidelity CSI reconstruction from limited pilot observations, they often suffer from high inference latency. To achieve accurate CSI estimation under stringent latency constraints, this paper proposes a null-space flow matching (FM) framework that leverages a range-null space decomposition to separate observation-informed and underdetermined channel components. Specifically, the pilot observations are used to regulate the observable range-space channel component, while an FM-based generative prior primarily resolves the ambiguous null-space degrees of freedom through iterative refinement. To further improve the robustness and efficiency of the proposed framework, we introduce a noise-aware adaptive correction strategy to suppress channel noise on the refinement trajectory, along with a power-law time schedule to better allocate the limited number of refinement steps. Experimental results demonstrate that our method achieves competitive normalized mean square error (NMSE) performance even under a strict latency budget of around 3 ms, while delivering a superior accuracy-latency tradeoff compared with both model-based and generative baselines.

cs.IT

HAL: Inducing Human-likeness in LLMs with Alignment

Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised training, rather than targeted alignment. We introduce Human Aligning LLMs (HAL), a framework for aligning language models to conversational human-likeness using an interpretable, data-driven reward. HAL derives explicit conversational traits from contrastive dialogue data, combines them into a compact scalar score, and uses this score as a transparent reward signal for alignment with standard preference optimization methods. Using this approach, we align models of varying sizes without affecting their overall performance. In large-scale Chatbot Arena-style human evaluations, a model aligned with HAL is more frequently perceived as human-like in conversation. Because HAL operates over explicit, interpretable traits, it enables inspection of alignment behavior and diagnosis of unintended effects. More broadly, HAL demonstrates how soft, qualitative properties of language--previously outside the scope for alignment--can be made measurable and aligned in an interpretable and explainable way.

cs.AI

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack sufficient reasoning ability for precise diagnosis. To address these limitations, we present DentalGPT, a specialized dental MLLM developed through high-quality domain knowledge injection and reinforcement learning. Specifically, the largest annotated multimodal dataset for dentistry to date was constructed by aggregating over 120k dental images paired with detailed descriptions that highlight diagnostically relevant visual features, making it the multimodal dataset with the most extensive collection of dental images to date. Training on this dataset significantly enhances the MLLM's visual understanding of dental conditions, while the subsequent reinforcement learning stage further strengthens its capability for multimodal complex reasoning. Comprehensive evaluations on intraoral and panoramic benchmarks, along with dental subsets of medical VQA benchmarks, show that DentalGPT achieves superior performance in disease classification and dental VQA tasks, outperforming many state-of-the-art MLLMs despite having only 7B parameters. These results demonstrate that high-quality dental data combined with staged adaptation provides an effective pathway for building capable and domain-specialized dental MLLMs.

cs.CV

DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle

Real-world enterprise data intelligence workflows encompass data engineering that turns raw sources into analytical-ready tables and data analysis that convert those tables into decision-oriented insights. We introduce DAComp, a benchmark of 210 tasks that mirrors these complex workflows. Data engineering (DE) tasks require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data analysis (DA) tasks pose open-ended business problems that demand strategic planning, exploratory analysis through iterative coding, interpretation of intermediate results, and the synthesis of actionable recommendations. Engineering tasks are scored through execution-based, multi-metric evaluation. Open-ended tasks are assessed by a reliable, experimentally validated LLM-judge, which is guided by hierarchical, meticulously crafted rubrics. Our experiments reveal that even state-of-the-art agents falter on DAComp. Performance on DE tasks is particularly low, with success rates under 20%, exposing a critical bottleneck in holistic pipeline orchestration, not merely code generation. Scores on DA tasks also average below 40%, highlighting profound deficiencies in open-ended reasoning and demonstrating that engineering and analysis are distinct capabilities. By clearly diagnosing these limitations, DAComp provides a rigorous and realistic testbed to drive the development of truly capable autonomous data agents for enterprise settings. Our data and code are available at https://da-comp.github.io

cs.CL

Revealing the Temporally Stable Bimodal Energy Distribution of FRB 20121102A with a Tripled Burst Set from AI Detections

Active repeating Fast Radio Bursts (FRBs), with their large number of bursts, burst energy distribution, and their potential energy evolution, offer critical insights into the FRBs emission mechanisms. Traditional pipelines search for bursts through conducting dedispersion trials and looking for signals above certain fluence thresholds, both of which could result in missing weak and narrow-band bursts. In order to improve the completeness of the burst set, we develop an End-to-end DedispersE-agnostic Nonparametric AI model (EDEN), which directly detect bursts from dynamic spectrum and is the first detection pipeline that operates without attempting dedispersion. We apply EDEN to archival FAST L-band observations during the extreme active phase of the repeating source FRB 20121102A, resulting in the largest burst set for any FRB to date, which contains 5,927 individual bursts, tripling the original burst set. The much enhanced completeness enables a refined analysis of the temporal behavior of energy distribution, revealing that the bimodal energy distribution remains stable over time. It is rather an intrinsic feature of the emission mechanisms than a consequence of co-evolving with burst rate.

astro-ph.HE

TransMPC: Transformer-based Explicit MPC with Variable Prediction Horizon

Traditional online Model Predictive Control (MPC) methods often suffer from excessive computational complexity, limiting their practical deployment. Explicit MPC mitigates online computational load by pre-computing control policies offline; however, existing explicit MPC methods typically rely on simplified system dynamics and cost functions, restricting their accuracy for complex systems. This paper proposes TransMPC, a novel Transformer-based explicit MPC algorithm capable of generating highly accurate control sequences in real-time for complex dynamic systems. Specifically, we formulate the MPC policy as an encoder-only Transformer leveraging bidirectional self-attention, enabling simultaneous inference of entire control sequences in a single forward pass. This design inherently accommodates variable prediction horizons while ensuring low inference latency. Furthermore, we introduce a direct policy optimization framework that alternates between sampling and learning phases. Unlike imitation-based approaches dependent on precomputed optimal trajectories, TransMPC directly optimizes the true finite-horizon cost via automatic differentiation. Random horizon sampling combined with a replay buffer provides independent and identically distributed (i.i.d.) training samples, ensuring robust generalization across varying states and horizon lengths. Extensive simulations and real-world vehicle control experiments validate the effectiveness of TransMPC in terms of solution accuracy, adaptability to varying horizons, and computational efficiency.

cs.RO

From Flat to Hierarchical: Evolving Tree-structured Thoughts for Fine-grained Alpha Mining

Alpha mining, aimed at discovering predictive return signals, is typically formulated as symbolic regression. Traditional symbolic methods suffer from search inefficiency and biased prior knowledge. Recently, Large Language Models (LLMs) have emerged as a promising alternative, automatically generating textual thoughts and executable codes to achieve both efficient and interpretable alpha mining. However, existing approaches mostly focus on leveraging LLM's reasoning and reflection capabilities, yet largely neglect the positional bias due to the flat thought representation which restricts efficiency and diversity of the search process. This paper introduces Tree-structured thought Evolution (TreEvo), which evolves hierarchically decomposed thoughts to expand the effective search space. In addition, we propose a set of evolutionary operators tailored to structured thoughts. Experiments on four real-market datasets demonstrate that TreEvo not only obtains competitive alphas with traditional methods in up to 200 times fewer evaluations, but also consistently outperforms LLM-driven EAs across all datasets by $14.31\%$ on average.

cs.CE

WideSearch: Benchmarking Agentic Broad Info-Seeking

From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberate humans from this tedious work. However, the capability of these agents to perform such "wide-context" collection reliably and completely remains largely unevaluated due to a lack of suitable benchmarks. To bridge this gap, we introduce WideSearch, a new benchmark engineered to evaluate agent reliability on these large-scale collection tasks. The benchmark features 200 manually curated questions (100 in English, 100 in Chinese) from over 15 diverse domains, grounded in real user queries. Each task requires agents to collect large-scale atomic information, which could be verified one by one objectively, and arrange it into a well-organized output. A rigorous five-stage quality control pipeline ensures the difficulty, completeness, and verifiability of the dataset. We benchmark over 10 state-of-the-art agentic search systems, including single-agent, multi-agent frameworks, and end-to-end commercial systems. Most systems achieve overall success rates near 0\%, with the best performer reaching just 5\%. However, given sufficient time, cross-validation by multiple human testers can achieve a near 100\% success rate. These results demonstrate that present search agents have critical deficiencies in large-scale information seeking, underscoring urgent areas for future research and development in agentic search. Our dataset, evaluation pipeline, and benchmark results have been publicly released at https://widesearch-seed.github.io/

cs.CL

Learning from Expert Factors: Trajectory-level Reward Shaping for Formulaic Alpha Mining

Reinforcement learning (RL) has successfully automated the complex process of mining formulaic alpha factors, for creating interpretable and profitable investment strategies. However, existing methods are hampered by the sparse rewards given the underlying Markov Decision Process. This inefficiency limits the exploration of the vast symbolic search space and destabilizes the training process. To address this, Trajectory-level Reward Shaping (TLRS), a novel reward shaping method, is proposed. TLRS provides dense, intermediate rewards by measuring the subsequence-level similarity between partially generated expressions and a set of expert-designed formulas. Furthermore, a reward centering mechanism is introduced to reduce training variance. Extensive experiments on six major Chinese and U.S. stock indices show that TLRS significantly improves the predictive power of mined factors, boosting the Rank Information Coefficient by 9.29% over existing potential-based shaping algorithms. Notably, TLRS achieves a major leap in computational efficiency by reducing its time complexity with respect to the feature dimension from linear to constant, which is a significant improvement over distance-based baselines.

cs.LG

Probing intermediate-mass black hole binaries with the Lunar Gravitational-wave Antenna

New concepts for observing the gravitational waves (GWs) using a detector on the Moon, such as the Lunar Gravitational-wave Antenna (LGWA), have gained increasing attention. By utilizing the Moon as a giant antenna, the LGWA is expected to detect GWs in the frequency range from 1 millihertz (mHz) to several hertz, with optimal sensitivity in the decihertz band. Despite the debated formation and evolution channel of intermediate-mass black holes (IMBHs) with masses in the range of $[10^2, 10^5]\ {\rm M_\odot}$, binary systems containing at least one IMBH are widely believed to generate GWs spanning from mHz to a few Hz, making them a key scientific target for the LGWA. We explore the detectability of IMBH binaries with the LGWA in this work. The LGWA is more sensitive to nearby binaries (i.e. with redshift $z\lesssim0.5$) with the primary mass $m_1 \in [10^4, 10^5] \ {\rm M_\odot}$, while it prefers distant binaries (i.e. $z \gtrsim 5$) with $m_1 \in [10^3, 10^4] \ {\rm M_\odot}$. Considering a signal-to-noise ratio threshold of 10, our results imply that the LGWA can detect IMBH binaries up to $z \sim \mathcal{O}(10)$. We further show that the LGWA can constrain the primary mass with relative errors $\lesssim 0.1\%$ for binaries at $z \lesssim 0.5$. Furthermore, we show that the IMBH binaries at $z \lesssim 0.1$ can be used to constrain redshift with relative errors $\lesssim 10\%$, and those with $m_1 \in [10^4, 10^5] \ {\rm M_\odot}$ can be localized by the LGWA to be within $\mathcal{O} (10)$ $\rm deg^2$.

astro-ph.HE

A practical Bayesian method for gravitational-wave ringdown analysis with multiple modes

Gravitational-wave (GW) ringdown signals from black holes (BHs) encode crucial information about the gravitational dynamics in the strong-field regime, which offers unique insights into BH properties. In the future, the improving sensitivity of GW detectors is to enable the extraction of multiple quasi-normal modes (QNMs) from ringdown signals. However, incorporating multiple modes drastically enlarges the parameter space, posing computational challenges to data analysis. Inspired by the $F$-statistic method in the continuous GW searches, we develope an algorithm, dubbed as FIREFLY, for accelerating the ringdown signal analysis. FIREFLY analytically marginalizes the amplitude and phase parameters of QNMs to reduce the computational cost and speed up the full-parameter inference from hours to minutes, while achieving consistent posterior and evidence. The acceleration becomes more significant when more QNMs are considered. Rigorously based on the principle of Bayesian inference and importance sampling, our method is statistically interpretable, flexible in prior choice, and compatible with various advanced sampling techniques, providing a new perspective for accelerating future GW data analysis.

gr-qc

Vetting quark-star models with gravitational waves in the hierarchical Bayesian framework

The recent discovery of gravitational waves (GWs) has opened a new avenue for investigating the equation of state (EOS) of dense matter in compact stars, which is an outstanding problem in astronomy and nuclear physics. In the future, next-generation (XG) GW detectors will be constructed, deemed to provide a large number of high-precision observations. We investigate the potential of constraining the EOS of quark stars (QSs) with high-precision measurements of mass $m$ and tidal deformability $\Lambda$ from the XG GW observatories. We adopt the widely-used bag model for QSs, consisting of four microscopic parameters: the effective bag constant $B_{\rm eff}$, the perturbative quantum chromodynamics correction parameter $a_4$, the strange quark mass $m_s$, and the pairing energy gap $\Delta$. With the help of hierarchical Bayesian inference, for the first time we are able to infer the EOS of QSs combining multiple GW observations. Using the top 25 loudest GW events in our simulation, we find that, the constraints on $B_{\rm eff}$ and $\Delta$ are tightened by several times, while $a_4$ and $m_s$ are still poorly constrained. We also study a simplified 2-dimensional (2-d) EOS model which was recently proposed in literature. The 2-d model is found to exhibit significant parameter-estimation biases as more GW events are analyzed, while the predicted $m$-$\Lambda$ relation remains consistent with the full model.

astro-ph.HE

QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE

Alpha factor mining aims to discover investment signals from the historical financial market data, which can be used to predict asset returns and gain excess profits. Powerful deep learning methods for alpha factor mining lack interpretability, making them unacceptable in the risk-sensitive real markets. Formulaic alpha factors are preferred for their interpretability, while the search space is complex and powerful explorative methods are urged. Recently, a promising framework is proposed for generating formulaic alpha factors using deep reinforcement learning, and quickly gained research focuses from both academia and industries. This paper first argues that the originally employed policy training method, i.e., Proximal Policy Optimization (PPO), faces several important issues in the context of alpha factors mining. Herein, a novel reinforcement learning algorithm based on the well-known REINFORCE algorithm is proposed. REINFORCE employs Monte Carlo sampling to estimate the policy gradient-yielding unbiased but high variance estimates. The minimal environmental variability inherent in the underlying state transition function, which adheres to the Dirac distribution, can help alleviate this high variance issue, making REINFORCE algorithm more appropriate than PPO. A new dedicated baseline is designed to theoretically reduce the commonly suffered high variance of REINFORCE. Moreover, the information ratio is introduced as a reward shaping mechanism to encourage the generation of steady alpha factors that can better adapt to changes in market volatility. Evaluations on real assets data indicate the proposed algorithm boosts correlation with returns by 3.83\%, and a stronger ability to obtain excess returns compared to the latest alpha factors mining methods, which meets the theoretical results well.

q-fin.CP

Interior Object Geometry via Fitted Frames

We propose a means of computing fitted frames on the boundary and in the interior of objects and using them to provide the basis for producing geometric features from them that are not only alignment-free but most importantly can be made to correspond locally across a population of objects. We describe a representation targeted for anatomic objects which is designed to enable this strong locational correspondence within object populations and thus to provide powerful object statistics. It accomplishes this by understanding an object as the diffeomorphic deformation of the closure of the interior of an ellipsoid and by using a skeletal representation fitted throughout the deformation to produce a model of the target object, where the object is provided initially in the form of a boundary mesh. Via classification performance on hippocampi shape between individuals with a disorder vs. others, we compare our method to two state-of-theart methods for producing object representations that are intended to capture geometric correspondence across a population of objects and to yield geometric features useful for statistics, and we show notably improved classification performance by this new representation, which we call the evolutionary s-rep. The geometric features that are derived from each of the representations, especially via fitted frames, are discussed.

cs.CV

Constraints of the maximum mass of quark stars based on post-merger evolutions

We semi-analytically investigate the post-merger evolution of the binary quark star merger. The effective-one-body method is employed to estimate the energy and angular momentum dissipation due to gravitational waves in the inspiral phase. Three major mechanisms of energy and angular momentum dissipation are considered in the post-merger phase: mass outflows, neutrinos, and gravitational waves. The proportion of each mechanism could be determined by baryon number, energy and angular momentum conservation laws as well as the equilibrium model for rotating quark stars. Applying this analysis to the GW170817 event suggests two important conclusions: 1) a remnant quark star whose mass is smaller than the maximum mass of a uniformly rotating quark star can collapse before its rotational energy is dissipated via electromagnetic radiation (i.e., $\sim 100\,\mathrm{s}$) as the angular momentum left in the remnant quark star might not be large enough to sustain the additional self-gravity of the supramassive quark star due to the angular momentum dissipation of mass outflows, neutrinos and gravitational waves; 2) considering a general quark star equation of state model, a constraint on the maximum mass of cold and non-rotating quark stars is found as $M_{\mathrm{TOV}}\lesssim2.35^{+0.07}_{-0.17}\,M_{\odot}$, assuming a delayed collapse occurred before a large fraction of the total rotational energy ($\color{blue} \gtrsim 10^{53}\,$erg) of the merger remnant was deposited into the merger environment for the GW170817 event. These constraints could be improved with future merger events, once there are more evidences on its post-merger evolution channel or information on the amount of post-merger gravitational wave and neutrino emissions inferred from the multi-messenger observations.

astro-ph.HE