SearcharxivSearch

arXiv subjects

Biao He

Publications and source records attributed to Biao He.

At least 19 recordsLinked to original sources

Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs

We investigated Agentic RL with large language models on the \textsc{TravelPlanner} benchmark. Our approach, \textsc{Planner-R1}, achieved a \textbf{56.9\%} final-pass rate with only 180 training queries, a $2.7\times$ improvement over GPT-5's $21.2\%$ baseline and the strongest agentic result on the public leaderboard. A central finding was that smaller models (8B) were highly responsive to reward shaping: with dense process-level signals, they reached competitive performance while being $3.5\times$ more compute-efficient and $1.5\times$ more memory-efficient than 32B models. Larger models were more robust under sparse rewards but exhibited smaller relative gains from shaping and higher variance across runs. While curriculum learning offered no significant benefit, shaped rewards consistently amplified learning dynamics, making 8B models the most efficient setting for agentic RL. Crucially, these gains did not come at the cost of overfitting: fine-tuned models mostly maintained or exceeded baseline performance on out-of-domain tasks, including \textsc{Multi-IF}, \textsc{NaturalPlan}, and $\tau$-\textsc{Bench}. These results establish reward shaping as a decisive lever for scaling agentic RL, highlight the competitive strength of smaller models, and demonstrate that efficiency can be achieved without sacrificing generalization.

cs.AI

Hybrid Machine Learning and Physics-based Modelling of Pedestrian Pushing Behaviours at Bottlenecks

In high-density crowds, close proximity between pedestrians makes the steady state highly vulnerable to disruption by pushing behaviours, potentially leading to serious accidents. However, the scarcity of experimental data has hindered systematic studies of its mechanisms and accurate modelling. Using behavioural data from bottleneck experiments, we investigate pedestrian heterogeneity in pushing tendencies, showing that pedestrians tend to push under high-motivation and in wider corridors. We introduce a spatial discretization method to encode neighbour states into feature vectors, serving together with pedestrian pushing tendencies as inputs to a random forest model for predicting pushing behaviours. Through comparing speed-headway relationships, we reveal that pushing behaviours correspond to an aggressive space-utilization movement strategy. Consequently, we propose a hybrid machine learning and physics-based model integrating pushing tendencies heterogeneity, pushing behaviours prediction, and dynamic movement strategies adjustment. Validations show that the hybrid model effectively reproduces experimental crowd dynamics and fits to incorporate additional behaviours.

physics.soc-ph

Analytic formula for the proton radioactivity spectroscopic factor

In the present work, we systematically study the spectroscopic factor of proton radioactivity ($S_p$) with $A>100$ using the deformed two-potential approach (D-TPA). It is found that there is a link between the quadrupole deformation parameter of proton emitter and $S_p$. Based on this result, we propose a simple analytic formula for estimating the spectroscopic factor of proton radioactivity. With the help of this formula, the calculated half-lives of proton radioactivity can reproduce the experimental data successfully within a factor of 2.77. Furthermore, we extend the D-TPA with this formula for evaluating the spectroscopic factor to predict the proton radioactivity half-lives of 12 proton radioactivity candidates whose radioactivity is energetically allowed or observed but not yet quantified in NUBASE2020. For comparison, the universal decay law for proton radioactivity (UDLP) and the new Geiger-Nuttall law (NG-N) are also used. It turns out that all of the predicted results are basically consistent with each other.

nucl-th

ODDFUZZ: Discovering Java Deserialization Vulnerabilities via Structure-Aware Directed Greybox Fuzzing

Java deserialization vulnerability is a severe threat in practice. Researchers have proposed static analysis solutions to locate candidate vulnerabilities and fuzzing solutions to generate proof-of-concept (PoC) serialized objects to trigger them. However, existing solutions have limited effectiveness and efficiency. In this paper, we propose a novel hybrid solution ODDFUZZ to efficiently discover Java deserialization vulnerabilities. First, ODDFUZZ performs lightweight static taint analysis to identify candidate gadget chains that may cause deserialization vulner-abilities. In this step, ODDFUZZ tries to locate all candidates and avoid false negatives. Then, ODDFUZZ performs directed greybox fuzzing (DGF) to explore those candidates and generate PoC testcases to mitigate false positives. Specifically, ODDFUZZ applies a structure-aware seed generation method to guarantee the validity of the testcases, and adopts a novel hybrid feedback and a step-forward strategy to guide the directed fuzzing. We implemented a prototype of ODDFUZZ and evaluated it on the popular Java deserialization repository ysoserial. Results show that, ODDFUZZ could discover 16 out of 34 known gadget chains, while two state-of-the-art baselines only identify three of them. In addition, we evaluated ODDFUZZ on real-world applications including Oracle WebLogic Server, Apache Dubbo, Sonatype Nexus, and protostuff, and found six previously unreported exploitable gadget chains with five CVEs assigned.

cs.CR

Improving Java Deserialization Gadget Chain Mining via Overriding-Guided Object Generation

Java (de)serialization is prone to causing security-critical vulnerabilities that attackers can invoke existing methods (gadgets) on the application's classpath to construct a gadget chain to perform malicious behaviors. Several techniques have been proposed to statically identify suspicious gadget chains and dynamically generate injection objects for fuzzing. However, due to their incomplete support for dynamic program features (e.g., Java runtime polymorphism) and ineffective injection object generation for fuzzing, the existing techniques are still far from satisfactory. In this paper, we first performed an empirical study to investigate the characteristics of Java deserialization vulnerabilities based on our manually collected 86 publicly known gadget chains. The empirical results show that 1) Java deserialization gadgets are usually exploited by abusing runtime polymorphism, which enables attackers to reuse serializable overridden methods; and 2) attackers usually invoke exploitable overridden methods (gadgets) via dynamic binding to generate injection objects for gadget chain construction. Based on our empirical findings, we propose a novel gadget chain mining approach, \emph{GCMiner}, which captures both explicit and implicit method calls to identify more gadget chains, and adopts an overriding-guided object generation approach to generate valid injection objects for fuzzing. The evaluation results show that \emph{GCMiner} significantly outperforms the state-of-the-art techniques, and discovers 56 unique gadget chains that cannot be identified by the baseline approaches.

cs.CR

Systematic study of two-proton radioactivity half-lives based on a modified Gamow-like model

In the present work, we systematically study the two-proton (2p) radioactivity half-lives of nuclei close to the proton drip line within a modified Gamow-like model. Using this model, the calculated 2p radioactivity half-lives can well reproduce the experimental data. Moreover, we use this model to predict the 2p radioactivity half-lives of 22 candidates whose 2p radioactivity is energetically allowed or observed but not yet quantied in evaluated nuclear properties table NUBASE2016. The predicted results are in good agreement with the ones obtained by using Gamow-like model, effective liquid drop model (ELDM), generalized liquid drop model (GLDM) as well as a four-parameter formula.

nucl-th

Systematic study on proton radioactivity of spherical proton emitters within two--potential approach

In the present work we systematically study the half--lives of proton radioactivity for spherical proton emitters with ${Z\ge 69}$ based on two--potential approach. While the nuclear potential of the emitted proton--daughter nucleus is adopted by a parameterized cosh type, the parameters of the depth and diffuseness for nuclear potential are determined by fitting experimental data of 32 spherical proton emitters. In order to reduce the deviations between experimental half-lives and calculated ones, we propose a simple analytic expression for formation probability of proton radioactivity with the same orbital angular momentum $l$. The results indicate that the formation probability can be simply described by a formula of $A_d^{1/3}$. Moreover, the linear relationship between the formation probability and the fragmentation potential also exists. The calculated half-lives can well reproduce the experimental data.

nucl-th

Systematic study of two-proton radioactivity half-lives within the two-potential approach with Skyrme-Hartree-Fock

In this work, we systematically study the two-proton($2p$) radioactivity half-lives using the two-potential approach while the nuclear potential is obtained by using Skyrme-Hartree-Fock approach with the Skyrme effective interaction of {SLy8}. For true $2p$ radioactivity($Q_{2p}$ $>$ 0 and $Q_p$ $< $0, where the $Q_p$ and $Q_{2p}$ are the released energy of the one-proton and two-proton radioactivity), the standard deviation between the experimental half-lives and our theoretical calculations is {0.701}. In addition, we extend this model to predict the half-lives of 15 possible $2p$ radioactivity candidates with $Q_{2p}$ $>$ 0 taken from the evaluated atomic mass table AME2016. The calculated results indicate that a clear linear relationship between the logarithmic $2p$ radioactivity half-lives $\rm{log}_{10}T_{1/2}$ and coulomb parameters [ ($Z_{d}^{0.8}$+$l^{0.25}$)$Q_{2p}^{-1/2}$] considered the effect of orbital angular momentum proposed by Liu $et$ $al$ [Chin. Phys. C \textbf{45}, 024108 (2021)] is also existed. For comparison, the generalized liquid drop model(GLDM), the effective liquid drop model(ELDM) and Gamow-like model are also used. Our predicted results are consistent with the ones obtained by the other models.

nucl-th

New Geiger-Nuttall law for two-proton radioactivity

In the present work, combining with the Geiger-Nuttall law, a two-parameter empirical formula is proposed to study the two-proton (2p) radioactivity. Using this formula, the calculated 2p radioactivity half-lives are in good agreement with the experimental data as well as the calculated ones obtained by Goncalves et al: ([Phys. Lett. B 774, 14 (2017)]) using the effective liquid drop model (ELDM), Sreeja et al: ([Eur. Phys. J. A 55, 33 (2019)]) using a four-parameter empirical formula and Cui et al: ([Phys. Rev. C 101: 014301 (2020)]) using a generalized liquid drop model (GLDM). In addition, this two-parameter empirical formula is extended to predict the half-lives of 22 possible 2p radioactivity candidates, whose the 2p radioactivity released energy Q2p>0, obtained from the latest evaluated atomic mass table AME2016. The predicted results have good consistency with ones using other theoretical models such as the ELDM, GLDM and four-parameter empirical formula.

nucl-th

Coverage Analysis of Relay Assisted Millimeter Wave Cellular Networks with Spatial Correlation

We propose a novel analytical framework for evaluating the coverage performance of a millimeter wave (mmWave) cellular network where idle user equipments (UEs) act as relays. In this network, the base station (BS) adopts either the direct mode to transmit to the destination UE, or the relay mode if the direct mode fails, where the BS transmits to the relay UE and then the relay UE transmits to the destination UE. To address the drastic rotational movements of destination UEs in practice, we propose to adopt selection combining at destination UEs. New expression is derived for the signal-to-interference-plus-noise ratio (SINR) coverage probability of the network. Using numerical results, we first demonstrate the accuracy of our new expression. Then we show that ignoring spatial correlation, which has been commonly adopted in the literature, leads to severe overestimation of the SINR coverage probability. Furthermore, we show that introducing relays into a mmWave cellular network vastly improves the coverage performance. In addition, we show that the optimal BS density maximizing the SINR coverage probability can be determined by using our analysis.

cs.IT

New Geiger-Nuttall law for proton radioactivity

In the present work considering the contributions of the daughter nuclear charge and the orbital angular momentum taken away by the emitted proton, we propose a two-parameter formula of new Geiger-Nuttall law for proton radioactivity. A set of universal parameters of this law is obtained by fitting 44 experimental data of proton emitters in the ground state and isomeric state. The calculated results can reproduce the experimental data well. For a comparison, the calculations performed using other theoretical methods, such as UDLP proposed by Qi, et al. [https://journals.aps.org/prc/abstract/10.1103/PhysRevC.85.011303], the CPPM-Guo2013 analyzed by our previous work [Deng, et al., https://link.springer.com/article/10.1140/epja/i2019-12728-0] and the modified Gamow-like model proposed by us [Chen, et al., https://iopscience.iop.org/article/10.1088/1361-6471/ab1a56] are also included. Meanwhile, we extend this new Geiger-Nuttall law to predict the proton radioactivity half-lives for $51 \leq Z \leq 91$ nuclei, whose proton radioactivity is energetically allowed or observed but not yet quantified in NUBASE2016.

nucl-th

An Analysis of Two-User Uplink Asynchronous Non-Orthogonal Multiple Access Systems

Recent studies have numerically demonstrated the possible advantages of the asynchronous non-orthogonal multiple access (ANOMA) over the conventional synchronous non-orthogonal multiple access (NOMA). The ANOMA makes use of the oversampling technique by intentionally introducing a timing mismatch between symbols of different users. Focusing on a two-user uplink system, for the first time, we analytically prove that the ANOMA with a sufficiently large frame length can always outperform the NOMA in terms of the sum throughput. To this end, we derive the expression for the sum throughput of the ANOMA as a function of signal-to-noise ratio (SNR), frame length, and normalized timing mismatch. Based on the derived expression, we find that users should transmit at full powers to maximize the sum throughput. In addition, we obtain the optimal timing mismatch as the frame length goes to infinity. Moreover, we comprehensively study the impact of timing error on the ANOMA throughput performance. Two types of timing error, i.e., the synchronization timing error and the coordination timing error, are considered. We derive the throughput loss incurred by both types of timing error and find that the synchronization timing error has a greater impact on the throughput performance compared to the coordination timing error.

cs.IT

Low-Complexity Reconfigurable MIMO for Millimeter Wave Communications

The performance of millimeter wave (mmWave) multiple-input multiple-output (MIMO) systems is limited by the sparse nature of propagation channels and the restricted number of radio frequency (RF) chains at transceivers. The introduction of reconfigurable antennas offers an additional degree of freedom on designing mmWave MIMO systems. This paper provides a theoretical framework for studying the mmWave MIMO with reconfigurable antennas. Based on the virtual channel model, we present an architecture of reconfigurable mmWave MIMO with beamspace hybrid analog-digital beamformers and reconfigurable antennas at both the transmitter and the receiver. We show that employing reconfigurable antennas can provide throughput gain for the mmWave MIMO. We derive the expression for the average throughput gain of using reconfigurable antennas in the system, and further derive the expression for the outage throughput gain for the scenarios where the channels are (quasi) static. Moreover, we propose a low-complexity algorithm for reconfiguration state selection and beam selection. Our numerical results verify the derived expressions for the throughput gains and demonstrate the near-optimal throughput performance of the proposed low-complexity algorithm.

cs.IT

On Secure Transmission Design: An Information Leakage Perspective

Information leakage rate is an intuitive metric that reflects the level of security in a wireless communication system, however, there are few studies taking it into consideration. Existing work on information leakage rate has two major limitations due to the complicated expression for the leakage rate: 1) the analytical and numerical results give few insights into the trade-off between system throughput and information leakage rate; 2) and the corresponding optimal designs of transmission rates are not analytically tractable. To overcome such limitations and obtain an in-depth understanding of information leakage rate in secure wireless communications, we propose an approximation for the average information leakage rate in the fixed-rate transmission scheme. Different from the complicated expression for information leakage rate in the literature, our proposed approximation has a low-complexity expression, and hence, it is easy for further analysis. Based on our approximation, the corresponding approximate optimal transmission rates are obtained for two transmission schemes with different design objectives. Through analytical and numerical results, we find that for the system maximizing throughput subject to information leakage rate constraint, the throughput is an upward convex non-decreasing function of the security constraint and much too loose security constraint does not contribute to higher throughput; while for the system minimizing information leakage rate subject to throughput constraint, the average information leakage rate is a lower convex increasing function of the throughput constraint.

cs.CR

Covert Wireless Communication with a Poisson Field of Interferers

In this paper, we study covert communication in wireless networks consisting of a transmitter, Alice, an intended receiver, Bob, a warden, Willie, and a Poisson field of interferers. Bob and Willie are subject to uncertain shot noise due to the ambient signals from interferers in the network. With the aid of stochastic geometry, we analyze the throughput of the covert communication between Alice and Bob subject to given requirements on the covertness against Willie and the reliability of decoding at Bob. We consider non-fading and fading channels. We analytically obtain interesting findings on the impacts of the density and the transmit power of the concurrent interferers on the covert throughput. That is, the density and the transmit power of the interferers have no impact on the covert throughput as long as the network stays in the interference-limited regime, for both the non-fading and the fading cases. When the interference is sufficiently small and comparable with the receiver noise, the covert throughput increases as the density or the transmit power of the concurrent interferers increases.

cs.IT

Millimeter Wave Communications with Reconfigurable Antennas

The highly sparse nature of propagation channels and the restricted use of radio frequency (RF) chains at transceivers limit the performance of millimeter wave (mmWave) multiple-input multiple-output (MIMO) systems. Introducing reconfigurable antennas to mmWave can offer an additional degree of freedom on designing mmWave MIMO systems. This paper provides a theoretical framework for studying the mmWave MIMO with reconfigurable antennas. We present an architecture of reconfigurable mmWave MIMO with beamspace hybrid analog-digital beamformers and reconfigurable antennas at both the transmitter and the receiver. We show that employing reconfigurable antennas can provide throughput gain for the mmWave MIMO. We derive the expression for the average throughput gain of using reconfigurable antennas, and further simplify the expression by considering the case of large number of reconfiguration states. In addition, we propose a low-complexity algorithm for the reconfiguration state and beam selection, which achieves nearly the same throughput performance as the optimal selection of reconfiguration state and beams by exhaustive search.

cs.IT

On the Design of Secure Non-Orthogonal Multiple Access Systems

This paper proposes a new design of non-orthogonal multiple access (NOMA) under secrecy considerations. We focus on a NOMA system where a transmitter sends confidential messages to multiple users in the presence of an external eavesdropper. The optimal designs of decoding order, transmission rates, and power allocated to each user are investigated. Considering the practical passive eavesdropping scenario where the instantaneous channel state of the eavesdropper is unknown, we adopt the secrecy outage probability as the secrecy metric. We first consider the problem of minimizing the transmit power subject to the secrecy outage and quality of service constraints, and derive the closed-form solution to this problem. We then explore the problem of maximizing the minimum confidential information rate among users subject to the secrecy outage and transmit power constraints, and provide an iterative algorithm to solve this problem. We find that the secrecy outage constraint in the studied problems does not change the optimal decoding order for NOMA, and one should increase the power allocated to the user whose channel is relatively bad when the secrecy constraint becomes more stringent. Finally, we show the advantage of NOMA over orthogonal multiple access in the studied problems both analytically and numerically.

cs.IT

Interference Alignment with Power Splitting Relays in Multi-User Multi-Relay Networks

In this paper, we study a multi-user multi-relay interference-channel network, where energy-constrained relays harvest energy from sources' radio frequency (RF) signals and use the harvested energy to forward the information to destinations. We adopt the interference alignment (IA) technique to address the issue of interference, and propose a novel transmission scheme with the IA at sources and the power splitting (PS) at relays. A distributed and iterative algorithm to obtain the optimal PS ratios is further proposed, aiming at maximizing the sum rate of the network. The analysis is then validated by simulation results. Our results show that the proposed scheme with the optimal design significantly improves the performance of the network.

cs.IT