SearcharxivSearch

arXiv subjects

Heyang Li

Publications and source records attributed to Heyang Li.

6 recordsLinked to original sources

Towards a new paradigm of scientific discovery with socialized artificial intelligence

Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particular phenomena. Computation extended inquiry into systems beyond direct observation, while data-intensive methods opened new spaces of pattern and prediction. Science now confronts a different frontier. The central challenge is no longer simply to produce more information, but to organize expanding knowledge, reasoning, and evidence into a coherent process of discovery. Here, we introduce Bridging Literature, Agents, and Zero-gap Experimentation (BLAZE), a paradigm of socialized scientific intelligence. BLAZE conceives AI not as an assistant for isolated research tasks, but as an organizational infrastructure for scientific discovery. It connects persistent knowledge, collective reasoning, empirical validation, and human judgment within a continuous research lifecycle, transforming fragmented activities into a cumulative process of inquiry, criticism, and revision. The central premise of BLAZE is that scientific intelligence does not arise from computation alone. It emerges from the sustained interaction among knowledge, hypotheses, experiments, and collective verification. By organizing humans and machines within a shared scientific process, BLAZE makes discovery more traceable, reproducible, and cumulative while preserving human creativity, judgment, and responsibility. Socialized scientific intelligence may provide a foundation for the next era of science. Its purpose is not to replace human discovery, but to extend the scale, depth, and continuity of collective scientific inquiry.

cs.AI

Demanding peer review is associated with higher impact in published science

Peer review shapes which scientific claims enter the published record, but its internal dynamics are hard to measure at scale because reviewer criticism and author revision are usually embedded in long, unstructured correspondence. Here we use a fixed-prompt large language model pipeline to convert the review correspondence of \textit{Nature Communications} papers published from 2017 to 2024 into structured reviewer--author interactions. We find that review pressure is concentrated in the first round and focused disproportionately on core claims rather than peripheral presentation. Higher average opinion strength is also associated with more reviewer disagreement, while review patterns vary little with broad team attributes, consistent with relatively impartial evaluation. Contrary to the intuition that stronger papers should pass review more smoothly, with greater reviewer--author agreement and less extensive revision, we find that stronger criticism, higher-quality comments, and greater revision burden are associated with higher later citation impact within accepted papers. We finally show that fields differ more in review style than in review length, pointing to disciplinary variation in how criticism is negotiated and resolved. These findings position open peer review not just as a gatekeeping mechanism but as a measurable record of how influential scientific claims are challenged, defended, and revised before entering the published record.

cs.DL

Answering Constraint Path Queries over Graphs

Constraints are powerful declarative constructs that allow users to conveniently restrict variable values that potentially range over an infinite domain. In this paper, we propose a constraint path query language over property graphs, which extends Regular Path Queries (RPQs) with SMT constraints on data attributes in the form of equality constraints and Linear Real Arithmetic (LRA) constraints. We provide efficient algorithms for evaluating such path queries over property graphs, which exploits optimization of macro-states (among others, using theory-specific techniques). In particular, we demonstrate how such an algorithm may effectively utilize highly optimized SMT solvers for resolving such constraints over paths. We implement our algorithm in MillenniumDB, an open-source graph engine supporting property graph queries and GQL. Our extensive empirical evaluation in a real-world setting demonstrates the viability of our approach.

cs.DB

Exact solutions for nonlinear trapped lee waves in the $\beta$-plane approximation

In this paper, we construct exact solutions that character three-dimensional, nonlinear trapped lee waves propagation superimposed on longitudinal atmospheric currents in the $\beta$-plane approximation. The solutions obtained are presented in Lagrangian coordinates, and are Gerstner-like solutions. In the process, we also derive the dispersion relation and analyze the density, pressure and the vorticity qualitatively.

math.DS

X-ray Multimodal Intrinsic-Speckle-Tracking

We develop X-ray Multi-modal Intrinsic-Speckle-Tracking (MIST), a form of X-ray speckle-tracking that is able to recover both the position-dependent phase shift and the position-dependent small-angle X-ray scattering (SAXS) signal of a phase object. MIST is based on combining a Fokker-Planck description of paraxial X-ray optics, with an optical-flow formalism for X-ray speckle-tracking. Only two images need to be taken in the presence of the sample, corresponding to two different transverse positions of the speckle-generating membrane, in order to recover both the refractive and local-SAXS properties of the sample. Like the optical-flow X-ray phase-retrieval method which it generalises, the MIST method implicitly rather than explicitly tracks both the transverse motion and the diffusion of speckles that is induced by the presence of a sample. Application to X-ray synchrotron data shows the method to be efficient, rapid and stable.

eess.IV

Branch lengths on Yule trees and the expected loss of phylogenetic diversity

Diversification is nested, and early models suggested this could lead to a great deal of evolutionary redundancy in the Tree of Life. This result is based on a particular set of branch lengths produced by the common coalescent, where pendant branches leading to tips can be very short compared to branches deeper in the tree. Here, we analyze alternative and more realistic Yule and birth-death models. We show how censoring at the present both makes average branches one half what we might expect and makes pendant and interior branches roughly equal in length. Although dependent on whether we condition on the size of the tree, its age, or both, these results hold both for the Yule model and for birth-death models with moderate extinction. Importantly, the rough equivalency in interior and exterior branch lengths means the loss of evolutionary history with loss of species can be roughly linear. Under these models, the Tree of Life may offer limited redundancy in the face of ongoing species loss.

q-bio.PE