SearcharxivSearch

arXiv subjects

Yishan Wu

Publications and source records attributed to Yishan Wu.

8 recordsLinked to original sources

FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judgments, which often overlook local reasoning gaps. While formal theorem provers like Lean offer a path to rigorous verification, using them to evaluate informal text requires solving locality and semantic mismatches: a prover might bypass a local flaw by proving an overly broad target, or validate an auto-formalized statement that drifts from the original mathematical intent. To address this, we introduce FaithSieve, a Lean-assisted framework for fine-grained evaluation of natural-language mathematical proofs. FaithSieve decomposes coarse proof steps into local reasoning units, extracts typed proof obligations, and verifies them through a formal evaluation agent. Formal validation is gated by semantic alignment scoring, so Lean evidence is incorporated only when the formal statement faithfully preserves the context, objects, and logical form of the original claim. We construct two expert-verified datasets, ProofLoc-Olympiad and ProofLoc-University, to benchmark first-error localization. On the 350-problem Olympiad dataset, FaithSieve using a GPT-5.4 backbone achieves 81.43% exact first-error accuracy, outperforming the direct-judging baseline of 72.29%. Furthermore, on the 200-problem ProofLoc-University benchmark spanning six advanced domains, FaithSieve reaches 84.5% exact accuracy, compared to 75.0% for the direct judge. Our work demonstrates that decomposing proofs into fine-grained units and grounding them with faithful formal evidence significantly improves reliable evaluation of natural-language reasoning.

cs.AI

The Boolean polynomial polytope with multiple choice constraints

We consider a class of $0$-$1$ polynomial programming termed multiple choice polynomial programming (MCPP) where the constraint requires exact one component per subset of the partition to be $1$ after all the entries are partitioned. Compared to the unconstrained counterpart, there are few polyhedral studies of MCPP in general form. This paper serves as the first attempt to propose a polytope associated with a hypergraph to study MCPP, which is the convex hull of $0$-$1$ vectors satisfying multiple choice constraints and production constraints. With the help of the decomposability property, we obtain an explicit half-space representation of the MCPP polytope when the underlying hypergraph is $α$-acyclic by induction on the number of hyperedges, which is an analogy of the acyclicity results on the multilinear polytope by Del Pia and Khajavirad (SIAM J Optim 28 (2018) 1049) when the hypergraph is $γ$-acyclic. We also present a necessary and sufficient condition for the inequalities lifted from the facet-inducing ones for the multilinear polytope to be still facet-inducing for the MCPP polytope. This result covers the particular cases by Bärmann, Martin and Schneider (SIAM J Optim 33 (2023) 2909).

math.OC

An ODE approach to multiple choice polynomial programming

We propose an ODE approach to solving multiple choice polynomial programming (MCPP) after assuming that the optimum point can be approximated by the expected value of so-called thermal equilibrium as usually did in simulated annealing. The explicit form of the feasible region and the affine property of the objective function are both fully exploited in transforming the MCPP problem into the ODE system. We also show theoretically that a local optimum of the former can be obtained from an equilibrium point of the latter. Numerical experiments on two typical combinatorial problems, MAX-$k$-CUT and the calculation of star discrepancy, demonstrate the validity of our ODE approach, and the resulting approximate solutions are of comparable quality to those obtained by the state-of-the-art heuristic algorithms but with much less cost. This paper also serves as the first attempt to use a continuous algorithm for approximating the star discrepancy.

math.OC

Measuring the Competitive Pressure of Academic Journals and the Competitive Intensity within Subjects

A journal's impact and similarity with rivals is closely related to its competitive intensity. A subject area can be considered as an ecological system of journals, and can then be measured using the competitive intensity concept from plant systems. Based on Journal Citation Reports data from 1997, 2000, 2005, 2010, and 2013, we calculated the mutual citation, cosine similarity, and competitive relationship matrices for mycology journals. We derived the mutual citation network for mycology according to Journal Citation Reports data from 2013. We calculated each journal's competitive pressure, and the competitive intensity for the subject. We found that competitive pressures are very variable among journals. Differences between a journal's absolute and relative influence are related to the competitive pressure. A more powerful journal has lower competitive pressure. New journals have more competitive pressure. If there are no other influences, the competition intensity of a subject will continue to increase. Furthermore, we found that if a subject has more journals, its competitive intensity decreases.

cs.DL

6G Downlink Transmission via Rate Splitting Space Division Multiple Access Based on Grouped Code Index Modulation

A novel rate splitting space division multiple access (SDMA) scheme based on grouped code index modulation (GrCIM) is proposed for the sixth generation (6G) downlink transmission. The proposed RSMA-GrCIM scheme transmits information to multiple user equipments (UEs) through the space division multiple access (SDMA) technique, and exploits code index modulation for rate splitting. Since the CIM scheme conveys information bits via the index of the selected Walsh code and binary phase shift keying (BPSK) signal, our RSMA scheme transmits the private messages of each user through the indices, and the common messages via the BPSK signal. Moreover, the Walsh code set is grouped into several orthogonal subsets to eliminate the interference from other users. A maximum likelihood (ML) detector is used to recovery the source bits, and a mathematical analysis is provided for the upper bound bit error ratio (BER) of each user. Comparisons are also made between our proposed scheme and the traditional SDMA scheme in spectrum utilization, number of available UEs, etc. Numerical results are given to verify the effectiveness of the proposed SDMA-GrCIM scheme.

cs.IT

Impact of JD Bernal Thoughts in the Science of Science upon China: Implications for Quantitative Studies of Science Today

John Desmond Bernal (1901-1970) was one of the most eminent scientists in molecular biology, and also regarded as the founding father of the Science of Science. His book The Social Function of Science laid the theoretical foundations for the discipline. In this article, we summarize four chief characteristics of his ideas in the Science of Science: the socio-historical perspective, theoretical models, qualitative and quantitative approaches, and studies of science planning and policy. China has constantly reformed its scientific and technological system based on research evidence of the Science of Science. Therefore, we analyze the impact of Bernal Science-of-Science thoughts on the development of Science of Science in China, and discuss how they might be usefully taken still further in quantitative studies of science.

physics.soc-ph

A Probe into Causes of Non-citation Based on Survey Data

Empirical analysis results about the possible causes leading to non-citation may help increase the potential of researchers' work to be cited and editorial staffs of journals to identify contributions with potential high quality. In this study, we conduct a survey on the possible causes leading to citation or non-citation based on a questionnaire. We then perform a statistical analysis to identify the major causes leading to non-citation in combination with the analysis on the data collected through the survey. Most respondents to our questionnaire identified eight major causes that facilitate easy citation of one's papers, such as research hotspots and novel topics of content, longer intervals after publication, research topics similar to my work, high quality of content, reasonable self-citation, highlighted title, prestigious authors, academic tastes and interests similar to mine.They also pointed out that the vast difference between their current and former research directions as the primary reason for their previously uncited papers. They feel that text that includes notes, comments, and letters to editors are rarely cited, and the same is true for too short or too lengthy papers. In comparison, it is easier for reviews, articles, or papers of intermediate length to be cited.

cs.DL

The Effects of Research Level and Article Type on the Differences between Citation Metrics and F1000 Recommendations

F1000 recommendations have been validated as a potential data source for research evaluation, but reasons for differences between F1000 Article Factor (FFa scores) and citations remain to be explored. By linking 28254 publications in F1000 to citations in Scopus, we investigated the effect of research level and article type on the internal consistency of assessments based on citations and FFa scores. It turns out that research level has little impact, while article type has big effect on the differences. These two measures are significantly different for two groups: non-primary research or evidence-based research publications are more highly cited rather than highly recommended, however, translational research or transformative research publications are more highly recommended by faculty members but gather relatively lower citations. This can be expected because citation activities are usually practiced by academic authors while the potential for scientific revolutions and the suitability for clinical practice of an article should be investigated from the practitioners' points of view. We conclude with a policy relevant recommendation that the application of bibliometric approaches in research evaluation procedures should include the proportion of three types of publications: evidence-based research, transformative research, and translational research. The latter two types are more suitable to be assessed through peer review.

cs.DL