SearcharxivSearch

arXiv subjects

Shengzhi Li

Publications and source records attributed to Shengzhi Li.

13 recordsLinked to original sources

AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing

Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce. Yet a review is only useful if acting on it leads to a measurable improvement in the paper. We present AppliedScientist, a closed-loop system that couples an autonomous AI scientist with an AI reviewer, and evaluate it by iteratively revising rejected papers from a range of research subfields. To mirror how human authors build on earlier drafts, the AI scientist has access to its previous versions during revision. To avoid bias from prior judgments, however, each review is generated independently, with the reviewer having no memory of earlier feedback or scores. We compare three revision settings: one initialized with the original venue reviews, one initialized with AI-generated reviews, and autonomous self-revision using the same fixed prompt in every round. Because the reviewer both guides and evaluates the revision, we also assess the human-initialized revisions using Stanford Reviewer as an independent evaluator. Reviewer-guided revision consistently improves more than fixed-prompt self-revision, and Stanford Reviewer also assigns higher scores to later revisions. AppliedScientist resolves 128 of 150 execution-related weaknesses (85.3%), but only 2 of 18 idea-related weaknesses (11.1%), suggesting that iterative revision is effective at improving experiments and implementation, but rarely changes concerns about novelty or significance.

cs.AI

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and correcting mistakes over multiple turns. We introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions. To make real interactions verifiable, we curate 109 repository-level tasks from 11,260 recorded sessions, selecting sessions with recoverable repository states, clear user goals, and observable outcomes. To replay these interactions across agents, we build a reactive LLM-based user simulator that preserves the original users' intents and provides feedback when the coding agent's progress requires it. To evaluate agents as collaborators, we measure both final repository correctness and the number of corrective feedback turns required during the interaction. Experiments with frontier coding agents show that stronger agents generally achieve higher final success rates while requiring fewer interventions, suggesting an improved user experience.

cs.SE

Loop Quantum Kaluza-Klein Cosmology and Inflation

We present the detailed analyses of five-dimensional loop quantum Kaluza-Klein cosmology based on the symmetric reduction of the connection formulation of the full theory. The previous results in a particular scenario are extended to more general cases. The effective scalar constraint for the geometric sector of the model is derived by the systematic semi-classical analysis in both the canonical and path-integral formulations, incorporating the quantum fluctuations as a subleading-order correction. The resulting effective scalar constraint not only exhibits the correct classical limit of the quantum system, but also serves as the basis for investigating the following three distinct effective scenarios through the incorporation of matter contributions: (i) vacuum, (ii) minimally coupling with a scalar field, and (iii) coupling with the dust. In all the three effective scenarios, the big bang and potential past big rip singularities in the classical model are naturally resolved by including the leading-order quantum correction of holonomies. Moreover, the visible universe undergoes a super-inflationary phase after overcoming the classical big bang singularity, during which the phenomenologically desired 55 e-folds can be achieved by appropriate initial conditions. In the case where the subleading-order quantum fluctuation term is included as a constant, the evolutions of the five-dimensional universe in all the three effective scenarios not only achieve sufficient inflation in the visible dimensions, but also exhibit re-collapse behaviors at certain large scales. Hence the cosmic inflation may originate from the interplay between compact extra dimensions and quantum geometric effects.

gr-qc

Effective dynamics of Janis-Newman-Winicour spacetime

The effective dynamics of the Janis-Newman-Winicour spacetime inspired by loop quantum gravity is studied. Two different schemes are considered to regularize the Hamiltonian constraint for the quantum dynamics. In the $μ_0$ scheme in which the quantum parameters are treated as constants, the equations of motion generated by the effective Hamiltonian are solved analytically. The resulting quantum-corrected effective spacetime obviously extends the effective spacetime previously obtained in the literature. In the new effective spacetime, the naked singularity and the central singularity presented in the classical JNW spacetime are resolved by a series of quantum bounces. In the scheme of choosing the quantum parameters as Dirac observables, the effective dynamics is also solved in the light of the solution in $μ_0$ scheme. It turns out that the resulting effective spacetime has singularities due to the appearance of the zero points of the time reparametrization functions. Hence, the effective theory in this scheme does not remain valid throughout the full spacetime.

gr-qc

PushupBench: Your VLM is not good at counting pushups

Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition counting. The best frontier model achieves 42.1\% exact accuracy; open-source 4B models score $\sim$6\%, matching supervised baselines. We show that accuracy alone misleads -- weaker models exploit the modal count rather than reason temporally. Fine-tuning on counting with 1k samples transfers to general video understanding: MVBench (+2.15), PerceptionTest (+1.88), TVBench (+4.54), suggesting counting is a proxy for broader temporal reasoning.PushupBench incorporated in \texttt{lmms-eval} (https://github.com/EvolvingLMMs-Lab/lmms-eval/pull/1262) and hosted on (pushupbench.com/)

cs.CV

IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR

Peer review relies on substantive, evidence-based questions, yet current LLMs generate surface-level queries that perform worse than human reviewer questions in expert evaluation. To address this gap, we curate a high-quality dataset of reviewer questions from OpenReview and conduct a human preference study where expert annotators evaluate question-paper pairs across three dimensions: effort, evidence, and grounding. From these annotations, we train IntelliReward, a reward model built from a frozen autoregressive LLM with trainable multi-head transformers. Validated against expert judgments, IntelliReward predicts reviewer-question quality better than API-based SFT baselines and provides scalable evaluation. We apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) with IntelliReward to train IntelliAsk, a question-generation model aligned with human standards of effortful, evidence-based critique. Human evaluations show IntelliAsk generates more grounded, substantive and effortful questions than strong baselines and reduces reliance on first-page content. We also find improvements on reasoning and writing benchmarks, suggesting reviewer-question quality correlates with broader capabilities. Compared to Qwen3-32B, IntelliAsk improves MuSR (68.3 vs 64.7 Acc) and WritingBench (8.31 vs 8.07). We release our code, filtered review dataset, expert annotations, IntelliAsk and IntelliReward to support automatic evaluation of grounding, effort, and evidence in LLM-generated review questions.

cs.CL

Deparametrization and quantization of scalar-tensor gravity and its cosmological model

The degree of freedom of the scalar field in scalar-tensor gravity is employed as "time" to deparametrize the Hamiltonian constraint of the theory. The deparametrized system is then nonperturbatively quantized by the approach of loop quantum gravity. This results in a discrete time evolution of the physical states with respect to the gravitational degree of freedom in the quantum theory. In the corresponding Brans-Dicke cosmological model, the physical solutions to the quantum Hamiltonian constraint are obtained in the light of the deparametrization. The quantum dynamics indicates that the classical big bang singularity is replaced by a quantum bounce.

gr-qc

Effective Dynamics of Loop Quantum Kaluza-Klein Cosmology

The five-dimensional loop quantum Kaluza-Klein cosmology is constructed based on the symmetric reduction of the connection formulation of the full theory. Through semiclassical analysis, the effective scalar constraint for the cosmological model coupled with a dust field is derived, incorporating the quantum fluctuations of geometry as a subleading order correction. It demonstrates that the quantum model has the correct classical limit. The explicit solutions to the equations of motion show that the big bang and past big rip singularities in the classical model are avoided by a quantum bounce and a quantum collapse respectively in the effective model. In a particular scenario, the dynamical compactification of the extra dimension is realized, while the observable four-dimensional universe transitions through three distinct epochs: (i) a super-inflationary phase generating 55 e-folds, (ii) a decelerated expansion era, and (iii) a late-time accelerated expansion phase driven by quantum fluctuations. These results suggest that both cosmic inflation and dark energy may originate from the interplay between the compact extra dimension and quantum geometric effects.

gr-qc

Loop Quantum Vector-Tensor Gravity and Its Spherically Symmetric Model

The Hamiltoinian analysis of the vector-tensor theory of gravity is performed. The resulting geometrical dynamics is reformulated into the connection dynamics, with the real SU(2)-connection serving as one of the configuration variables. This formulation allows us to extend the loop quantization scheme of general relativity to the vector-tensor theory, thereby rigorously constructing its quantum kinematical framework. The scalar constraint is promoted to a well-defined operator in the vertex Hilbert space, to represent quantum dynamics. Moreover, the spherically symmetric model of the vector-tensor theory is obtained by the symmetric reduction. Following the general deparametrization strategy for theories with diffeomorphism invariance, the spherically symmetric model can be fully deparametrized in terms of the degrees of freedom of the vector field. The corresponding reduced phase space quantization is carried out. The physical Hamiltonian generating relative evolution is promoted to a well-defined operator on the physical Hilbert space.

gr-qc

Abstract2Appendix: Academic Reviews Enhance LLM Long-Context Capabilities

Large language models (LLMs) have shown remarkable performance across various tasks, yet their ability to handle long-context reading remains challenging. This study explores the effectiveness of leveraging high-quality academic peer review data for fine-tuning LLMs to enhance their long-context capabilities. We compare the Direct Preference Optimization (DPO) method with the Supervised Fine-Tuning (SFT) method, demonstrating DPO's superiority and data efficiency. Our experiments show that the fine-tuned model achieves a 4.04-point improvement over phi-3 and a 2.6\% increase on the Qasper benchmark using only 2000 samples. Despite facing limitations in data scale and processing costs, this study underscores the potential of DPO and high-quality data in advancing LLM performance. Additionally, the zero-shot benchmark results indicate that aggregated high-quality human reviews are overwhelmingly preferred over LLM-generated responses, even for the most capable models like GPT-4o. This suggests that high-quality human reviews are extremely rich in information, reasoning, and long-context retrieval, capabilities that even the most advanced models have not fully captured. These findings highlight the high utility of leveraging human reviews to further advance the field.

cs.CL

Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Multi-modal large language models (MLLMs) are expected to support multi-turn queries of interchanging image and text modalities in production. However, the current MLLMs trained with visual-question-answering (VQA) datasets could suffer from degradation, as VQA datasets lack the diversity and complexity of the original text instruction datasets with which the underlying language model was trained. To address this degradation, we first collect a lightweight, 5k-sample VQA preference dataset where answers were annotated by Gemini for five quality metrics in a granular fashion and investigate standard Supervised Fine-tuning, rejection sampling, Direct Preference Optimization (DPO) and SteerLM algorithms. Our findings indicate that with DPO, we can surpass the instruction-following capabilities of the language model, achieving a 6.73 score on MT-Bench, compared to Vicuna's 6.57 and LLaVA's 5.99. This enhancement in textual instruction-following capability correlates with boosted visual instruction performance (+4.9\% on MM-Vet, +6\% on LLaVA-Bench), with minimal alignment tax on visual knowledge benchmarks compared to the previous RLHF approach. In conclusion, we propose a distillation-based multi-modal alignment model with fine-grained annotations on a small dataset that restores and boosts MLLM's language capability after visual instruction tuning.

cs.CL

SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs

In this work, we present SciGraphQA, a synthetic multi-turn question-answer dataset related to academic graphs. SciGraphQA is 13 times larger than ChartVQA, the previously largest chart-visual question-answering dataset. It is also the largest open-sourced chart VQA dataset with non-synthetic charts. To build our dataset, we selected 290,000 Computer Science or Machine Learning ArXiv papers published between 2010 and 2020, and then used Palm-2 to generate 295K samples of open-vocabulary multi-turn question-answering dialogues about the graphs. As context, we provided the text-only Palm-2 with paper title, abstract, paragraph mentioning the graph, and rich text contextual data from the graph itself, obtaining dialogues with an average 2.23 question-answer turns for each graph. We asked GPT-4 to assess the matching quality of our question-answer turns given the paper's context, obtaining an average rating of 8.7/10 on our 3K test set. We evaluated the 0-shot capability of the most popular MLLM models such as LLaVa, mPLUGowl, BLIP-2, and openFlamingo's on our dataset, finding LLaVA-13B being the most performant with a CIDEr score of 0.08. We further enriched the question prompts for LLAVA by including the serialized data tables extracted from the graphs using the DePlot model, boosting LLaVA's 0-shot CIDEr to 0.15. To verify the validity of our dataset, we also fine-tuned LLaVa using our dataset, reaching a substantially higher CIDEr score of 0.26. We anticipate further accuracy improvement by including segmentation mask tokens and leveraging larger LLM backbones coupled with emergent prompting techniques. Our code and data are open-sourced.

cs.CL

Connection Dynamics of Reduced 5-dimensional Kaluza-Klein Theory and Its Deparametrization

The connection dynamics of the 5-dimensional Kaluza-Klein theory reduced on 4-dimensional spacetime is obtained by performing the Hamiltonian analysis and canonical transformations. Deparametrization is achieved in the spherically symmetric model of the theory without introducing additional matter fields beyond the 5-dimensional gravity. Thus the physical time evolution and the physical spatial coordinate can be provided by the geometrical degrees of freedom in the higher dimensional spacetime.

gr-qc