SearcharxivSearch

arXiv subjects

Linyi Zou

Publications and source records attributed to Linyi Zou.

4 recordsLinked to original sources

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration decisions, and computational resources. Pioneering autoresearch agents, despite great success, still lack mechanisms for continuity, recovery from dead ends, and value-driven compute allocation, which inherently undermines overall search efficiency, wastes computational resources, and lowers the chance of ultimate success. To bridge this gap, we introduce ScienceFlow, an end-to-end autoresearch agent framework that organizes long-horizon research work into research segments grounded in executable workspaces. It represents research progress as recoverable executable states, enabling efficient exploration, revision, and execution. Transitions between research segments are governed by Executable-State Transition through Re-Anchoring (ESTRA), which selects either the live state or an archived state as the next anchor and determines whether to continue or redirect the research trajectory. An evidence-aware execution controller allocates resources to physical jobs based on resource availability, remaining budget, and validated progress. We evaluate ScienceFlow on tasks spanning machine learning, scientific modeling, and mathematical optimization. Results on diverse long-horizon benchmarks demonstrate its ability to sustain effective research processes, highlighted by a SOTA 70.22 percent Any-Medal score on the full MLE-bench within a 24-hour budget, outperforming prior reported results by 4.92 percentage points. The efficacy of ScienceFlow further demonstrates that efficient state management, adaptive exploration, and objective-aligned execution are critical for scaling autonomous research beyond short-horizon interactions.

cs.AI

Bayesian Mendelian randomization testing of interval causal null hypotheses: ternary decision rules and loss function calibration

Our approach to Mendelian Randomization (MR) analysis is designed to increase reproducibility of causal effect "discoveries" by: (i) using a Bayesian approach to inference; (ii) replacing the point null hypothesis with a region of practical equivalence consisting of values of negligible magnitude for the effect of interest, while exploiting the ability of Bayesian analysis to quantify the evidence of the effect falling inside/outside the region; (iii) rejecting the usual binary decision logic in favour of a ternary logic where the hypothesis test may result in either an acceptance or a rejection of the null, while also accommodating an "uncertain" outcome. We present an approach to calibration of the proposed method via loss function, which we use to compare our approach with a frequentist one. We illustrate the method with the aid of a study of the causal effect of obesity on risk of juvenile myocardial infarction.

stat.ME

Bayesian Mendelian randomization with study heterogeneity and data partitioning for large studies

Background: Mendelian randomization (MR) is a useful approach to causal inference from observational studies when randomised controlled trials are not feasible. However, study heterogeneity of two association studies required in MR is often overlooked. When dealing with large studies, recently developed Bayesian MR is limited by its computational expensiveness. Methods: We addressed study heterogeneity by proposing a random effect Bayesian MR model with multiple exposures and outcomes. For large studies, we adopted a subset posterior aggregation method to tackle the problem of computation. In particular, we divided data into subsets and combine estimated subset causal effects obtained from the subsets". The performance of our method was evaluated by a number of simulations, in which part of exposure data was missing. Results: Random effect Bayesian MR outperformed conventional inverse-variance weighted estimation, whether the true causal effects are zero or non-zero. Data partitioning of large studies had little impact on variations of the estimated causal effects, whereas it notably affected unbiasedness of the estimates with weak instruments and high missing rate of data. Our simulation results indicate that data partitioning is a good way of improving computational efficiency, for little cost of decrease in unbiasedness of the estimates, as long as the sample size of subsets is reasonably large. Conclusions: We have further advanced Bayesian MR by including random effects to explicitly account for study heterogeneity. We also adopted a subset posterior aggregation method to address the issue of computational expensiveness of MCMC, which is important especially when dealing with large studies. Our proposed work is likely to pave the way for more general model settings, as Bayesian approach itself renders great flexibility in model constructions.

stat.ME

Overlapping-sample Mendelian randomisation with multiple exposures: A Bayesian approach

Background: Mendelian randomization (MR) has been widely applied to causal inference in medical research. It uses genetic variants as instrumental variables (IVs) to investigate putative causal relationship between an exposure and an outcome. Traditional MR methods have dominantly focussed on a two-sample setting in which IV-exposure association study and IV-outcome association study are independent. However, it is not uncommon that participants from the two studies fully overlap (one-sample) or partly overlap (overlapping-sample). Methods: We proposed a method that is applicable to all the three sample settings. In essence, we converted a two- or overlapping- sample problem to a one-sample problem where data of some or all of the individuals were incomplete. Assume that all individuals were drawn from the same population and unmeasured data were missing at random. Then the unobserved data were treated au pair with the model parameters as unknown quantities, and thus, could be imputed iteratively conditioning on the observed data and estimated parameters using Markov chain Monte Carlo. We generalised our model to allow for pleiotropy and multiple exposures and assessed its performance by a number of simulations using four metrics: mean, standard deviation, coverage and power. Results: Higher sample overlapping rate and stronger instruments led to estimates with higher precision and power. Pleiotropy had a notably negative impact on the estimates. Nevertheless, overall the coverages were high and our model performed well in all the sample settings. Conclusions: Our model offers the flexibility of being applicable to any of the sample settings, which is an important addition to the MR literature which has restricted to one- or two- sample scenarios. Given the nature of Bayesian inference, it can be easily extended to more complex MR analysis in medical research.

stat.ME