SearcharxivSearch

arXiv subjects

Yanqi Luo

Publications and source records attributed to Yanqi Luo.

11 recordsLinked to original sources

Tomographic Phase Imaging with Randomized Probe Imaging

We demonstrate tomographic phase imaging of cubic gold nanoparticles with sub-100 nm resolution requiring just a single far-field coherent diffraction pattern per projection. By using randomized probe imaging (RPI), a real-space amplitude and phase image can be reconstructed from a single diffraction pattern. These phase images are then fed into a standard tomographic workflow to retrieve the 3D volume. Nanoscale X-ray phase tomography has often relied on ptychography, which requires slow 2D scanning for each projection. By using RPI, the data collection is greatly sped up at the cost of some resolution. We compare an RPI-tomography volume to a ptychography-tomography volume of the same sample, finding high consistency between the two methods. To the best of our knowledge, this is the first application of RPI to tomography.

physics.optics

CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing

Scientific data processing often requires task-specific algorithms or AI models, creating a barrier for domain scientists who need to analyze their data but may not have extensive computing or image-processing expertise. This barrier is especially pronounced when data are noisy, have a high dynamic range, are sparsely labeled, or are only loosely specified. We introduce CVEvolve, an autonomous agentic harness with a zero-code interface for scientific data-processing algorithm discovery. CVEvolve combines a multi-round search strategy with tools for code execution, evaluation implementation, history management, holdout testing, and optional inspection of scientific data and visual outputs. The search alternates between discovery and improvement actions, and uses lineage-aware stochastic candidate sampling to balance exploration and exploitation. We demonstrate CVEvolve on X-ray fluorescence microscopy image registration, Bragg peak detection, high-energy diffraction microscopy image segmentation, and hybrid analytical-learning-based affine registration. Across these tasks, CVEvolve discovers algorithms that improve over baseline methods, while holdout test tracking helps identify candidates that generalize better than later over-optimized alternatives. These results show that zero-code, autonomous LLM-powered algorithm development can help domain scientists turn unstructured scientific image data into practical algorithms and downstream scientific discoveries.

cs.AI

EAA: Automating materials characterization with vision language model agents

We present Experiment Automation Agents (EAA), a vision-language-model-driven agentic system designed to automate complex experimental microscopy workflows. EAA integrates multimodal reasoning, tool-augmented action, and optional long-term memory to support both autonomous procedures and interactive user-guided measurements. Built on a flexible task-manager architecture, the system enables workflows ranging from fully agent-driven automation to logic-defined routines that embed localized LLM queries. EAA further provides a modern tool ecosystem with two-way compatibility for Model Context Protocol (MCP), allowing instrument-control tools to be consumed or served across applications. We demonstrate EAA at an imaging beamline at the Advanced Photon Source, including automated zone plate focusing, natural language-described feature search, and interactive data acquisition. These results illustrate how vision-capable agents can enhance beamline efficiency, reduce operational burden, and lower the expertise barrier for users.

cs.AI

FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration

AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize codebases in response to evolving software ecosystems. Traditionally, such migrations have relied on rule-based systems and human intervention. With the advent of powerful large language models (LLMs), AI-driven agentic frameworks offer a promising alternative-but their effectiveness has not been systematically evaluated. In this paper, we introduce FreshBrew, a novel benchmark for evaluating AI agents on project-level Java migrations, with a specific focus on measuring an agent's ability to preserve program semantics and avoid reward hacking, which we argue requires projects with high test coverage for a rigorous and reliable evaluation. We benchmark several state-of-the-art LLMs, and compare their performance against established rule-based tools. Our evaluation of AI agents on this benchmark of 228 repositories shows that the top-performing model, Gemini 2.5 Flash, can successfully migrate 52.3 percent of projects to JDK 17. Our empirical analysis reveals novel insights into the critical strengths and limitations of current agentic approaches, offering actionable insights into their real-world applicability. Our empirical study reveals failure modes of current AI agents in realistic Java modernization tasks, providing a foundation for evaluating trustworthy code-migration systems. By releasing FreshBrew, we aim to facilitate rigorous, reproducible evaluation and catalyze progress in AI-driven codebase modernization.

cs.SE

WALT: Web Agents that Learn Tools

Web agents promise to automate complex browser tasks, but current methods remain brittle -- relying on step-by-step UI interactions and heavy LLM reasoning that break under dynamic layouts and long horizons. Humans, by contrast, exploit website-provided functionality through high-level operations like search, filter, and sort. We introduce WALT (Web Agents that Learn Tools), a framework that reverse-engineers latent website functionality into reusable invocable tools. Rather than hypothesizing ad-hoc skills, WALT exposes robust implementations of automations already designed into websites -- spanning discovery (search, filter, sort), communication (post, comment, upvote), and content management (create, edit, delete). Tools abstract away low-level execution: instead of reasoning about how to click and type, agents simply call search(query) or create(listing). This shifts the computational burden from fragile step-by-step reasoning to reliable tool invocation. On VisualWebArena and WebArena, WALT achieves higher success with fewer steps and less LLM-dependent reasoning, establishing a robust and generalizable paradigm for browser automation.

cs.CV

SCUBA: Salesforce Computer Use Benchmark

We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task instances derived from real user interviews, spanning three primary personas, platform administrators, sales representatives, and service agents. The tasks test a range of enterprise-critical abilities, including Enterprise Software UI navigation, data manipulation, workflow automation, information retrieval, and troubleshooting. To ensure realism, SCUBA operates in Salesforce sandbox environments with support for parallel execution and fine-grained evaluation metrics to capture milestone progress. We benchmark a diverse set of agents under both zero-shot and demonstration-augmented settings. We observed huge performance gaps in different agent design paradigms and gaps between the open-source model and the closed-source model. In the zero-shot setting, open-source model powered computer-use agents that have strong performance on related benchmarks like OSWorld only have less than 5\% success rate on SCUBA, while methods built on closed-source models can still have up to 39% task success rate. In the demonstration-augmented settings, task success rates can be improved to 50\% while simultaneously reducing time and costs by 13% and 16%, respectively. These findings highlight both the challenges of enterprise tasks automation and the promise of agentic solutions. By offering a realistic benchmark with interpretable evaluation, SCUBA aims to accelerate progress in building reliable computer-use agents for complex business software ecosystems.

cs.AI

MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources

We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk-mitigated sourcing strategy that combines public-domain and permissively licensed text (e.g., CC-BY/Apache) with carefully justified low-risk additions (e.g., government works and EU TDM-eligible sources). MixtureVitae adopts a simple, single-stage pretraining recipe that integrates a large proportion of permissive synthetic instruction and reasoning data-signals typically introduced during post-training and generally scarce in permissive web corpora. We categorize all sources into a three-tier scheme that reflects varying risk levels and provide shard-level provenance metadata to enable risk-aware usage. In controlled experiments using the open-sci-ref training protocol (fixed architectures and hyperparameters; 50B and 300B token budgets across 130M-1.7B parameters), models trained on MixtureVitae consistently outperform other permissive datasets across a suite of standard benchmarks, and at the 1.7B-parameters/300B-tokens setting, they surpass FineWeb-Edu and approach DCLM late in training. Performance is particularly strong on MMLU and on math and code benchmarks: a 1.7B model pretrained on 300B MixtureVitae tokens matches or exceeds a strong 1.7B instruction-tuned baseline on GSM8K, HumanEval, and MBPP, despite using over 36 times fewer tokens (300B vs. ~11T). Supported by a thorough decontamination analysis, these results show that permissive-first data with high instruction and reasoning density, tiered by licensing and provenance-related risk, can provide a practical and risk-mitigated foundation for training capable LLMs, reducing reliance on broad web scrapes without sacrificing competitiveness. Code: https://github.com/ontocord/mixturevitae

cs.CL

Diversity Enhances an LLM's Performance in RAG and Long-context Task

The rapid advancements in large language models (LLMs) have highlighted the challenge of context window limitations, primarily due to the quadratic time complexity of the self-attention mechanism (\(O(N^2)\), where \(N\) denotes the context window length). This constraint impacts tasks such as retrieval-augmented generation (RAG) in question answering (Q\&A) and long context summarization. A common approach involves selecting content with the highest similarity to the query; however, this often leads to redundancy and the exclusion of diverse yet relevant information. Building on principles from Maximal Marginal Relevance (MMR) and Farthest Point Sampling (FPS), we integrate diversity into the content selection process. Our findings reveal that incorporating diversity substantially increases the recall of selecting relevant sentences or chunks before LLM-based Q\&A and summarization. These results highlight the importance of maintaining diversity in future LLM applications to further improve summarization and Q\&A outcomes.

cs.CL

Imaging real-time amorphization of hybrid perovskite solar cells under electrical biasing

Perovskite solar cells have drawn much attention in recent years, owing to its world-record setting photovoltaic performances. Despite its promising use in tandem applications and flexible devices, its practicality is still limited by its structural instability often arising from ion migration and defect formation. While it is generally understood that ion instability is a primary cause for degradation, there is still a lack of direct evidence of structural transformation at the atomistic scale. Such an understanding is crucial to evaluate and pin-point how such instabilities are induced relative to external perturbations such as illumination or electrical bias with time, allowing researchers to devise effective strategies to mitigate them. Here, we designed an in-situ TEM setup to enable real-time observation of amorphization in double cation mixed perovskite materials under electrical biasing at 1 V. It is found that amorphization occurs along the (001) and (002) planes, which represents the observation of in-situ facet-dependent amorphization of a perovskite crystal. To reverse the degradation, the samples were heated at 50 oC and was found to recrystallize, effectively regaining its performance losses. This work is vital toward understanding fundamental ion-migration phenomena and address instability challenges of perovskite optoelectronics.

cond-mat.mtrl-sci

Residual Nanoscale Strain in Cesium Lead Bromide Perovskite Reduces Stability and Alters Local Band Gap

Using nanoprobe X-ray diffraction microscopy, we investigate the relationship between residual strains from crystal growth in CsPbBr$_3$ thin film crystals, their stability, and local bandgap. We find that out-of-plane compressive strain that arises from cooldown from crystallization is detrimental to material stability under X-ray irradiation. We also find that the optical photoluminescence red shifts as a result of the out-of-plane compressive strain. The sensitivity of bandgap to strain suggests possible applications such as stress-sensitive sensors. Mosaicity, the formation of small misorientations in neighboring crystalline domains we observe in some CsPbBr$_3$ single crystals, indicates the significant variations in crystal quality that can occur even in single-crystal halide perovskites. The nano-diffraction results suggest that reducing local strains is a necessary path to enhance the stability of perovskite optoelectronic materials and devices from light-emitting diodes to high-energy detectors.

cond-mat.mtrl-sci

Homogenization of Halide Distribution and Carrier Dynamics in Alloyed Organic-Inorganic Perovskites

Perovskite solar cells have shown remarkable efficiencies beyond 22%, through organic and inorganic cation alloying. However, the role of alkali-metal cations is not well-understood. By using synchrotron-based nano-X-ray fluorescence and complementary measurements, we show that when adding RbI and/or CsI the halide distribution becomes homogenous. This homogenization translates into long-lived charge carrier decays, spatially homogenous carrier dynamics visualized by ultrafast microscopy, as well as improved photovoltaic device performance. We find that Rb and K phase-segregate in highly concentrated aggregates. Synchrotron-based X-ray-beam-induced current and electron-beam-induced current of solar cells show that Rb clusters do not contribute to the current and are recombination active. Our findings bring light to the beneficial effects of alkali metal halides in perovskites, and point at areas of weakness in the elemental composition of these complex perovskites, paving the way to improved performance in this rapidly growing family of materials for solar cell applications.

cond-mat.mtrl-sci