SearcharxivSearch

arXiv subjects

Xingjian Hu

Publications and source records attributed to Xingjian Hu.

11 recordsLinked to original sources

Accelerating Scientific Research with Gemini in the Real-World

We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materials science, Co-Scientist interfaced with a semi-automated chemical vapor deposition reactor to design a safe precursor route for MXenes; experimental execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, although further experiments are needed to confirm the atomic structure. Leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, it also tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. In biology, Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across inducer (IPTG) gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphological measurements. In computer science, Co-Scientist autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench (Hard and Professional) while reducing potential clinical harm under blinded physician evaluation. Finally, a double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews demonstrates that Co-Scientist's reliability modules reduce hallucination and plagiarism while improving research safety. Together, these results demonstrate progress toward closed-loop multi-agent scientific AI systems capable of accelerating real-world scientific discovery.

cs.AI

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail to differentiate question difficulty, limiting their ability to effectively distinguish models' capabilities. To address this limitation, we propose RankLLM, a novel framework designed to quantify both question difficulty and model competency. RankLLM introduces difficulty as the primary criterion for differentiation, enabling a more fine-grained evaluation of LLM capabilities. RankLLM's core mechanism facilitates bidirectional score propagation between models and questions. The core intuition of RankLLM is that a model earns a competency score when it correctly answers a question, while a question's difficulty score increases when it challenges a model. Using this framework, we evaluate 30 models on 35,550 questions across multiple domains. RankLLM achieves 90% agreement with human judgments and consistently outperforms strong baselines such as IRT. It also exhibits strong stability, fast convergence, and high computational efficiency, making it a practical solution for large-scale, difficulty-aware LLM evaluation.

cs.CL

NARRA-Gym for Evaluating Interactive Narrative Agents

Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limited: existing evaluations often focus on static prompts, isolated story generations, or post-hoc ratings, and therefore miss whether models can jointly manage story generation, long-context state and pacing, character simulation, empathic personalization, and story-grounded artifacts. We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and logs the full model-in-the-loop trajectory, including story construction, memory updates, planning, pacing interventions, and optional artifact synthesis. We evaluate nine frontier LLMs using a controlled LLM-as-judge sweep over eight benchmark personas and a human evaluation in which participants rate customized model outputs. Our results show substantial variation across models, personas, and evaluation dimensions: models that produce fluent stories can still fail on robustness, user experience, or resistance-sensitive personalization. These findings suggest that interactive narrative offers a useful benchmark for evaluating long-horizon, user-adaptive LLM behavior beyond isolated story quality.

cs.CL

GraphPL: Leveraging GNN for Efficient and Robust Modalities Imputation in Patchwork Learning

Current research on distributed multi-modal learning typically assumes that clients can access complete information across all modalities, which may not hold in practice. In this paper, we explore patchwork learning, in which the modalities available to different clients vary, and the objective is to impute the missing modalities for each client in an unsupervised manner. Existing methods are shown not to fully utilize the modality information as they tend to rely on only a subset of the observed modalities. To address this issue, we propose GraphPL, which combines graph neural networks with patchwork learning to flexibly integrate all observed modalities and remains robust with noisy inputs. Experimental results show that GraphPL achieves SOTA performance on benchmark datasets. Our results on real-world distributed electronic health record dataset show GraphPL learns strong downstream features and enables tasks like disease prediction via superior modality imputation.

cs.LG

Critical Self-Similar Markov Trees

Recently introduced and studied in arXiv:2407.07888, a self-similar Markov tree (ssMt) is a random decorated tree that vastly generalises the fragmentation tree. We study here the critical case that was left aside in arXiv:2407.07888. Borrowing techniques from branching random walk, in particular the recent result of Aïdékon--Hu--Shi arXiv:2409.01048, we can complete the picture by constructing critical ssMt, computing their fractal dimension and studying their associated harmonic and length measures using spinal decomposition.

math.PR

Scaling limits of critical FK-decorated random planar maps with $q=4$

We establish the first scaling limit for FK($q$)-weighted planar maps in the critical case $q=4$, resolving a problem that has remained open since Sheffield's seminal work arXiv:1108.2241. In that work, Sheffield proved a scaling limit for $q<4$ via the celebrated hamburger-cheeseburger bijection, which initiated the peanosphere (mating-of-trees) approach to Liouville quantum gravity. We prove that, at criticality, the associated burger count $\mathcal{S}$ and discrepancy $\mathcal{D}$ satisfy \[ \left(\frac{\mathcal{S}_{\lfloor nt \rfloor}}{\sqrt{n}}, \frac{\log(n)}{{2π}\sqrt{n}} \mathcal{D}_{\lfloor nt \rfloor}\right)_{t\in\mathbb{R}} \stackrel{\text{d}}{\longrightarrow} (B^1_t, B^2_{t})_{t\in\mathbb{R}}, \] where $B^1$ and $B^2$ are independent two-sided Brownian motions. To the best of our knowledge, no conjecture for the correct discrepancy scaling factor had previously been formulated. Matching the limiting process with the critical mating of trees arXiv:2109.00275, we establish the first rigorous planar map convergence towards CLE$_4$ and critical ($γ=2$) Liouville quantum gravity, in the peanosphere sense. Our proof is based on a novel approach that reveals the exactly solvable nature of the model through a correspondence with the (bicoloured) fully packed loop-$O(2)$ model on triangulations, and yields critical geometric exponents matching the predictions of conformal field theory.

math.PR

The scaling limit of the volume of loop O(n) quadrangulations

We study the volume of rigid loop-$O(n)$ quadrangulations with a boundary of length $2p$ in the non-generic critical regime. We prove that, as the half-perimeter $p$ goes to infinity, the volume scales in distribution to an explicit random variable. This limiting random variable is described in terms of the multiplicative cascades of Chen, Curien and Maillard arXiv:1702.06916, or alternatively (in the dilute case) as the law of the area of a unit-boundary $γ$-quantum disc, as determined by Ang and Gwynne arXiv:1903.09120, for suitable $γ$. Our arguments go through a classification of the map into several regions, where we rule out the contribution of bad regions to be left with a tractable portion of the map. One key observable for this classification is a Markov chain which explores the nested loops around a size-biased vertex pick in the map, making explicit the spinal structure of the discrete multiplicative cascade. We stress that our techniques enable us to include the boundary case $n=2$, that we define rigorously, and where the nested cascade structure is that of a critical branching random walk. In that case the scaling limit is given by the limit of the derivative martingale and is inverse-exponentially distributed, which answers a conjecture of arXiv:2005.06372v2.

math.PR

Zero-shot Autonomous Microscopy for Scalable and Intelligent Characterization of 2D Materials

Characterization of atomic-scale materials traditionally requires human experts with months to years of specialized training. Even for trained human operators, accurate and reliable characterization remains challenging when examining newly discovered materials such as two-dimensional (2D) structures. This bottleneck drives demand for fully autonomous experimentation systems capable of comprehending research objectives without requiring large training datasets. In this work, we present ATOMIC (Autonomous Technology for Optical Microscopy & Intelligent Characterization), an end-to-end framework that integrates foundation models to enable fully autonomous, zero-shot characterization of 2D materials. Our system integrates the vision foundation model (i.e., Segment Anything Model), large language models (i.e., ChatGPT), unsupervised clustering, and topological analysis to automate microscope control, sample scanning, image segmentation, and intelligent analysis through prompt engineering, eliminating the need for additional training. When analyzing typical MoS2 samples, our approach achieves 99.7% segmentation accuracy for single layer identification, which is equivalent to that of human experts. In addition, the integrated model is able to detect grain boundary slits that are challenging to identify with human eyes. Furthermore, the system retains robust accuracy despite variable conditions including defocus, color temperature fluctuations, and exposure variations. It is applicable to a broad spectrum of common 2D materials-including graphene, MoS2, WSe2, SnSe-regardless of whether they were fabricated via chemical vapor deposition or mechanical exfoliation. This work represents the implementation of foundation models to achieve autonomous analysis, establishing a scalable and data-efficient characterization paradigm that fundamentally transforms the approach to nanoscale materials research.

cond-mat.mtrl-sci

SketchRef: a Multi-Task Evaluation Benchmark for Sketch Synthesis

Sketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research. However, the field lacks a unified benchmark to evaluate the performance of various synthesis methods. To address this, we propose SketchRef, the first comprehensive multi-task evaluation benchmark for sketch synthesis. SketchRef fully leverages the shared characteristics between sketches and reference photos. It introduces two primary tasks: category prediction and structural consistency estimation, the latter being largely overlooked in previous studies. These tasks are further divided into five sub-tasks across four domains: animals, common things, human body, and faces. Recognizing the inherent trade-off between recognizability and simplicity in sketches, we are the first to quantify this balance by introducing a recognizability calculation method constrained by simplicity, mRS, ensuring fair and meaningful evaluations. To validate our approach, we collected 7,920 responses from art enthusiasts, confirming the effectiveness of our proposed evaluation metrics. Additionally, we evaluate the performance of existing sketch synthesis methods on our benchmark, highlighting their strengths and weaknesses. We hope this study establishes a standardized benchmark and offers valuable insights for advancing sketch synthesis algorithms.

cs.CV

TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression Recognition

Handwritten Mathematical Expression Recognition (HMER) has extensive applications in automated grading and office automation. However, existing sequence-based decoding methods, which directly predict $\LaTeX$ sequences, struggle to understand and model the inherent tree structure of $\LaTeX$ and often fail to ensure syntactic correctness in the decoded results. To address these challenges, we propose a novel model named TAMER (Tree-Aware Transformer) for handwritten mathematical expression recognition. TAMER introduces an innovative Tree-aware Module while maintaining the flexibility and efficient training of Transformer. TAMER combines the advantages of both sequence decoding and tree decoding models by jointly optimizing sequence prediction and tree structure prediction tasks, which enhances the model's understanding and generalization of complex mathematical expression structures. During inference, TAMER employs a Tree Structure Prediction Scoring Mechanism to improve the structural validity of the generated $\LaTeX$ sequences. Experimental results on CROHME datasets demonstrate that TAMER outperforms traditional sequence decoding and tree decoding models, especially in handling complex mathematical structures, achieving state-of-the-art (SOTA) performance.

cs.CV

SegHist: A General Segmentation-based Framework for Chinese Historical Document Text Line Detection

Text line detection is a key task in historical document analysis facing many challenges of arbitrary-shaped text lines, dense texts, and text lines with high aspect ratios, etc. In this paper, we propose a general framework for historical document text detection (SegHist), enabling existing segmentation-based text detection methods to effectively address the challenges, especially text lines with high aspect ratios. Integrating the SegHist framework with the commonly used method DB++, we develop DB-SegHist. This approach achieves SOTA on the CHDAC, MTHv2, and competitive results on HDRC datasets, with a significant improvement of 1.19% on the most challenging CHDAC dataset which features more text lines with high aspect ratios. Moreover, our method attains SOTA on rotated MTHv2 and rotated HDRC, demonstrating its rotational robustness. The code is available at https://github.com/LumionHXJ/SegHist.

cs.CL