SearcharxivSearch

arXiv subjects

Jiaqiang Li

Publications and source records attributed to Jiaqiang Li.

6 recordsLinked to original sources

Atria Dawn: The Dawn of Agentic Superintelligence

As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.

cs.AI

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents

cs.AI

Sparse Equation Matching: A Derivative-Free Learning for General-Order Dynamical Systems

Equation discovery is a fundamental learning task for uncovering the underlying dynamics of complex systems, with wide-ranging applications in areas such as brain connectivity analysis, climate modeling, gene regulation, and physical simulation. However, many existing approaches rely on accurate derivative estimation and are limited to first-order dynamical systems, restricting their applicability in real-world scenarios. In this work, we propose Sparse Equation Matching (SEM), a unified framework that encompasses several existing equation discovery methods under a common formulation. SEM introduces an integral-based sparse regression approach using Green's functions, enabling derivative-free estimation of differential operators and their associated driving functions in general-order dynamical systems. The effectiveness of SEM is demonstrated through extensive simulations, benchmarking its performance against derivative-based approaches. We then apply SEM to electroencephalographic (EEG) data recorded during multiple oculomotor tasks, collected from 52 participants in a brain-computer interface experiment. Our method identifies active brain regions across participants and reveals task-specific connectivity patterns. These findings offer valuable insights into brain connectivity and the underlying neural mechanisms.

cs.LG

An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint

Recent work has demonstrated the remarkable potential of Large Language Models (LLMs) in test-time scaling. By making models think before answering, they are able to achieve much higher accuracy with extra inference computation. However, in many real-world scenarios, models are used under time constraints, where an answer should be given within a certain output length. It is unclear whether and how the reasoning ability of different LLMs remain effective under strict constraints. We take a first look at this problem by conducting an in-depth empirical study. Specifically, we test 30 LLMs on common reasoning datasets under a wide range of output length budgets, and we analyze the correlation between the inference accuracy and various properties including model type, model size, prompt style, etc. We also consider the mappings between token budgets and actual on-device latency budgets. The results have demonstrated several interesting findings regarding the budget-aware LLM reasoning ability that differ from the unconstrained situation, e.g. the optimal choices of either model size or prompt style change under different budgets. These findings offer timely evaluation to this area and practical guidance for users to deploy LLMs under real-world latency constraints.

cs.AI

Relaxation Critical Dynamics in Measurement-induced Phase Transitions

Measurement-induced phase transition (MIPT) describes the nonanalytical change of the entanglement entropy resulting from the interplay between measurement and unitary evolution. In this paper, we investigate the relaxation critical dynamics near the MIPT for different initial states in a one-dimensional quantum circuit. Specifically, when the initial state is in the volume-law phase with vanishing measurement probability, we find that the half-chain entanglement entropy $S$ decays as $S\propto t^{-1}$ with the coefficients proportional to the size of the system in the short-time stage; In contrast, when the initial state is the product state, $S$ increases with time as $S\propto \ln{t}$, consistent with previous studies. Despite these contrasting behaviors, we develop a unified scaling form to describe these scaling behaviors for different initial states where the off-critical-point effects can also be incorporated. This framework offers significant advantages for experimental MIPT detection. Our novel scheme, leveraging relaxation dynamical scaling, drastically reduces post-selection overhead, and can eliminate it completely with trackable classical simulation.

quant-ph

Driven Critical Dynamics in Measurement-induced Phase Transitions

Measurement-induced phase transitions (MIPT), characterizing abrupt changes in entanglement properties in quantum many-body systems subjected to unitary evolution with interspersed projective measurements, have garnered increasing interest. In this work, we generalize the Kibble-Zurek (KZ) driven critical dynamics that has achieved great success in traditional quantum and classical phase transitions to MIPT. By linearly changing the measurement probability $p$ to cross the critical point $p_c$ with driving velocity $R$, we identify the dynamic scaling relation of the entanglement entropy $S$ versus $R$ at $p_c$. For decreasing $p$ from the area-law phase, $S$ satisfies $S\propto \ln R$; while for increasing $p$ from the volume-law phase, $S$ satisfies $S\propto R^{1/r}$ in which $r=z+1/ν$ with $z$ and $ν$ being the dynamic and correlation length exponents, respectively. Moreover, we find that the driven dynamics from the volume-law phase violates the adiabatic-impulse scenario of the KZ mechanism. In spite of this, a unified finite-time scaling (FTS) form can be developed to describe these scaling behaviors. Besides, the dynamic scaling of the entanglement entropy of an auxiliary qubit $S_Q$ is also investigated to further confirm the universality of the FTS form. By successfully establishing the driven dynamic scaling theory of this newfashioned entanglement transition, we bring a new fundamental perspective into MIPT that can be detected in fast-developing quantum computers.

quant-ph