SearcharxivSearch

arXiv subjects

Andy Cheng

Publications and source records attributed to Andy Cheng.

2 recordsLinked to original sources

ArchEval: Measuring AI Agents as Computer Architects

Computer architecture has long used benchmarks to make progress measurable. LLM agents create a different measurement problem: success is not merely writing code or tuning parameters. The agent must interpret workloads, choose mechanisms, use simulators, predict performance, satisfy hard constraints, and decide which feasible design is worth evaluating. This paper introduces ArchEval, a benchmark and platform for evaluating LLM agents on computer architecture design and optimization. It contains 20 challenges across CPU core mechanisms, system architecture, memory systems, accelerators, and compute-in-memory, backed by eight simulators. Each challenge is posed under three settings: L1 full harness, with repeated simulator feedback; L2 simulator-code container, where simulator source is available but the agent must assemble its own workflow; and L3 agent-only, with no runnable feedback before submission. Each run reports baseline-normalized verifier performance and records the full trajectory, connecting results to workload analysis, simulator-tool use, prediction, constraint handling, and artifact integrity. Initial results show a sharp boundary in current agents. With L1 support, all four evaluated agents reach or exceed baseline and improve real designs across diverse simulators. Removing support exposes weaknesses: many agents fail to turn simulator source into useful experiments, and L3 predictions often disagree with verifier results. In L3, only GPT-5.5 + Codex remains above baseline, reaching 1.21x geomean performance and a 65% win rate; the other three fall below baseline. Even GPT-5.5 + Codex has only a 15% performance-modeling pass rate. ArchEval frames today's agents as useful optimization assistants rather than autonomous architects, and identifies capabilities needed next: simulator-tool use, calibrated prediction, pre-feedback judgment, and useful mechanism discovery.

cs.AR

Characterization of the ejecta from NASA/DART impact on Dimorphos: observations and Monte Carlo models

The NASA/DART (Double Asteroid Redirection Test) spacecraft successfully crashed on Dimorphos, the secondary component of the binary (65803) Didymos system. Following the impact, a large dust cloud was released, and a long-lasting dust tail was developed. We have extensively monitored the dust tail from the ground and from the Hubble Space Telescope (HST). We provide a characterization of the ejecta dust properties, i.e., particle size distribution and ejection speeds, ejection geometric parameters, and mass, by combining both observational data sets, and by using Monte Carlo models of the observed dust tail. The differential size distribution function that best fits the imaging data was a broken power-law, having a power index of --2.5 for particles of r$\le$ 3 mm, and of --3.7 for larger particles. The particles range in sizes from 1 $\mu$m up to 5 cm. The ejecta is characterized by two components, depending on velocity and ejection direction. The northern component of the double tail, observed since October 8th 2022, might be associated to a secondary ejection event from impacting debris on Didymos, although it is also possible that this feature results from the binary system dynamics alone. The lower limit to the total dust mass ejected is estimated at $\sim$6$\times$10$^6$ kg, half of this mass being ejected to interplanetary space.

astro-ph.EP