SearcharxivSearch

arXiv subjects

Ingrida Semenec

Publications and source records attributed to Ingrida Semenec.

3 recordsLinked to original sources

World-Time Compute with Verified Code World Models

LLMs generalize across a domain only after seeing many real, labeled examples, which most domains lack. We study a way to manufacture it cheaply. When a domain's dynamics can be written as code, one template instantiates into many world models: executable, verifiable programs over symbolic state, each an inexhaustible source of exactly-labeled trajectories. Fine-tuning an LLM on trajectories through many such worlds, which we call world-time compute, a training-time analogue of test-time compute, lifts generalization to held-out worlds it never trained on (synthesized world families). Gains are largest where capability is scarcest: +29 points at 0.5B; the largest model's lift is within noise, consistent with saturation. Labels can be trusted because the worlds are verified code: synthesized-then-checked dynamics are exact over 20-step rollouts and answer 10x out-of-distribution probes exactly (100%), whereas per-step LLM and MLP predictors compound error and collapse. Unlike domain randomization, each world is independently authored and verified; a corrupted-label control shows label exactness, not task variety, drives the gains. On real benchmarks (ARC-AGI grids, List Functions, CLRS) the same lever holds as per-world test-time training. On List Functions the harder cross-world form holds: one adapter trained on 128 disjoint worlds reaches 40% on held-out worlds versus 6% for a corrupted-label control (+34 points, CI [29, 39]). The gain is a saturating regularity, not a law: largest for few-step reasoning and small/weak models, fading for long chains, perception-induced tasks, and saturated tasks; cross-task transfer is weak without shared skill. Worlds are authored and served by OpenWorld, a zero-dependency framework (companion paper). Scope: symbolic state; pixel-native domains remain territory of learned models. All code, recipes, and this manuscript regenerate from one repository.

cs.LG

From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems

Agentic AI systems capable of generating full-stack web applications from natural language prompts ("prompt- to-app") represent a significant shift in software development. However, evaluating these systems remains challenging, as visual polish, functional correctness, and user trust are often misaligned. As a result, it is unclear how existing prompt-to-app tools compare under realistic, human-centered evaluation criteria. In this paper, we introduce a human-centered benchmark for evaluating prompt-to-app systems and conduct a large-scale comparative study of three widely used platforms: Replit, Bolt, and Firebase Studio. Using a diverse set of 96 prompts spanning common web application tasks, we generate 288 unique application artifacts. We evaluate these systems through a large-scale human-rater study involving 205 participants and 1,071 quality-filtered pairwise comparisons, assessing task-based ease of use, visual appeal, perceived completeness, and user trust. Our results show that these systems are not interchangeable: Firebase Studio consistently outperforms competing platforms across all human-evaluated dimensions, achieving the highest win rates for ease of use, trust, visual appeal, and visual appropriateness. Bolt performs competitively on visual appeal but trails Firebase on usability and trust, while Replit underperforms relative to both across most metrics. These findings highlight a persistent gap between visual polish and functional reliability in prompt-to-app systems and demonstrate the necessity of interactive, task-based evaluation. We release our benchmark framework, prompt set, and generated artifacts to support reproducible evaluation and future research in agentic application generation.

cs.HC

Mineral Detection of Neutrinos and Dark Matter. A Whitepaper

Minerals are solid state nuclear track detectors - nuclear recoils in a mineral leave latent damage to the crystal structure. Depending on the mineral and its temperature, the damage features are retained in the material from minutes (in low-melting point materials such as salts at a few hundred degrees C) to timescales much larger than the 4.5 Gyr-age of the Solar System (in refractory materials at room temperature). The damage features from the $O(50)$ MeV fission fragments left by spontaneous fission of $^{238}$U and other heavy unstable isotopes have long been used for fission track dating of geological samples. Laboratory studies have demonstrated the readout of defects caused by nuclear recoils with energies as small as $O(1)$ keV. This whitepaper discusses a wide range of possible applications of minerals as detectors for $E_R \gtrsim O(1)$ keV nuclear recoils: Using natural minerals, one could use the damage features accumulated over $O(10)$ Myr$-O(1)$ Gyr to measure astrophysical neutrino fluxes (from the Sun, supernovae, or cosmic rays interacting with the atmosphere) as well as search for Dark Matter. Using signals accumulated over months to few-years timescales in laboratory-manufactured minerals, one could measure reactor neutrinos or use them as Dark Matter detectors, potentially with directional sensitivity. Research groups in Europe, Asia, and America have started developing microscopy techniques to read out the $O(1) - O(100)$ nm damage features in crystals left by $O(0.1) - O(100)$ keV nuclear recoils. We report on the status and plans of these programs. The research program towards the realization of such detectors is highly interdisciplinary, combining geoscience, material science, applied and fundamental physics with techniques from quantum information and Artificial Intelligence.

astro-ph.IM