SearcharxivSearch

arXiv subjects

Ido Levy

Publications and source records attributed to Ido Levy.

At least 19 recordsLinked to original sources

Governance by Construction for Generalist Agents

Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by construction. Systems must specify which actions are allowed, when human oversight is required, and what information may be exposed, without rebuilding the agent for each domain. This demo presents CUGA's policy system, a modular policy-as-code layer that composes with a generalist LLM agent to deliver predictable, auditable, and compliance-aware behavior in compound workflows without model fine-tuning. We present a runtime governance architecture that enforces policy interventions at every critical stage of execution. Rather than passively constraining behavior, policies intercept the agent at five structural checkpoints: upstream of planning (Intent Guard), within the system prompt to steer reasoning (Playbook), at the tool-call boundary to enforce proper usage (Tool Guide), outside the reasoning loop as a Human-in-the-Loop gate for high-risk actions (Tool Approvals), and at the output stage to filter and structure the final response (Output Formatter). Together, these stages embed governance continuously across the agent's execution pipeline rather than treating it as an afterthought. Using a healthcare scenario and a multi-layered enforcement intervention, the demo shows dynamic playbook injection for structured tool-sequence enforcement, intent guards that block malicious or accidental harmful requests, and human-in-the-loop tool approval checkpoints for potentially destructive actions. The artifact illustrates how typed governance primitives enable faster, safer deployment of enterprise agentic systems while improving policy adherence and execution consistency.

cs.AI

TabAgent: A Framework for Replacing Agentic Generative Components with Tabular-Textual Classifiers

Agentic systems, AI architectures that autonomously execute multi-step workflows to achieve complex goals, are often built using repeated large language model (LLM) calls for closed-set decision tasks such as routing, shortlisting, gating, and verification. While convenient, this design makes deployments slow and expensive due to cumulative latency and token usage. We propose TabAgent, a framework for replacing generative decision components in closed-set selection tasks with a compact textual-tabular classifier trained on execution traces. TabAgent (i) extracts structured schema, state, and dependency features from trajectories (TabSchema), (ii) augments coverage with schema-aligned synthetic supervision (TabSynth), and (iii) scores candidates with a lightweight classifier (TabHead). On the long-horizon AppWorld benchmark, TabAgent maintains task-level success while eliminating shortlist-time LLM calls, reducing latency by approximately 95% and inference cost by 85-91%. Beyond tool shortlisting, TabAgent generalizes to other agentic decision heads, establishing a paradigm for learned discriminative replacements of generative bottlenecks in production agent architectures.

cs.CL

AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems

We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes fifteen failure-detection tools and two root-cause analysis modules that jointly uncover weaknesses across input handling, prompt design, and output generation. It integrates lightweight rule-based checks with LLM-as-a-judge assessments to support structured incident detection, classification, and repair. We applied the framework to IBM CUGA, evaluating its performance on the AppWorld and WebArena benchmarks. The analysis revealed recurrent planner misalignments, schema violations, brittle prompt dependencies, and more. Based on these insights, we refined both prompting and coding strategies, maintaining CUGA's benchmark results while enabling mid-sized models such as Llama 4 and Mistral Medium to achieve notable accuracy gains, substantially narrowing the gap with frontier models. Beyond quantitative validation, we conducted an exploratory study that fed the framework's diagnostic outputs and agent description into an LLM for self-reflection and prioritization. This interactive analysis produced actionable insights on recurring failure patterns and focus areas for improvement, demonstrating how validation itself can evolve into an agentic, dialogue-driven process. These results show a path toward scalable, quality assurance, and adaptive validation in production agentic systems, offering a foundation for more robust, interpretable, and self-improving agentic architectures.

cs.AI

Textual Planning with Explicit Latent Transitions

Planning with LLMs is bottlenecked by token-by-token generation and repeated full forward passes, making multi-step lookahead and rollout-based search expensive in latency and compute. We propose EmbedPlan, which replaces autoregressive next-state generation with a lightweight transition model operating in a frozen language embedding space. EmbedPlan encodes natural language state and action descriptions into vectors, predicts the next-state embedding, and retrieves the next state by nearest-neighbor similarity, enabling fast planning computation without fine-tuning the encoder. We evaluate next-state prediction across nine classical planning domains using six evaluation protocols of increasing difficulty: interpolation, plan-variant, extrapolation, multi-domain, cross-domain, and leave-one-out. Results show near-perfect interpolation performance but a sharp degradation when generalization requires transfer to unseen problems or unseen domains; plan-variant evaluation indicates generalization to alternative plans rather than memorizing seen trajectories. Overall, frozen embeddings support within-domain dynamics learning after observing a domain's transitions, while transfer across domain boundaries remains a bottleneck.

cs.CL

Hybrid superinductance with Al/InAs

We report microwave spectroscopy of Josephson junctions chains made from an epitaxial Al/InAs heterostructure. The chains exhibit superinductance, with characteristic wave impedance exceeding $R_{Q} = \hbar/(2e)^{2}$. The planar nature of the junctions results in a large plasma frequency, with no measurable deviations from ideal dispersion up to $12~\mathrm{GHz}$. Internal quality factors decrease sharply with frequency, which we describe with a simple loss model. The possibility of a loss mechanism intrinsic to the superconductor-semiconductor junction is considered.

cond-mat.mes-hall

From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production

Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business value. This path is complicated by fragmented frameworks, slow development, and the absence of standardized evaluation practices. Generalist agents have emerged as a promising direction, excelling on academic benchmarks and offering flexibility across task types, applications, and modalities. Yet, evidence of their use in production enterprise settings remains limited. This paper reports IBM's experience developing and piloting the Computer Using Generalist Agent (CUGA), which has been open-sourced for the community (https://github.com/cuga-project/cuga-agent). CUGA adopts a hierarchical planner--executor architecture with strong analytical foundations, achieving state-of-the-art performance on AppWorld and WebArena. Beyond benchmarks, it was evaluated in a pilot within the Business-Process-Outsourcing talent acquisition domain, addressing enterprise requirements for scalability, auditability, safety, and governance. To support assessment, we introduce BPO-TA, a 26-task benchmark spanning 13 analytics endpoints. In preliminary evaluations, CUGA approached the accuracy of specialized agents while indicating potential for reducing development time and cost. Our contribution is twofold: presenting early evidence of generalist agents operating at enterprise scale, and distilling technical and organizational lessons from this initial pilot. We outline requirements and next steps for advancing research-grade architectures like CUGA into robust, enterprise-ready systems.

cs.AI

Strongly anharmonic flux-tunable transmon based on InAs-Al 2D heterostructure

The gatemon qubits, made of transparent superconducting-semiconducting Josephson junctions, typically have even weaker anharmonicity than the opaque AlOx-junction transmons. However, flux-frustrated gatemons can acquire a much stronger anharmonicity, originating from the interference of the higher-order harmonics of the supercurrent. Here we investigate this effect of enhanced anharmonicity in split-junction gatemon devices based on InAs-Al 2D heterostructure. We find that anharmonicity in excess of 100% can be routinely achieved at the half-integer flux sweet-spot without any need for electrical gating or excessive sensitivity to the offset charge noise. We verified that such intrinsically large anharmonicity enables our devices to be driven coherently with raw Rabi frequencies exceeding 100 MHz, without any pulse shaping, simplifying implementation and control compared to traditional gatemons and transmons. Furthermore, by analyzing a relatively high-resolution spectroscopy of the device transitions as a function of flux, we were able to extract fine details of the current-phase relation, to which transport measurements would hardly be sensitive. The strong anharmonicity of our anharmonic tunable transmons, along with their bare-bones design, can prove to be a precious resource that transparent superconducting-semiconducting junctions bring to quantum information processing.

cond-mat.mes-hall

Towards Enterprise-Ready Computer Using Generalist Agent

This paper presents our ongoing work toward developing an enterprise-ready Computer Using Generalist Agent (CUGA) system. Our research highlights the evolutionary nature of building agentic systems suitable for enterprise environments. By integrating state-of-the-art agentic AI techniques with a systematic approach to iterative evaluation, analysis, and refinement, we have achieved rapid and cost-effective performance gains, notably reaching a new state-of-the-art performance on the WebArena and AppWorld benchmarks. We detail our development roadmap, the methodology and tools that facilitated rapid learning from failures and continuous system refinement, and discuss key lessons learned and future challenges for enterprise adoption.

cs.DC

Unsupervised Translation of Emergent Communication

Emergent Communication (EC) provides a unique window into the language systems that emerge autonomously when agents are trained to jointly achieve shared goals. However, it is difficult to interpret EC and evaluate its relationship with natural languages (NL). This study employs unsupervised neural machine translation (UNMT) techniques to decipher ECs formed during referential games with varying task complexities, influenced by the semantic diversity of the environment. Our findings demonstrate UNMT's potential to translate EC, illustrating that task complexity characterized by semantic diversity enhances EC translatability, while higher task complexity with constrained semantic variability exhibits pragmatic EC, which, although challenging to interpret, remains suitable for translation. This research marks the first attempt, to our knowledge, to translate EC without the aid of parallel data.

cs.CL

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprises can trust. To integrate these agents into critical workflows, safety and trustworthiness (ST) are prerequisite conditions for adoption. We introduce \textbf{\textsc{ST-WebAgentBench}}, a configurable and easily extensible suite for evaluating web agent ST across realistic enterprise scenarios. Each of its 222 tasks is paired with ST policies, concise rules that encode constraints, and is scored along six orthogonal dimensions (e.g., user consent, robustness). Beyond raw task success, we propose the \textit{Completion Under Policy} (\textit{CuP}) metric, which credits only completions that respect all applicable policies, and the \textit{Risk Ratio}, which quantifies ST breaches across dimensions. Evaluating three open state-of-the-art agents reveals that their average CuP is less than two-thirds of their nominal completion rate, exposing critical safety gaps. By releasing code, evaluation templates, and a policy-authoring interface, \href{https://sites.google.com/view/st-webagentbench/home}{\textsc{ST-WebAgentBench}} provides an actionable first step toward deploying trustworthy web agents at scale.

cs.AI

Machine learning analysis of structural data to predict electronic properties in near-surface InAs quantum wells

Semiconductor crosshatch patterns in thin film heterostructures form as a result of strain relaxation processes and dislocation pile-ups during growth of lattice mismatched materials. Due to their connection with the internal misfit dislocation network, these crosshatch patterns are a complex fingerprint of internal strain relaxation and growth anisotropy. Therefore, this mesoscopic fingerprint not only describes the residual strain state of a near-surface quantum well, but also could provide an indicator of the quality of electron transport through the material. Here, we present a method utilizing computer vision and machine learning to analyze AFM crosshatch patterns that exhibits this correlation. Our analysis reveals optimized electron transport for moderate values of $\lambda$ (crosshatch wavelength) and $\epsilon$ (crosshatch height), roughly 1 $\mu$m and 4 nm, respectively, that define the average waveform of the pattern. Simulated 2D AFM crosshatch patterns are used to train a machine learning model to correlate the crosshatch patterns to dislocation density. Furthermore, this model is used to evaluate the experimental AFM images and predict a dislocation density based on the crosshatch waveform. Predicted dislocation density, experimental AFM crosshatch data, and experimental transport characterization are used to train a final model to predict 2D electron gas mean free path. This model shows electron scattering is strongly correlated with elastic effects (e.g. dislocation scattering) below 200 nm $\lambda_{MFP}$.

cond-mat.mes-hall

From Grounding to Planning: Benchmarking Bottlenecks in Web Agents

General web-based agents are increasingly essential for interacting with complex web environments, yet their performance in real-world web applications remains poor, yielding extremely low accuracy even with state-of-the-art frontier models. We observe that these agents can be decomposed into two primary components: Planning and Grounding. Yet, most existing research treats these agents as black boxes, focusing on end-to-end evaluations which hinder meaningful improvements. We sharpen the distinction between the planning and grounding components and conduct a novel analysis by refining experiments on the Mind2Web dataset. Our work proposes a new benchmark for each of the components separately, identifying the bottlenecks and pain points that limit agent performance. Contrary to prevalent assumptions, our findings suggest that grounding is not a significant bottleneck and can be effectively addressed with current techniques. Instead, the primary challenge lies in the planning component, which is the main source of performance degradation. Through this analysis, we offer new insights and demonstrate practical suggestions for improving the capabilities of web agents, paving the way for more reliable agents.

cs.AI

Microwave Andreev bound state spectroscopy in a semiconductor-based Planar Josephson junction

By coupling a semiconductor-based planar Josephson junction to a superconducting resonator, we investigate the Andreev bound states in the junction using dispersive readout techniques. Using electrostatic gating to create a narrow constriction in the junction, our measurements unveil a strong coupling interaction between the resonator and the Andreev bound states. This enables the mapping of isolated tunable Andreev bound states, with an observed transparency of up to 99.94\% along with an average induced superconducting gap of $\sim 150 \mu$eV. Exploring the gate parameter space further elucidates a non-monotonic evolution of multiple Andreev bound states with varying gate voltage. Complimentary tight-binding calculations of an Al-InAs planar Josephson junction with strong Rashba spin-orbit coupling provide insight into possible mechanisms responsible for such behavior. Our findings highlight the subtleties of the Andreev spectrum of Josephson junctions fabricated on superconductor-semiconductor heterostructures and offering potential applications in probing topological states in these hybrid platforms.

cond-mat.mes-hall

The infrastructure powering IBM's Gen AI model development

AI Infrastructure plays a key role in the speed and cost-competitiveness of developing and deploying advanced AI models. The current demand for powerful AI infrastructure for model training is driven by the emergence of generative AI and foundational models, where on occasion thousands of GPUs must cooperate on a single training job for the model to be trained in a reasonable time. Delivering efficient and high-performing AI training requires an end-to-end solution that combines hardware, software and holistic telemetry to cater for multiple types of AI workloads. In this report, we describe IBM's hybrid cloud infrastructure that powers our generative AI model development. This infrastructure includes (1) Vela: an AI-optimized supercomputing capability directly integrated into the IBM Cloud, delivering scalable, dynamic, multi-tenant and geographically distributed infrastructure for large-scale model training and other AI workflow steps and (2) Blue Vela: a large-scale, purpose-built, on-premises hosting environment that is optimized to support our largest and most ambitious AI model training tasks. Vela provides IBM with the dual benefit of high performance for internal use along with the flexibility to adapt to an evolving commercial landscape. Blue Vela provides us with the benefits of rapid development of our largest and most ambitious models, as well as future-proofing against the evolving model landscape in the industry. Taken together, they provide IBM with the ability to rapidly innovate in the development of both AI models and commercial offerings.

cs.DC

Gatemonium: A Voltage-Tunable Fluxonium

We present a new style of fluxonium qubit, gatemonium, based on an all superconductorsemiconductor hybrid platform. The linear inductance is achieved using six hundred planar Al-InAs Josephson junctions (JJs) in series. By tuning the single junction with a gate voltage, we demonstrate electrostatic control of the effective Josephson energy, tuning the weight of the fictitious phase particle. One and two-tone spectroscopy of the gatemonium transitions further reveal details of the hybrid plasmon-fluxon spectrum. Accounting for the nonsinusoidal current-phase relation of the single junction, we fit the measured spectra to extract charging and inductive energies. We conduct time domain characterization of the plasmon modes in a second gatemonium device with different charging energy and JJ array inductance. We discuss future directions for this platform in gate voltage-tunable, high plasma frequency, enhanced impedance junction arrays, and enhanced coherence times for voltage tunable architectures.

cond-mat.mes-hall

Molecular beam epitaxy growth of superconducting tantalum germanide

Developing new material platforms for use in superconductor-semiconductor hybrid structures is desirable due to limitations caused by intrinsic microwave losses present in commonly used III/V material systems. With the recent reports on tantalum superconducting qubits that show improvements over the Nb and Al counterparts, exploring Ta as an alternative superconductor in hybrid material systems is promising. Here, we study the growth of Ta on semiconducting Ge (001) substrates grown via molecular beam epitaxy. We show that at a growth temperature of 400$^{\circ}$C the Ta diffuses into the Ge matrix in a self-limiting nature resulting in smooth and abrupt surfaces and interfaces with roughness on the order of 3-7 \r{A} as measured by atomic force microscopy and x-ray reflectivity. The films are found to be a mixture of Ta$_{5}$Ge$_{3}$ and TaGe$_{2}$ binary alloys and form a native oxide that seems to form a sharp interface with the underlying film. These films are superconducting with a $T_{C}\sim 1.8-2$K and $H_{C}^{\perp} \sim 1.88T$, $H_{C}^{\parallel} \sim 5.1T$. These results show this tantalum germanide film to be promising for future superconducting quantum information platforms.

cond-mat.mes-hall

Structural and magnetic properties of molecular beam epitaxy (MnSb2Te4)x(Sb2Te3)1-x topological materials with exceedingly high Curie temperature

Tuning magnetic properties of magnetic topological materials is of interest to realize elusive physical phenomena such as quantum anomalous hall effect (QAHE) at higher temperatures and design topological spintronic devices. However, current topological materials exhibit Curie temperature (TC) values far below room temperature. In recent years, significant progress has been made to control and optimize TC, particularly through defect engineering of these structures. Most recently we showed evidence of TC values up to 80K for (MnSb2Te4)x(Sb2Te3)1-x, where x is greater than or equal to 0.7 and less than or equal to 0.85, by controlling the compositions and Mn content in these structures. Here we show further enhancement of the TC, as high as 100K, by maintaining high Mn content and reducing the growth rate from 0.9 nm/min to 0.5 nm/min. Derivative curves reveal the presence of two TC components contributing to the overall value and propose TC1 and TC2 have distinct origins: excess Mn in SLs and Mn in Sb2-yMnyTe3QLs alloys, respectively. In pursuit of elucidating the mechanisms promoting higher Curie temperature values in this system, we show evidence of structural disorder where Mn is occupying not only Sb sites but also Te sites, providing evidence of significant excess Mn and a new crystal structure:(Mn1+ySb2-yTe4)x(Sb2-yMnyTe3)1-x. Our work shows progress in understanding how to control magnetic defects to enhance desired magnetic properties and the mechanism promoting these high TC in magnetic topological materials such as (Mn1+ySb2-yTe4)x(Sb2-yMnyTe3)1-x.

cond-mat.mtrl-sci

Characterizing losses in InAs two-dimensional electron gas-based gatemon qubits

The tunnelling of cooper pairs across a Josephson junction (JJ) allow for the nonlinear inductance necessary to construct superconducting qubits, amplifiers, and various other quantum circuits. An alternative approach using hybrid superconductor-semiconductor JJs can enable superconducting qubit architectures with all electric control. Here we present continuous-wave and time-domain characterization of gatemon qubits and coplanar waveguide resonators based on an InAs two-dimensional electron gas. We show that the qubit undergoes a vacuum Rabi splitting with a readout cavity and we drive coherent Rabi oscillations between the qubit ground and first excited states. We measure qubit relaxation times to be $T_1 =$ 100 ns over a 1.5 GHz tunable band. We detail the loss mechanisms present in these materials through a systematic study of the quality factors of coplanar waveguide resonators. While various loss mechanisms are present in III-V gatemon circuits we detail future directions in enhancing the relaxation times of qubit devices on this platform.

quant-ph