SearcharxivSearch

arXiv subjects

Thanh Nguyen

Publications and source records attributed to Thanh Nguyen.

At least 19 recordsLinked to original sources

REST API Testing with Verified LLM-Inferred Dependencies and Response-Driven Refinement

Testing RESTful APIs requires generating sequences of API calls that satisfy dependencies among operations, parameters, and runtime-created resources. Recent LLM-based approaches infer such dependencies and generate test sequences from OpenAPI specifications, but they often treat LLM-inferred relationships as correct without execution-based validation. This can introduce spurious dependencies, miss feasible operation chains, and produce infeasible tests. In this paper, we propose APIPilot}, an execution-validated framework for REST API testing. APIPilot first derives candidate producer-consumer dependencies from OpenAPI specifications using structural heuristics and LLM-based semantic reasoning. It then treats these dependencies as hypotheses and validates them through concrete API executions before using them for test generation. The validated dependencies are organized into a dependency graph from which APIPilot constructs coverage-aware workflows via bounded top-k graph traversal, separating semantic dependency inference from sequence construction. To improve subsequent tests, APIPilot further performs response-driven refinement: runtime responses are analyzed to update resource pools, adjust input-generation constraints, and prune or revise invalid dependency mappings. Empirical evaluation on 16 real-world REST API services shows that APIPilot achieves 92.3% operation coverage, up to 58.6% code coverage, and an 88.1% workflow execution success rate, outperforming both LLM-based and traditional REST API testing baselines. APIPilot also detects 197 unique 5xx failures and specification-execution mismatches, demonstrating the benefit of grounding dependency inference in execution feedback.

cs.SE

Deterministic single-electron trapping on solid neon using engineered dielectric surface geometry

Levitating electron qubit on the surface of solid neon has recently emerged as a promising and intrinsically noise-resilient platform for quantum information processing. Their ultra-clean, inert environment suppresses conventional decoherence pathways associated with lattice disorder, charge traps, and nuclear-spin baths that limit coherence in semiconductor qubits. Yet, uncontrolled surface features such as bumps, valleys, and electrode-defined gaps can bind electrons unintentionally, contributing charge noise and inducing spin-orbit coupling mediated decoherence. To address this challenge, we propose an engineered interface in which a dielectric layer is deposited beneath the solid neon to provide an atomically smooth template, eliminating surface-roughness induced trapping. By selectively etching this dielectric layer at desired qubit locations, deterministic potential minima can be engineered to reliably capture electrons while suppressing unwanted surface bound states. We perform large-scale Schrodinger and Poisson simulation to compare the existing and proposed strategies of electron trapping on neon, obtaining good agreement with recent experimental measurements.

quant-ph

Kani: A Model Checker for Rust

Rust's ownership type system prevents memory errors in safe code, but certain desirable properties remain orthogonal to compilation: the soundness of unsafe operations (e.g., raw pointer dereferences), functional correctness, and absence of runtime panics. We present Kani, an open-source model checker for Rust that pushes bounded model checking beyond bug-finding to provide correctness guarantees for these properties. Kani compiles proof harnesses from Rust's Mid-level Intermediate Representation (MIR) into CBMC's bit-precise verification engine, automatically checking a comprehensive set of safety properties with no user annotation. To extend verification from bounded to unbounded, Kani provides a specification language comprising function contracts, loop contracts, quantifiers, and function stubbing. We demonstrate feasibility through case studies on industrial Rust projects, where contracts upgraded verification from panic-freedom to functional correctness, uncovering six previously unknown bugs. Kani operates at scale in production CI, with over 16,000 harnesses verified per code change in the Rust standard library verification campaign.

cs.SE

Verifying the Rust Standard Library

Rust's type system prevents many classes of memory errors, yet its standard library relies heavily on unsafe code whose correctness is validated through testing, including dynamic checks under Miri, but lacks static verification. We present what is, to the best of our knowledge, the largest verification campaign reported for a software library: an open, crowdsourced effort that integrates complementary verification tools into the continuous integration of a verification repository forked from the Rust standard library. We analyze the campaign's effectiveness, discuss the practical value of machine-checked proofs for a subset of undefined behaviors (e.g., out-of-bounds access, null and dangling pointer dereferences, and use of uninitialized memory), and frame the remaining obstacles as open challenges for the formal-methods community.

cs.LO

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning

Diffusion-based Q-learning has emerged as a powerful paradigm for offline reinforcement learning, but its reliance on multi-step denoising makes both training and inference computationally expensive and brittle. Recent efforts to accelerate diffusion Q-learning toward single-step action generation typically introduce auxiliary networks, policy distillation, or multi-phase training, which frequently compromise simplicity, stability, or performance. To address these limitations, we introduce Bootstrapped Flow Q-Learning (BFQ), a novel framework that enables accurate single-step action generation during both training and inference, without auxiliary networks or distillation procedures. BFQ adopts a divide-and-conquer view of the displacement vector along the flow path: it begins by learning short-range displacements that can be accurately estimated from the Flow Matching marginal velocity, and bootstraps these components to directly learn a noise-to-action mapping in a single step. This formulation eliminates multi-step denoising, resulting in a learning procedure that is substantially faster, simpler, and more robust. Extensive D4RL evaluations show that BFQ improves performance while significantly reducing computational cost compared to multi-step diffusion baselines, demonstrating that single-step action generation suffices for high-performance offline Reinforcement Learning.

cs.LG

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, how good these systems are actually, especially compared to human reviewers at catching scientific gaps, remains poorly understood. In this work, we introduce PRISM (Peer Review Intelligence via Structured Multi-dimensional assessment), a benchmarking framework that evaluates review quality across four dimensions: Depth of Analysis, Novelty Assessment,Flaw Identification & Major Issues Prioritization, and Multi-dimensional Constructiveness. Unlike most existing evaluations based on surface-level metrics like ROUGE and BLEU, or unconstrained LLM-as-a-judge prompting that conflates fluency with rigor, PRISM grounds each dimension in argument mining, retrieval-augmented verification, and consensus-based scoring. We apply PRISM to benchmark five leading automated reviewer systems and human reviewers on a stratified corpus of reviews from ICLR, ICML, and NeurIPS. The results reveal that LLMs can match or beat human reviewers on individual dimensions: comparable depth of analysis, stronger novelty verification, and highly accurate critique prioritization. However, no single system consistently matches the balanced performance of the human baseline across all dimensions at once. Each exhibits a distinct specialization profile with characteristic blind spots -- failure modes that aggregate metrics miss entirely. The implication is that LLM reviewers are best understood as targeted supplements to human review, effective within specific dimensions, but unreliable as standalone replacements. Our demo and key results can be found at https://khanhthanhdev.github.io/prism-page/.

cs.CL

Large Transverse Thermoelectric Effect in Weyl Semimetal TaIrTe$_4$ Engineered for Photodetection

Anomalous local photocurrent generation via second-order nonlinear and thermoelectric responses is a signature of many topological semimetals. The emergence of these photocurrents is inherently linked to symmetry breaking and anisotropy of their crystal lattices. Studies of type-II Weyl semimetals of group C$_{2v}$ (WTe$_2$, MoTe$_2$, TaIrTe$_4$) have reported anomalous, nonlocal photocurrents localized to crystals edges or far from electrodes, which are highly dependent on the geometry of the material sample. While originally attributed to a nonlinear charge current response, it was recently shown that these currents could instead be attributed to the anisotropic Seebeck coefficients of the materials. Here, we confirm that anomalous photocurrents observed in TaIrTe$_4$ under either visible or far-infrared far-field illumination originate from the large transverse thermoelectric effect. We engineer the mutual orientation of crystal edges and electrodes as well as the thermal environment of TaIrTe$_4$ to control and amplify its spatial photocurrent response. We show that substrate engineering can locally enhance photocurrent. This framework of thermal device engineering can enable broadband photo detection schemes by leveraging spectral and spatial dependence of photocurrents for applications like wavefront sensing, beam positioning, and edge detection.

cond-mat.mtrl-sci

Probing Non-Equilibrium Grain Boundary Dynamics with XPCS and Domain-Adaptive Machine Learning

Grain-boundary (GB) dynamics control the stability, mechanical, and functional response of nanocrystalline materials, but direct experimental access to their slow non-equilibrium motion has been limited. Here we establish X-ray photon correlation spectroscopy (XPCS), combined with domain-adaptive machine learning, as a quantitative probe of GB dynamics. Temperature- and grain-size-dependent two-time XPCS measurements in nanocrystalline silicon reveal pronounced departures from time-translation invariance, showing that GB relaxation can remain far from equilibrium over experimental timescales. However, direct extraction of quantitative physical information from these high-dimensional, noisy fluctuation maps faces a significant challenge. To overcome this barrier, we develop a semi-supervised learning framework that transfers physical parameter labels from continuum simulations to unlabeled experimental XPCS maps through domain-adaptive representation alignment. This AI-augmented approach enables the extraction of key kinetic parameters, including bulk diffusivity, GB stiffness, and effective GB concentration, directly from experimental XPCS measurements. Our results show how machine learning can transform indirect fluctuation signals into quantitative materials dynamics, providing a general route to study non-equilibrium defect motion in solids.

cond-mat.mtrl-sci

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substantial Functional Requirement Documents (FRDs) and repository context to be ingested in a single workflow, making specification quality the execution fuel for autonomous delivery. This report formalises Spec-Driven Agentic Development (SDAD) as a synthesis of disciplined up-front formalisation and high-velocity implementation: intent capture, machine-readable specification, agentic synthesis, and independent multi-agent verification under human sign-off. We revisit the historical pendulum between Waterfall and Agile, introduce AI-code as a fourth production paradigm, and compare Human-Agile (circa 2020) with Agentic-SDAD (circa 2026) across artefacts, cadence, accountability, and security posture. Beyond process description, we extend the model to team role metamorphosis (engineer, QA, platform, and product functions), quantitative governance (Ambiguity Tax, Spec Fidelity, SER, and TCI_agentic with repair multiplier phi), and pragmatic adoption via hybrid estimation and a staged migration blueprint. Industrial and research evidence on AI-augmented testing and verification is integrated to motivate separation between synthesis and release authority. Overall, the paper argues that agentic speed does not eliminate engineering discipline; it relocates discipline upstream into specification precision, explicit gates, and auditable provenance.

cs.AI

Ordinal Lindahl Equilibrium for Voting

The core is a central concept in multi-winner social choice, ensuring that no coalition of voters can support an alternative outcome whose size or cost exceeds the group's share of the electorate. This idea originates from the Lindahl equilibrium in classical public goods theory. Yet Lindahl equilibria may fail to exist when voters have ordinal preferences over a finite set of outcomes and monetary transfers are not allowed. We introduce Lindahl Equilibrium with Ordinal Preferences (LEO), extending the equilibrium framework to discrete collective choice. Using LEO, we construct randomized outcomes that satisfy (approximate) core constraints for a probabilistic set of voters, while ensuring that each voter is represented with high probability. We also provide a deterministic approximate core guarantee with a factor of 6.24, improving on the previous bound of 32. In structured environments, these outcomes can be computed efficiently. Overall, our results extend classical equilibrium concepts, providing a normative foundation for proportional representation and practical algorithms for applications in voting and fair machine learning.

cs.GT

One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning

Diffusion Q-Learning (DQL) has established diffusion policies as a high-performing paradigm for offline reinforcement learning, but its reliance on multi-step denoising for action generation renders both training and inference slow and fragile. Existing efforts to accelerate DQL toward one-step denoising typically rely on auxiliary modules or policy distillation, sacrificing either simplicity or performance. It remains unclear whether a one-step policy can be trained directly without such trade-offs. To this end, we introduce One-Step Flow Q-Learning (OFQL), a novel framework that enables effective one-step action generation during both training and inference, without auxiliary modules or distillation. OFQL reformulates the DQL policy within the Flow Matching (FM) paradigm but departs from conventional FM by learning an average velocity field that directly supports accurate one-step action generation. This design removes the need for multi-step denoising and backpropagation-through-time updates, resulting in substantially faster and more robust learning. Extensive experiments on the D4RL benchmark show that OFQL, despite generating actions in a single step, not only significantly reduces computation during both training and inference but also outperforms multi-step DQL by a large margin. Furthermore, OFQL surpasses all other baselines, achieving state-of-the-art performance in D4RL.

cs.LG

Experimental and computational studies of the hydrogenation of carbon disulfide (CS2) on ice analogues

Carbon disulfide (CS$_2$) is one of the sulfur-bearing species expected to be present in the interstellar medium (ISM). In this study, we investigated the surface reactions of solid CS$_2$ with hydrogen (H) atoms on amorphous solid water (ASW) using laboratory experiments supported by computational calculations. Our results show that CS$_2$ reacts with H atoms through quantum tunneling in the initial step, followed by successive H addition reactions, with or without activation barriers, on icy surfaces. These processes lead to the formation of several sulfur-bearing species, including hydrogen sulfide (H$_2$S), methyl mercaptan (CH$_3$SH), and small amounts of dithioformic acid (HC(S)SH) and methanedithiol (CH$_2$(SH)$_2$). The observed reactivity of CS$_2$ with H atoms provides a plausible explanation for the non-detection of CS$_2$ in interstellar ices. Furthermore, the efficient hydrogenation of the complex molecules derived from CS$_2$, namely HC(S)SH and CH$_2$(SH)$_2$, suggests that these species could be easily undergone with H atoms to produce other S-bearing species under ISM conditions.

astro-ph.GA

Automatic Identification of Compounds in Molecular Mixtures from Liquid-Phase Infrared Spectra

Interpreting spectroscopy data is a critical bottleneck in automating chemical research and industrial characterization. Particularly within infrared (IR) spectroscopy, identifying compounds in complex, liquid-phase chemical mixtures largely relies on expert knowledge, as variable peak assignment, broadening, and shifts hinder data-driven methods. Here, we show that an algorithmic approach can identify components in both simulated and experimental mixture spectra with high accuracy despite nonlinearities in liquid-phase IR data. The method is comprehensively benchmarked with a dataset of over 44,000 simulated liquid-phase IR spectra for mixtures and achieves up to 90% accuracy in identifying molecular components across a dataset of binary and ternary liquid mixtures. Our strategy is robust to perturbation of spectra, and its accuracy is capped by near-identical liquid-phase IR spectra that limit the resolution of chemical identification, imposing theoretical limits on achieving perfect accuracy in structure identification. Finally, we apply the method to automatically interpret IR spectra in experimental settings, correctly identifying the components of nearly all samples within a blind study. This work provides tools and data to advance automated chemical laboratories through algorithmic interpretation of liquid-phase IR spectra of mixtures.

cond-mat.mtrl-sci

Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning

Offline reinforcement learning (RL) has garnered significant interest due to its safe and easily scalable paradigm. However, training under this paradigm presents its own challenge: the extrapolation error stemming from out-of-distribution (OOD) data. Existing methodologies have endeavored to address this issue through means like penalizing OOD Q-values or imposing similarity constraints on the learned policy and the behavior policy. Nonetheless, these approaches are often beset by limitations such as being overly conservative in utilizing OOD data, imprecise OOD data characterization, and significant computational overhead. To address these challenges, this paper introduces an Uncertainty-Aware Rank-One Multi-Input Multi-Output (MIMO) Q Network framework. The framework aims to enhance Offline Reinforcement Learning by fully leveraging the potential of OOD data while still ensuring efficiency in the learning process. Specifically, the framework quantifies data uncertainty and harnesses it in the training losses, aiming to train a policy that maximizes the lower confidence bound of the corresponding Q-function. Furthermore, a Rank-One MIMO architecture is introduced to model the uncertainty-aware Q-function, \TP{offering the same ability for uncertainty quantification as an ensemble of networks but with a cost nearly equivalent to that of a single network}. Consequently, this framework strikes a harmonious balance between precision, speed, and memory efficiency, culminating in improved overall performance. Extensive experimentation on the D4RL benchmark demonstrates that the framework attains state-of-the-art performance while remaining computationally efficient. By incorporating the concept of uncertainty quantification, our framework offers a promising avenue to alleviate extrapolation errors and enhance the efficiency of offline RL.

cs.LG

Matching with Committee Preferences

We study a many-to-one matching model inspired by school choice, where schools evaluate applicants using multiple rankings rather than a single priority order. We model each school's evaluation with social choice criteria to reflect the school's internal ranking process. In particular, we define acceptable choices as candidates ranked above a top percentile of the accepted cohort by a sufficient number of evaluators. Stability is then defined in terms of acceptability: accepted candidates must receive strong support, while rejected candidates receive at most weak support. Since exact acceptability and stability may not exist, we construct approximately stable outcomes using a new equilibrium concept that combines matching with a Lindahl equilibrium over ordinal preferences, providing a flexible, equilibrium-based framework for committee-based matching markets.

cs.GT

Systematic Trend-Following with Adaptive Portfolio Construction: Enhancing Risk-Adjusted Alpha in Cryptocurrency Markets

Cryptocurrency markets exhibit pronounced momentum effects and regime-dependent volatility, presenting both opportunities and challenges for systematic trading strategies. We propose AdaptiveTrend, a multi-component algorithmic trading framework that integrates high-frequency trend-following on 6-hour intervals with monthly adaptive portfolio construction and asymmetric long-short capital allocation. Our framework introduces three key innovations: (1) a dynamic trailing stop mechanism calibrated to intra-day volatility regimes, (2) a rolling Sharpe-ratio-based asset selection procedure with market-capitalization-aware filtering, and (3) a theoretically motivated asymmetric 70/30 long-short allocation scheme grounded in the empirical positive drift of crypto markets. Through extensive out-of-sample backtesting across 150+ cryptocurrency pairs over a 36-month evaluation window (2022-2024), AdaptiveTrend achieves an annualized Sharpe ratio of 2.41, a maximum drawdown of -12.7%, and a Calmar ratio of 3.18, significantly outperforming benchmark trend-following strategies (TSMOM, time-series momentum) and equal-weighted buy-and-hold portfolios. We further conduct rigorous robustness analyses including parameter sensitivity, transaction cost modeling, and regime-conditional performance decomposition, demonstrating the strategy's resilience across bull, bear, and sideways market conditions.

cs.CE

Sharing with Frictions: Limited Transfers and Costly Inspections

The radio spectrum suitable for commercial wireless services is limited. A portion of the radio spectrum has been reserved for institutions using it for non-commercial purposes such as federal agencies, defense, public safety bodies and scientific institutions. In order to operate efficiently, these incumbents need clean spectrum access. However, commercial users also want access, and granting them access may materially interfere with the existing activity of the incumbents. Conventional market based mechanisms for allocating scarce resources in this context are problematic. Allowing direct monetary transfers to and from public or scientific institutions risks distorting their non-commercial mission. Moreover, often only the incumbent knows the exact value of the interference it experiences, and, likewise, only commercial users can predict accurately the expected monetary outcome from sharing the resource. Thus, our problem is to determine the efficient allocation of resources in the presence of private information without the use of direct monetary transfers. The problem is not unique to spectrum. Other resources that governments hold in trust share the same feature. We propose a novel mechanism design formulation of the problem, characterize the optimal mechanism and describe some of its qualitative properties.

econ.TH

Practitioner Insights on Fairness Requirements in the AI Development Life Cycle: An Interview Study

Nowadays, Artificial Intelligence (AI), particularly Machine Learning (ML) and Large Language Models (LLMs), is widely applied across various contexts. However, the corresponding models often operate as black boxes, leading them to unintentionally act unfairly towards different demographic groups. This has led to a growing focus on fairness in AI software recently, alongside the traditional focus on the effectiveness of AI models. Through 26 semi-structured interviews with practitioners from different application domains and with varied backgrounds across 23 countries, we conducted research on fairness requirements in AI from software engineering perspective. Our study assesses the participants' awareness of fairness in AI / ML software and its application within the Software Development Life Cycle (SDLC), from translating fairness concerns into requirements to assessing their arising early in the SDLC. It also examines fairness through the key assessment dimensions of implementation, validation, evaluation, and how it is balanced with trade-offs involving other priorities, such as addressing all the software functionalities and meeting critical delivery deadlines. Findings of our thematic qualitative analysis show that while our participants recognize the aforementioned AI fairness dimensions, practices are inconsistent, and fairness is often deprioritized with noticeable knowledge gaps. This highlights the need for agreement with relevant stakeholders on well-defined, contextually appropriate fairness definitions, the corresponding evaluation metrics, and formalized processes to better integrate fairness into AI/ML projects.

cs.SE