SearcharxivSearch

arXiv subjects

Richard Kang

Publications and source records attributed to Richard Kang.

8 recordsLinked to original sources

Registry-Governed Agent Lifecycle:Completing EDDOps with Evaluation-DrivenRegistration, Promotion, and Retirement on AWS AgentCore

Enterprise adoption of LLM agents requires model selection methods that balance quality, reliability, safety, latency, and cost. Evaluation-Driven Development and Operations (EDDOps) positions evaluation as a continuous governing function across the agent lifecycle rather than a terminal checkpoint. This paper presents a practitioner-oriented instantiation of EDDOps on AWS Bedrock AgentCore and proposes a cost-to-performance framework for selecting foundation models in enterprise agent architectures. We make three contributions: a conceptual synthesis explaining why traditional TDD/BDD methods are insufficient for non-deterministic LLM agents; an architectural mapping of the EDDOps reference architecture onto AgentCore Runtime, Evaluations, Agent Registry, and CloudWatch observability; and an empirical cost-to-performance decision framework validated through a proof-of-concept comparing three foundation models across two deployment paths. Using trace data from 30 single-turn invocations across six agents, 9 multi-turn evaluations, and registry-integrated governance, we show how evaluation evidence can convert model selection from a benchmark-ranking exercise into a governed economic decision. The results suggest that managed agent platforms can support EDDOps when they provide trace-native observability, pluggable evaluator frameworks, and governed registry-based discovery.

cs.SE

Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express

Agent interoperability protocols (MCP, A2A, ACP, ANP, and ERC-8004) have rapidly matured to enable identity, capability discovery, tool access, and message exchange between autonomous agents. However, as enterprises deploy heterogeneous agent fleets that must make collective decisions under governance constraints, a question arises: can these protocols support governed agent communities, or only task-oriented coordination? We present a systematic gap analysis applying a six-dimension governance requirements taxonomy (membership, deliberation, voting, dissent preservation, human escalation, and audit/replay) derived from organizational theory, multi-agent systems literature, and enterprise governance standards. We analyze each protocol's specification against this taxonomy, classifying capabilities as Supported, Partial, or Absent. The resulting gap matrix reveals that voting and dissent preservation are universally absent across all five protocols, deliberation is absent or at most partial, and no protocol encodes the full set of primitives required for governed agent communities. We distinguish extensible gaps (addressable through protocol extension mechanisms) from structural gaps (requiring a new architectural layer) and assess time-sensitivity based on observed protocol evolution velocity. The analysis establishes that agent community governance constitutes a missing architectural layer above current interoperability standards, not a missing feature within them.

cs.MA

Governed AI-Assisted Engineering: Graduated Human Oversight for Agentic Code Generation in Regulated Domains

The adoption of agentic AI coding systems -- where autonomous agents generate, review, test, and deploy code with minimal human intervention -- creates a governance challenge in regulated industries. Existing frameworks address AI-assisted development maturity or the productivity-reliability tension but offer no mechanism for calibrating human oversight intensity to regulatory impact. We present the Governed AI-Assisted Engineering (GAIE) framework, a three-tier graduated human oversight model for agentic code generation in regulated domains. GAIE introduces the Oversight Classification Model (OCM), a deterministic decision function that classifies code generation tasks by regulatory impact, customer proximity, reversibility, and data sensitivity to route them through one of three oversight tiers: human-in-the-loop (strategic functions), human-over-the-loop (customer-impacting), or automated-with-monitoring (internal). Each tier defines required evidence artifacts for compliance auditability. We map GAIE against the Bank of Thailand's 2025 AI risk-management policy and demonstrate cross-jurisdiction applicability to MAS (Singapore), NIST AI RMF, ISO/IEC 42001, and the EU AI Act. Evaluation through regulatory coverage analysis, comparative framework analysis, and analytical productivity modeling suggests that graduated oversight preserves 84--97% of agentic coding velocity (central estimate: 91%) while maintaining compliance evidence coverage for regulated functions. GAIE contributes a framework that explicitly bridges AI-assisted development maturity with regulatory governance through proportionate human oversight.

cs.HC

Extending orbital-optimized density functional theory to L-edge XPS and beyond: Spin-orbit coupling via non-orthogonal quasi-degenerate perturbation theory

Quantum mechanical calculations of core electron binding energies (CEBEs) leading to 2p hole states are relevant to interpreting L-edge x-ray photo-electron spectroscopy (XPS), as well as higher edges. Orbital-optimized density functional theory (OO-DFT) accurately predicts K-edge CEBEs but is challenged by the presence of significant spin-orbit coupling (SOC) at L- and higher edges. To extend OO-DFT to L-edges and higher, our method utilizes scalar-relativistic, spin-restricted OO-DFT to construct a minimal, quasi-degenerate basis of core-hole states corresponding to a chosen inner-shell (e.g. ionizing all six possible 2p spin orbitals). Non-orthogonal configuration interaction (NOCI) is then used to make the matrix elements of the full Hamiltonian including SOC in this quasi-degenerate model space of determinants. Using a screened 1-electron SOC operator parametrized with the Dirac-Coulomb-Breit (DCB) Hamiltonian results in doublet splitting (DS) values for 3rd row atoms that are nearly in quantitative agreement with experiment. The resulting NOCI eigenvalues are shifted by the average of the (scalar) OO-DFT CEBEs to yield CEBEs (split by SOC) corrected for dynamic correlation. Comparing calculations on gas phase molecules with experimental results establishes that NO-QDPT with the SCAN functional (NO-QDPT/SCAN), using the DCB screened 1-electron SOC operator is accurate to about 0.2 eV for L-edge CEBEs of molecules containing 3rd row atoms. However, this NO-QDPT approach becomes less accurate for 4th-row elements starting in the middle of the 3d transition metal series, especially as the atomic number increases.

physics.chem-ph

Moonwalk: Advancing Gait-Based User Recognition on Wearable Devices with Metric Learning

Personal devices have adopted diverse authentication methods, including biometric recognition and passcodes. In contrast, headphones have limited input mechanisms, depending solely on the authentication of connected devices. We present Moonwalk, a novel method for passive user recognition utilizing the built-in headphone accelerometer. Our approach centers on gait recognition; enabling users to establish their identity simply by walking for a brief interval, despite the sensor's placement away from the feet. We employ self-supervised metric learning to train a model that yields a highly discriminative representation of a user's 3D acceleration, with no retraining required. We tested our method in a study involving 50 participants, achieving an average F1 score of 92.9% and equal error rate of 2.3%. We extend our evaluation by assessing performance under various conditions (e.g. shoe types and surfaces). We discuss the opportunities and challenges these variations introduce and propose new directions for advancing passive authentication for wearable devices.

cs.HC

Vision-Based Hand Gesture Customization from a Single Demonstration

Hand gesture recognition is becoming a more prevalent mode of human-computer interaction, especially as cameras proliferate across everyday devices. Despite continued progress in this field, gesture customization is often underexplored. Customization is crucial since it enables users to define and demonstrate gestures that are more natural, memorable, and accessible. However, customization requires efficient usage of user-provided data. We introduce a method that enables users to easily design bespoke gestures with a monocular camera from one demonstration. We employ transformers and meta-learning techniques to address few-shot learning challenges. Unlike prior work, our method supports any combination of one-handed, two-handed, static, and dynamic gestures, including different viewpoints, and the ability to handle irrelevant hand movements. We implement three real-world applications using our customization method, conduct a user study, and achieve up to 94% average recognition accuracy from one demonstration. Our work provides a viable path for vision-based gesture customization, laying the foundation for future advancements in this domain.

cs.HC

Advancing Location-Invariant and Device-Agnostic Motion Activity Recognition on Wearable Devices

Wearable sensors have permeated into people's lives, ushering impactful applications in interactive systems and activity recognition. However, practitioners face significant obstacles when dealing with sensing heterogeneities, requiring custom models for different platforms. In this paper, we conduct a comprehensive evaluation of the generalizability of motion models across sensor locations. Our analysis highlights this challenge and identifies key on-body locations for building location-invariant models that can be integrated on any device. For this, we introduce the largest multi-location activity dataset (N=50, 200 cumulative hours), which we make publicly available. We also present deployable on-device motion models reaching 91.41% frame-level F1-score from a single model irrespective of sensor placements. Lastly, we investigate cross-location data synthesis, aiming to alleviate the laborious data collection tasks by synthesizing data in one location given data from another. These contributions advance our vision of low-barrier, location-invariant activity recognition systems, catalyzing research in HCI and ubiquitous computing.

cs.HC

Relativistic Orbital Optimized Density Functional Theory for Accurate Core-Level Spectroscopy

Core-level spectra of 1s electrons of elements heavier than Ne show significant relativistic effects. We combine advances in orbital optimized DFT (OO-DFT) with the spin-free exact two-component (X2C) model for scalar relativistic effects, to study K-edge spectra of third period elements. OO-DFT/X2C is found to be quite accurate at predicting energies, yielding $\sim 0.5$ eV RMS error vs experiment with the modern SCAN (and related) functionals. This marks a significant improvement over the $>50$ eV deviations that are typical for the popular time-dependent DFT (TDDFT) approach. Consequently, experimental spectra are quite well reproduced by OO-DFT/X2C, sans empirical shifts for alignment. OO-DFT/X2C combines high accuracy with ground state DFT cost and is thus a promising route for computing core-level spectra of third period elements. We also explored K and L edges of 3d transition metals to identify limitations of the OO-DFT/X2C approach in modeling the spectra of heavier atoms.

physics.chem-ph