SearcharxivSearch

arXiv subjects

William Wong

Publications and source records attributed to William Wong.

7 recordsLinked to original sources

Computer Use at the Edge of the Statistical Precipice

Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naive aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, and multifactorial variability) and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework pairing Wilson score intervals with hierarchical bootstrap, producing confidence intervals that correctly account for the nested structure of CUA benchmarks, as we empirically demonstrate. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.

cs.SE

Detecting misinformation through Framing Theory: the Frame Element-based Model

In this paper, we delve into the rapidly evolving challenge of misinformation detection, with a specific focus on the nuanced manipulation of narrative frames - an under-explored area within the AI community. The potential for Generative AI models to generate misleading narratives underscores the urgency of this problem. Drawing from communication and framing theories, we posit that the presentation or 'framing' of accurate information can dramatically alter its interpretation, potentially leading to misinformation. We highlight this issue through real-world examples, demonstrating how shifts in narrative frames can transmute fact-based information into misinformation. To tackle this challenge, we propose an innovative approach leveraging the power of pre-trained Large Language Models and deep neural networks to detect misinformation originating from accurate facts portrayed under different frames. These advanced AI techniques offer unprecedented capabilities in identifying complex patterns within unstructured data critical for examining the subtleties of narrative frames. The objective of this paper is to bridge a significant research gap in the AI domain, providing valuable insights and methodologies for tackling framing-induced misinformation, thus contributing to the advancement of responsible and trustworthy AI technologies. Several experiments are intensively conducted and experimental results explicitly demonstrate the various impact of elements of framing theory proving the rationale of applying framing theory to increase the performance in misinformation detection.

cs.CL

Optimizing Industrial HVAC Systems with Hierarchical Reinforcement Learning

Reinforcement learning (RL) techniques have been developed to optimize industrial cooling systems, offering substantial energy savings compared to traditional heuristic policies. A major challenge in industrial control involves learning behaviors that are feasible in the real world due to machinery constraints. For example, certain actions can only be executed every few hours while other actions can be taken more frequently. Without extensive reward engineering and experimentation, an RL agent may not learn realistic operation of machinery. To address this, we use hierarchical reinforcement learning with multiple agents that control subsets of actions according to their operation time scales. Our hierarchical approach achieves energy savings over existing baselines while maintaining constraints such as operating chillers within safe bounds in a simulated HVAC control environment.

cs.LG

Irreducible representations of the symmetric groups from slash homologies of p-complexes

In the 40s, Mayer introduced a construction of (simplicial) $p$-complex by using the unsigned boundary map and taking coefficients of chains modulo $p$. We look at such a $p$-complex associated to an $(n-1)$-simplex; in which case, this is also a $p$-complex of representations of the symmetric group of rank $n$ - specifically, of permutation modules associated to two-row compositions. In this article, we calculate the so-called slash homology - a homology theory introduced by Khovanov and Qi - of such a $p$-complex. We show that every non-trivial slash homology group appears as an irreducible representation associated to two-row partitions, and how this calculation leads to a basis of these irreducible representations given by the so-called $p$-standard tableaux.

math.RT

Perverse equivalence in $SL(2,q)$

We will be defining a type of perverse equivalence that always corresponds to a derived equivalence with two-term tilting complexes. We are going to show that the tilting considered by Okuyama and Yoshii for the proof of Broué's conjecture for $SL(2,q)$ in defining characteristic is a composition of such perverse equivalences.

math.RT

Thin-film transistor electrical performance of hybrid MoS 2 -P3HT semiconductor layers

The hole carrier field-effect mobility of hybrid molybdenum disulfide (MoS2) nanoparticles suspended in poly(3-hexylthiophene) (P3HT) thin film transistor (TFT) was found to be enhanced when it compared to P3HT-only TFTs. The improvement in the hole charge transport was found to be a function of the concentration of MoS2 in P3HT with high MoS2 concentrations resulting in an increase in the on-current of the device. Moreover, Au has a high work function of 5.1 eV which is suitable with the HOMO level of P3HT. We find that MoS2 and Au have the proper energy level for hole transport and injection.

physics.app-ph

Some perverse equivalences of $SL(2,q)$ in its defining characteristic

In this article, we study the modular representations of the special linear group of degree two over a finite field in defining characteristic. In particular, we study the automorphisms of derived category of representations. We have been able to obtain a new type of autoequivalence. This autoequivalence has some uncommon features. It is more conveniently conceived and proved using the representation theory of its Brauer correspondence but at the same time it can be very neatly described, using a type of derived equivalence called perverse equivalence, in global settings.

math.RT