SearcharxivSearch

arXiv subjects

Yanwei Huang

Publications and source records attributed to Yanwei Huang.

10 recordsLinked to original sources

KEditVis: A Visual Analytics System for Knowledge Editing of Large Language Models

Large Language Models (LLMs) demonstrate exceptional capabilities in factual question answering, yet they sometimes provide incorrect responses. To address this issue, knowledge editing techniques have emerged as effective methods for correcting factual information in LLMs. However, typical knowledge editing workflows struggle with identifying the optimal set of model layers for editing and rely on summary indicators that provide insufficient guidance. This lack of transparency hinders effective comparison and identification of optimal editing strategies. In this paper, we present KEditVis, a novel visual analytics system designed to assist users in gaining a deeper understanding of knowledge editing through interactive visualizations, improving editing outcomes, and discovering valuable insights for the future development of knowledge editing algorithms. With KEditVis, users can select appropriate layers as the editing target, explore the reasons behind ineffective edits, and perform more targeted and effective edits. Our evaluation, including usage scenarios, expert interviews, and a user study, validates the effectiveness and usability of the system.

cs.HC

Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring

LLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge, such as dataset schemas, domain conventions, and task-specific requirements, that isn't explicitly provided. This results in frequently erroneous scripts that require users to repeatedly clarify their intent. Additionally, users struggle to validate generated scripts because they cannot verify whether the model correctly applied implicit knowledge. We present Cerebra, an interactive NL-to-SQL tool that aligns implicit knowledge between users and LLMs during SQL authoring. Cerebra automatically retrieves implicit knowledge from historical SQL scripts based on user instructions, presents this knowledge in an interactive tree view for code review, and supports iterative refinement to improve generated scripts. To evaluate the effectiveness and usability of Cerebra, we conducted a user study with 16 participants, demonstrating its improved support for customized SQL authoring. The source code of Cerebra is available at https://github.com/zjuidg/CHI26-Cerebra.

cs.HC

Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeek

Web AI agents such as ChatGPT Agent and GenSpark are increasingly used for routine web-based tasks, yet they still rely on text-based input prompts, lack proactive detection of user intent, and offer no support for interactive data analysis and decision making. We present WebSeek, a mixed-initiative browser extension that enables users to discover and extract information from webpages to then flexibly build, transform, and refine tangible data artifacts-such as tables, lists, and visualizations-all within an interactive canvas. Within this environment, users can perform analysis-including data transformations such as joining tables or creating visualizations-while an in-built AI both proactively offers context-aware guidance and automation, and reactively responds to explicit user requests. An exploratory user study (N=15) with WebSeek as a probe reveals participants' diverse analysis strategies, underscoring their desire for transparency and control during human-AI collaboration.

cs.HC

Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI

Despite their increasing capabilities, text-to-image generative AI systems are known to produce biased, offensive, and otherwise problematic outputs. While recent advancements have supported testing and auditing of generative AI, existing auditing methods still face challenges in supporting effectively explore the vast space of AI-generated outputs in a structured way. To address this gap, we conducted formative studies with five AI auditors and synthesized five design goals for supporting systematic AI audits. Based on these insights, we developed Vipera, an interactive auditing interface that employs multiple visual cues including a scene graph to facilitate image sensemaking and inspire auditors to explore and hierarchically organize the auditing criteria. Additionally, Vipera leverages LLM-powered suggestions to facilitate exploration of unexplored auditing directions. Through a controlled experiment with 24 participants experienced in AI auditing, we demonstrate Vipera's effectiveness in helping auditors navigate large AI output spaces and organize their analyses while engaging with diverse criteria.

cs.HC

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language models (LLMs). While recent data augmentation techniques have facilitated the creation of large-scale benchmarks, existing LLM-generated datasets often suffer from limited reliability, diversity, and scalability. To address these challenges, we introduce PuzzleClone, a formal framework for synthesizing verifiable data at scale using a novel DSL-driven approach. Our approach features three key innovations: (1) encoding seed puzzles into structured logical specifications, (2) generating scalable variants through systematic variable and constraint randomization, and (3) ensuring validity via a reproduction mechanism. Applying PuzzleClone, we construct PC-83K, a benchmark comprising over 83K diverse and programmatically validated puzzles. The generated puzzles span a wide spectrum of difficulty and formats, posing significant challenges to current state-of-the-art models. Experimental results show that post training (SFT and RL) on PC-83K yields substantial improvements not only on the testset but also on various logic and mathematical benchmarks. Post training raises average performance on PC-83K from 14.5 to 66.0 and delivers consistent improvements across 7 logic and mathematical benchmarks up to 18.4 absolute percentage points (SATBench from 51.6 to 70.0). Our code and data are available at https://github.com/HiThink-Research/PuzzleClone.

cs.AI

Vipera: Towards systematic auditing of generative text-to-image models at scale

Generative text-to-image (T2I) models are known for their risks related such as bias, offense, and misinformation. Current AI auditing methods face challenges in scalability and thoroughness, and it is even more challenging to enable auditors to explore the auditing space in a structural and effective way. Vipera employs multiple visual cues including a scene graph to facilitate image collection sensemaking and inspire auditors to explore and hierarchically organize the auditing criteria. Additionally, it leverages LLM-powered suggestions to facilitate exploration of unexplored auditing directions. An observational user study demonstrates Vipera's effectiveness in helping auditors organize their analyses while engaging with diverse criteria.

cs.HC

StructVizor: Interactive Profiling of Semi-Structured Textual Data

Data profiling plays a critical role in understanding the structure of complex datasets and supporting numerous downstream tasks, such as social media analytics and financial fraud detection. While existing research predominantly focuses on structured data formats, a substantial portion of semi-structured textual data still requires ad-hoc and arduous manual profiling to extract and comprehend its internal structures. In this work, we propose StructVizor, an interactive profiling system that facilitates sensemaking and transformation of semi-structured textual data. Our tool mainly addresses two challenges: a) extracting and visualizing the diverse structural patterns within data, such as how information is organized or related, and b) enabling users to efficiently perform various wrangling operations on textual data. Through automatic data parsing and structure mining, StructVizor enables visual analytics of structural patterns, while incorporating novel interactions to enable profile-based data wrangling. A comparative user study involving 12 participants demonstrates the system's usability and its effectiveness in supporting exploratory data analysis and transformation tasks.

cs.HC

Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts

Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring.

cs.HC

How to obtain complex transition dipole moments satisfying crystal symmetry and periodicity from ab-initio calculations

Transition dipole moments (TDM) between energy bands of solids deserve special attention nowadays as intense lasers can easily drive non-adiabatic transitions of excited electron wave packets across the Brillouin zones. The TDM is required to be continuous, satisfying crystal symmetry, and periodic at zone boundaries. While present day ab-initio algorithms are powerful in calculating band structures of solids, they all introduced random phases into the eigenfunctions at each crystal momentum k. In this paper, we show how to choose a ``smooth-periodic'' gauge where TDMs can be smooth versus k, preserving crystal symmetry, as well as maintaining periodic at zone boundaries. Based on band structure and TDMs in the ``smooth-periodic'' gauge calculated from ab-initio algorithms, we revisit high-order harmonic generation from MgO which exhibits inversion symmetry and ZnO which has broken symmetry. The symmetry properties of TDMs with respect to k ensure the absence of even-order harmonics in system with inversion symmetry, while the TDM in the `smooth-periodic'' gauge for ZnO is shown to enhance even harmonics that were underestimated in previous simulations. These results reveal the importance of correctly treating the complex TDMs in nonlinear laser-solid interactions which has been elusive so far.

physics.optics

Dissociation products and structures of solid H2S at strong compression

Hydrogen sulfides have recently received a great deal of interest due to the record high superconducting temperatures of up to 203 K observed on strong compression of dihydrogen sulfide (H2S). A joint theoretical and experimental study is presented in which decomposition products and structures of compressed H2S are characterized, and their superconducting properties are calculated. In addition to the experimentally known H2S and H3S phases, our first-principles structure searches have identified several energetically competitive stoichiometries that have not been reported previously; H2S3, H3S2, and H4S3. In particular, H4S3 is predicted to be thermodynamically stable within a large pressure range of 25-113 GPa. High-pressure room-temperature X-ray diffraction measurements confirm the presence of H3S and H4S3 through decomposition of H2S that emerge at 27 GPa and coexist with residual H2S, at least up to the highest pressure studied in our experiments of 140 GPa. Electron-phonon coupling calculations show that H4S3 has a small Tc of below 2 K, and that H2S is mainly responsible for the observed superconductivity of samples prepared at low temperature (<100K).

cond-mat.supr-con