SearcharxivSearch

arXiv subjects

Aobo Yang

Publications and source records attributed to Aobo Yang.

11 recordsLinked to original sources

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make claims about a closed model? We evaluate surrogate fidelity at the prediction, attribution, and representation levels. For binary classification tasks, log-odds provide an API-compatible scalar readout of the model's representation space, and leave-one-out attributions provide insight into model behavior. Across eleven models spanning four families (Llama, Qwen, GPT, and Gemini), we find that prediction fidelity substantially overstates attribution fidelity: models that agree on what the answer is often disagree on why. We document an access-validity inversion: white-box signals like attention patterns and perturbation magnitudes are highly stable across models but only weakly predictive of causal attributions, which black-box input ablations capture by design. Mechanistic insight does not automatically transfer to closed targets, and prediction-level agreement is insufficient to warrant such transfer. Code and results are available at https://github.com/facebookresearch/surrogate.

cs.LG

Dflow-SUR: Enhancing Generative Aerodynamic Inverse Design using Differentiation Throughout Flow Matching

Generative inverse design requires incorporating physical constraints to ensure that generated designs are both reliable and accurate. However, we observe that current state-of-the-art energy-based methods suffer from an asynchronous phenomenon, where the optimization of the physical loss is constrained by the flow matching inference process. To overcome this limitation, we introduce Dflow-SUR, a differentiation strategy that separates the optimization of the physical loss from the flow matching inference. Compared to the most advanced energy-based baseline, Dflow-SUR achieves a reduction in physical loss by four orders of magnitude, while also cutting wall-clock time by 74% on the airfoil case. Additionally, it increases the mean lift-to-drag ratio by 11.8% over traditional Latin-hypercube sampling in wing design. Beyond improvements in accuracy and efficiency, Dflow-SUR offers three additional practical advantages: (i) enhanced control over guidance, (ii) lower surrogate uncertainty, and (iii) greater robustness to hyper-parameter tuning. Together, these results demonstrate that Dflow-SUR is a highly promising framework, providing both scalability and high fidelity for generative aerodynamic design.

cs.CE

Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning

Large Language Models (LLMs) exhibit impressive capabilities but often hallucinate, confidently providing incorrect answers instead of admitting ignorance. Prior work has shown that models encode linear representations of their own knowledge and that activation steering can reduce hallucinations. These approaches, however, require real-time monitoring and intervention during inference. We introduce Contrastive Activation Steering for Amortized Learning (CASAL), an efficient algorithm that connects interpretability with amortized optimization. CASAL directly bakes the benefits of activation steering into model's weights. Once trained, LLMs answer questions they know while abstaining from answering those they do not. CASAL's light-weight design requires training only a submodule of a single transformer layer and yet reduces hallucination by 30%-40% across multiple short-form QA benchmarks. CASAL is 30x more compute-efficient and 20x more data-efficient than strong LoRA-based baselines such as SFT and DPO, boosting its practical applicability in data scarce domains. Importantly, CASAL also generalizes effectively to out-of-distribution (OOD) domains. We showcase CASAL's flexibility in mitigating hallucinations in both text-only and vision-language models. To our knowledge, CASAL is the first steering-based training method that has been shown to be effective for both dense and Mixture-of-Experts (MoE) models. CASAL represents a promising step forward for applying interpretability-inspired method for practical deployment in production systems.

cs.CL

FuncGenFoil: Airfoil Generation and Editing Model in Function Space

Aircraft manufacturing is the jewel in the crown of industry, in which generating high-fidelity airfoil geometries with controllable and editable representations remains a fundamental challenge. Existing deep learning methods, which typically rely on predefined parametric representations (e.g., B\'ezier) or discrete point sets, face an inherent trade-off between expressive power and resolution adaptability. To tackle this challenge, we introduce FuncGenFoil, a novel function-space generative model that directly reconstructs airfoil geometries as function curves. Our method inherits the advantages of arbitrary-resolution sampling and smoothness from parametric functions, as well as the strong expressiveness of discrete point-based representations. Empirical evaluations demonstrate that FuncGenFoil improves upon state-of-the-art methods in airfoil generation, achieving a relative 74.4% reduction in label error and a 23.2% increase in diversity on the AF-200K dataset. Our results highlight the advantages of function-space modeling for aerodynamic shape optimization, offering a powerful and flexible framework for high-fidelity airfoil design.

cs.LG

Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks

Recent advancements in Large Language Models (LLMs) have sparked widespread concerns about their safety. Recent work demonstrates that safety alignment of LLMs can be easily removed by fine-tuning with a few adversarially chosen instruction-following examples, i.e., fine-tuning attacks. We take a further step to understand fine-tuning attacks in multilingual LLMs. We first discover cross-lingual generalization of fine-tuning attacks: using a few adversarially chosen instruction-following examples in one language, multilingual LLMs can also be easily compromised (e.g., multilingual LLMs fail to refuse harmful prompts in other languages). Motivated by this finding, we hypothesize that safety-related information is language-agnostic and propose a new method termed Safety Information Localization (SIL) to identify the safety-related information in the model parameter space. Through SIL, we validate this hypothesis and find that only changing 20% of weight parameters in fine-tuning attacks can break safety alignment across all languages. Furthermore, we provide evidence to the alternative pathways hypothesis for why freezing safety-related parameters does not prevent fine-tuning attacks, and we demonstrate that our attack vector can still jailbreak LLMs adapted to new languages.

cs.CL

The Llama 3 Herd of Models

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.

cs.AI

Using Captum to Explain Generative Language Models

Captum is a comprehensive library for model explainability in PyTorch, offering a range of methods from the interpretability literature to enhance users' understanding of PyTorch models. In this paper, we introduce new features in Captum that are specifically designed to analyze the behavior of generative language models. We provide an overview of the available functionalities and example applications of their potential for understanding learned associations within generative language models.

cs.CL

A Smart Adaptively Reconfigurable DC Battery for Higher Efficiency of Electric Vehicle Drive Trains

This paper proposes a drive train topology with a dynamically reconfigurable dc battery, which breaks hard-wired batteries into smaller subunits. It can rapidly control the output voltage and even contribute to voltage shaping of the inverter. Based upon the rapid development of low-voltage transistors and modular circuit topologies in the recent years, the proposed technology uses recent 48 V power electronics to achieve higher-voltage output and reduce losses in electric vehicle (EV) drive trains. The fast switching capability and low loss of low-voltage field effect transistors (FET) allow sharing the modulation with the main drive inverter. As such, the slower insulated-gate bipolar transistors (IGBT) of the inverter can operate at ideal duty cycle and aggressively reduced switching, while the adaptive dc battery provides an adjustable voltage and all common-mode contributions at the dc link with lower loss. Up to 2/3 of the switching of the main inverter is avoided. At high drive speeds and thus large modulation indices, the proposed converter halves the total loss compared to using the inverter alone; at lower speeds and thus smaller modulation indices, the advantage is even more prominent because of the dynamically lowered dc-link voltage. Furthermore, it can substantially reduce the distortion, particularly at lower modulation indices, e.g., down to 1/2 compared to conventional space-vector modulation and even 1/3 for discontinuous pulse-width modulation with hard-wired battery. Other benefits include alleviated insulation stress for motor windings, active battery balancing, and eliminating the vulnerability of large hard-wired battery packs to weak cells. We demonstrate the proposed motor drive on a 3-kW setup with eight battery modules.

eess.SY

Comparative Explanations of Recommendations

As recommendation is essentially a comparative (or ranking) process, a good explanation should illustrate to users why an item is believed to be better than another, i.e., comparative explanations about the recommended items. Ideally, after reading the explanations, a user should reach the same ranking of items as the system's. Unfortunately, little research attention has yet been paid on such comparative explanations. In this work, we develop an extract-and-refine architecture to explain the relative comparisons among a set of ranked items from a recommender system. For each recommended item, we first extract one sentence from its associated reviews that best suits the desired comparison against a set of reference items. Then this extracted sentence is further articulated with respect to the target user through a generative model to better explain why the item is recommended. We design a new explanation quality metric based on BLEU to guide the end-to-end training of the extraction and refinement components, which avoids generation of generic content. Extensive offline evaluations on two large recommendation benchmark datasets and serious user studies against an array of state-of-the-art explainable recommendation algorithms demonstrate the necessity of comparative explanations and the effectiveness of our solution.

cs.IR

Study on S2 Flow Path Design and Three-dimensional Numerical Simulation Parameter Calibration in Axial Compressor

Aerodynamic design process of multi - stage axial flow compressor usually uses the way that combines the S2 flow design and three-dimensional CFD numerical simulation analysis. Based on Mr. Wu Zhonghua's " Three-dimensional flow theory ", aiming at the S2 flow design matching parameters and the three-dimensional CFD numerical simulation data, through autonomous programming, the S2 design parameter distribution and the corresponding CGNS format CFD calculation results are extracted. Then make the comparative analysis of the two and provide the modification suggestion of the design. The examples have been tested by the comparison and correction of the eight-stage axial flow compressor calculation and finally improve the design performance of the compressor design. The design adiabatic efficiency increases by 0.5%. The surge margin increases by more than 5%. The validity and feasibility of the method are verified.

math.OC

Explanation as a Defense of Recommendation

Textual explanations have proved to help improve user satisfaction on machine-made recommendations. However, current mainstream solutions loosely connect the learning of explanation with the learning of recommendation: for example, they are often separately modeled as rating prediction and content generation tasks. In this work, we propose to strengthen their connection by enforcing the idea of sentiment alignment between a recommendation and its corresponding explanation. At training time, the two learning tasks are joined by a latent sentiment vector, which is encoded by the recommendation module and used to make word choices for explanation generation. At both training and inference time, the explanation module is required to generate explanation text that matches sentiment predicted by the recommendation module. Extensive experiments demonstrate our solution outperforms a rich set of baselines in both recommendation and explanation tasks, especially on the improved quality of its generated explanations. More importantly, our user studies confirm our generated explanations help users better recognize the differences between recommended items and understand why an item is recommended.

cs.IR