SearcharxivSearch

arXiv subjects

Abhinav Rao

Publications and source records attributed to Abhinav Rao.

9 recordsLinked to original sources

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be reversed through limited realignment. We systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representations throughout training. Although we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with apparent rapid realignment largely disappearing after controlling for response-length differences. We further find that previously reported mechanistic signatures, including representational phase transitions in LoRA space, do not consistently correlate with behavioral misalignment across training. Our results suggest that current evidence for EM is less robust than previously claimed and highlight the need for evaluation protocols that carefully control for these surface level dataset artifacts to identify the robustness of the EM phenomenon.

cs.CL

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We introduce NormAd, an evaluation framework to assess LLMs' cultural adaptability, specifically measuring their ability to judge social acceptability across varying levels of cultural norm specificity, from abstract values to explicit social norms. As an instantiation of our framework, we create NormAd-Eti, a benchmark of 2.6k situational descriptions representing social-etiquette related cultural norms from 75 countries. Through comprehensive experiments on NormAd-Eti, we find that LLMs struggle to accurately judge social acceptability across these varying degrees of cultural contexts and show stronger adaptability to English-centric cultures over those from the Global South. Even in the simplest setting where the relevant social norms are provided, the best LLMs' performance (< 82\%) lags behind humans (> 95\%). In settings with abstract values and country information, model performance drops substantially (< 60\%), while human accuracy remains high (> 90\%). Furthermore, we find that models are better at recognizing socially acceptable versus unacceptable situations. Our findings showcase the current pitfalls in socio-cultural reasoning of LLMs which hinder their adaptability for global audiences.

cs.CL

Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin

This paper presents the work of restoring punctuation for ASR transcripts generated by multilingual ASR systems. The focus languages are English, Mandarin, and Malay which are three of the most popular languages in Singapore. To the best of our knowledge, this is the first system that can tackle punctuation restoration for these three languages simultaneously. Traditional approaches usually treat the task as a sequential labeling task, however, this work adopts a slot-filling approach that predicts the presence and type of punctuation marks at each word boundary. The approach is similar to the Masked-Language Model approach employed during the pre-training stages of BERT, but instead of predicting the masked word, our model predicts masked punctuation. Additionally, we find that using Jieba1 instead of only using the built-in SentencePiece tokenizer of XLM-R can significantly improve the performance of punctuating Mandarin transcripts. Experimental results on English and Mandarin IWSLT2022 datasets and Malay News show that the proposed approach achieved state-of-the-art results for Mandarin with 73.8% F1-score while maintaining a reasonable F1-score for English and Malay, i.e. 74.7% and 78% respectively. Our source code that allows reproducing the results and building a simple web-based application for demonstration purposes is available on Github.

cs.CL

[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs

We introduce two paradoxes concerning jailbreak of foundation models: First, it is impossible to construct a perfect jailbreak classifier, and second, a weaker model cannot consistently detect whether a stronger (in a pareto-dominant sense) model is jailbroken or not. We provide formal proofs for these paradoxes and a short case study on Llama and GPT4-o to demonstrate this. We discuss broader theoretical and practical repercussions of these results.

cs.CL

Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these attacks and their mitigations. We bridge this gap by proposing a formalism and a taxonomy of known (and possible) jailbreaks. We survey existing jailbreak methods and their effectiveness on open-source and commercial LLMs (such as GPT-based models, OPT, BLOOM, and FLAN-T5-XXL). We further discuss the challenges of jailbreak detection in terms of their effectiveness against known attacks. For further analysis, we release a dataset of model outputs across 3700 jailbreak prompts over 4 tasks.

cs.CL

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a global scale. When provided with an ethical policy, an LLM should be capable of making decisions that are ethically consistent to the policy. We develop a framework that integrates moral dilemmas with moral principles pertaining to different foramlisms of normative ethics, and at different levels of abstractions. Initial experiments with GPT-x models shows that while GPT-4 is a nearly perfect ethical reasoner, the models still have bias towards the moral values of Western and English speaking societies.

cs.CL

MALITE: Lightweight Malware Detection and Classification for Constrained Devices

Today, malware is one of the primary cyberthreats to organizations. Malware has pervaded almost every type of computing device including the ones having limited memory, battery and computation power such as mobile phones, tablets and embedded devices like Internet-of-Things (IoT) devices. Consequently, the privacy and security of the malware infected systems and devices have been heavily jeopardized. In recent years, researchers have leveraged machine learning based strategies for malware detection and classification. Malware analysis approaches can only be employed in resource constrained environments if the methods are lightweight in nature. In this paper, we present MALITE, a lightweight malware analysis system, that can classify various malware families and distinguish between benign and malicious binaries. MALITE converts a binary into a gray scale or an RGB image and employs low memory and battery power consuming as well as computationally inexpensive malware analysis strategies. We have designed MALITE-MN, a lightweight neural network based architecture and MALITE-HRF, an ultra lightweight random forest based method that uses histogram features extracted by a sliding window. We evaluate the performance of both on six publicly available datasets (Malimg, Microsoft BIG, Dumpware10, MOTIF, Drebin and CICAndMal2017), and compare them to four state-of-the-art malware classification techniques. The results show that MALITE-MN and MALITE-HRF not only accurately identify and classify malware but also respectively consume several orders of magnitude lower resources (in terms of both memory as well as computation capabilities), making them much more suitable for resource constrained environments.

cs.CR

Printable, castable, nanocrystalline cellulose-epoxy composites exhibiting hierarchical nacre-like toughening

Due to their exceptional mechanical and chemical properties and their natural abundance, cellulose nanocrystals (CNCs) are promising building blocks of sustainable polymer composites. However, the rapid gelation of CNC dispersions has generally limited CNC-based composites to low CNC fractions, in which polymer remains the dominant phase. Here we report on the formulation and processing of crosslinked CNC-epoxy composites with a CNC fraction exceeding 50 wt.%. The microstructure comprises sub-micrometer aggregates of CNCs crosslinked to polymer, which are analogous to the lamellar structure of nacre and promotes toughening mechanisms associated with bulk ductile behavior, despite the brittle behavior of the aggregates at the nanoscale. At 63 wt.% CNCs, the composites exhibit a hardness of 0.66 GPa and a fracture toughness of 5.2 MPa.m$^{1/2}$. The hardness of this all-organic material is comparable to aluminum alloys, and the fracture toughness at the centimeter scale is comparable to that of wood cell wall. We show that CNC-epoxy composite objects can be shaped from the gel precursors by direct-write printing and by casting, while the cured composites can be machined into complex 3D shapes. The formulation, processing route, and the insights on toughening mechanisms gained from our multiscale approach can be applied broadly to highly loaded nanocomposites.

physics.app-ph

Shear Melting and Recovery of Crosslinkable Cellulose Nanocrystal-Polymer Gels

Cellulose nanocrystals (CNC) are naturally-derived nanostructures of growing importance for the production of composites having attractive mechanical properties, and offer improved sustainability over purely petroleum-based alternatives. Fabrication of CNC composites typically involves extrusion of CNC suspensions and gels in a variety of solvents, in the presence of additives such as polymers and curing agents. However, most studies so far have focused on aqueous CNC gels, yet the behavior of CNC-polymer gels in organic solvents is important to their wider processability. Here, we study the rheological behavior of composite polymer-CNC gels in dimethylformamide, which include additives for both UV and thermal crosslinking. Using rheometry coupled with in-situ infrared spectroscopy, we show that under external shear, CNC-polymer gels display progressive and irreversible failure of the hydrogen bond network that is responsible for their pronounced elastic properties. In the absence of cross-linking additives, the polymer-CNC gels show negligible recovery upon cessation of flow, while the presence of additives allows the gels to recover via van der Waals interactions. By exploring a broad range of shear history and CNC concentrations, we construct master curves for the temporal evolution of the viscoelastic properties of the polymer-CNC gels, illustrating universality of the observed dynamics with respect to gel composition and flow conditions. We therefore find that polymer-CNC composite gels display a number of the distinctive features of colloidal glasses and, strikingly, that their response to the flow conditions encountered during processing can be tuned by chemical additives. These findings have implications for processing of dense CNC-polymer composites in solvent casting, 3D printing, and other manufacturing techniques.

cond-mat.soft