SearcharxivSearch

arXiv subjects

Kaile Wang

Publications and source records attributed to Kaile Wang.

At least 19 recordsLinked to original sources

An X-ray Absorption-Line Survey of the Circumgalactic Medium of M31 with XMM-Newton

X-ray absorption line spectroscopy provides a powerful method for probing the extended hot circumgalactic medium (CGM) of individual galaxies. In this study, we present the first survey of \osv and \oeit absorption lines toward the hot halo of M\,31, using $15$ and $18$ background active galactic nuclei, with impact parameters ranging from $R_{\rm imp}\sim300$--$730$~kpc and $\sim190$--$730$~kpc, respectively. We find a marginal excess absorption above the expected Milky Way (MW) foreground toward the innermost sightlines, consistent with an additional hot-CGM contribution associated with M\,31. Comparing sightlines inside and outside $R_{\rm imp}=300$--$350$~kpc, the excess corresponds to a mean equivalent width of $\sim4$--$9$~m\AA\ for the inner \osv sightlines at $R_{\rm imp}\sim310$~kpc and $\sim17$~m\AA\ for the inner \oeit sightlines at $R_{\rm imp}\sim200$~kpc, with a nominal combined significance level of $2.2$--$2.3\,\sigma$. The hot-CGM mass inferred from this excess is highly model dependent, especially on the assumed CGM boundary, and therefore does not yet provide a robust baryon census. Attributing the excess entirely to hot gas confined within approximately the virial radius of M\,31 would require a CGM mass exceeding the nominal ``missing'' baryon budget. The inferred mass decreases to a few $\times10^{11}~M_\odot$ when the assumed CGM boundary is extended to $400-500$~kpc. Deeper observations, both through longer exposures of existing sightlines and the inclusion of additional inner-halo targets, will be essential to robustly constrain the properties of the hot CGM of M\,31.

astro-ph.GA

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse. Current safety evaluations, however, operate in natural language and cannot determine whether a model-generated amino acid sequence is biological gibberish or a computational risk signal. To address this evaluation blind spot, we introduce SPIKE-Bench, coupling 631 curated toxin-design prompts across seven functional categories with the SPIKE funnel, a three-stage protocol that filters output through compliance, biological plausibility, and predicted toxicity, producing stage-level diagnostics and an aggregate function-aware metric: the Functional Harmfulness Rate (FHR). An audit of 32 LLMs reveals that most models freely comply with toxin-design requests; FHR is driven primarily by biological generation capability rather than safety alignment, reaching 50.7%; and Refusal Rate fails to predict functional risk. As a first step toward mitigation, we provide BioSafe-Guard, a domain-specialized classifier that substantially reduces predicted functional risk while preserving benign utility. We release SPIKE-Bench and BioSafe-Guard at https://github.com/PKU-Alignment/SPIKE-Bench to support more rigorous biosecurity evaluation of LLMs.

q-bio.QM

A Galactic Perspective on the (Unremarkable) Relative Refractory Depletion Observed in the Sun

Over the last two decades, the Sun has been observed to be depleted in refractory elements as a function of elemental condensation temperature (\tcond) relative to $\sim80\%$ of its counterparts. We assess the impact of Galactic chemical evolution (GCE) on refractory element--\tcond\ trends for 109,500 unique solar analogs from the GALAH, APOGEE, Gaia RVS, and \cite{bedell18} surveys. We find that a star's \feh\ and \alphafe\ are predictive of its \tcond\ slope (\rsq\ = $15 \pm 5\%$, $23 \pm 10\%$ respectively) while \teff\ and \logg\ contribute more weakly (\rsq\ = $9 \pm 5\%$, $13 \pm 16\%$). The Sun's abundance pattern resembles that of more metal-rich (0.1 dex) and $α$-depleted stars ($-0.02$ dex), suggesting a connection to broader GCE trends. To more accurately model stars' nucleosynthetic signatures, we apply the K-process model from \cite{Griffith24}, which casts each star's abundance pattern as a linear combination of core-collapse and Type Ia supernovae contributions. We find the Sun appears chemically ordinary in this framework, consistent with the intrinsic population scatter expected from stellar nucleosynthesis. We show that refractory element--\tcond\ trends arise because elements with higher \tcond\ have higher contributions from core-collapse supernovae. Refractory element depletion trends primarily reflect nucleosynthetic enrichment patterns shaped by GCE and local ISM inhomogeneities, with these processes accounting for $> 90\%$ of the observed variation within $2σ$. This work highlights how abundance diversity due to local and global chemical enrichment complicates the interpretation of population-scale planet-related chemical signatures in current datasets.

astro-ph.SR

Virtual Polarization Modulation: Enabling CSI-Free DCO-OFDM over Dynamic OWC Channels

In dynamically varying optical wireless communication (OWC) links, conventional quadrature amplitude modulation (QAM) in optical orthogonal frequency-division multiplexing (OFDM) requires frequent channel estimation and equalization, incurring pilot overhead and processing latency. This paper proposes a virtual polarization modulation (VPM)-based direct-current-biased optical OFDM (DCO-OFDM) scheme that maps each data symbol onto the three-dimensional Stokes space and places its corresponding Jones vector across two adjacent OFDM subcarriers. Using a rotation-based analytical framework, closed-form symbol error rate (SER) expressions are derived for arbitrary spherical constellations, along with upper and lower bounds and high signal-to-noise ratio (SNR) approximations. The framework is further extended to practical OWC scenarios with frequency-selective channels and atmospheric turbulence. Monte Carlo (MC) simulations validate the theoretical results. The results show that under practical OWC impairments, VPM outperforms QAM with least-squares (LS) channel estimation and minimum mean square error (MMSE) equalization. At a target SER of $10^{-5}$, 16-VPM achieves SNR gains of approximately 7.5 dB and 4 dB over equalized 16-QAM and 8-QAM, respectively, in frequency-selective channels, and a 6 dB advantage over equalized 16-QAM under atmospheric turbulence. By eliminating the need for channel state information, the proposed VPM-based DCO-OFDM provides a robust and low-latency solution for dynamic OWC links.

physics.optics

Mergers Fall Short: Non-merger Channels Required for Galactic Heavy Element Production

Since the discovery of the binary neutron star merger GW170817 and its associated kilonova, neutron star mergers have been established as a key production channel for r-process elements in the Universe. However, various lines of evidence, including observations of r-process abundances inferred from stellar spectra of Milky Way disk stars, suggest that additional channels are needed to fully account for r-process element enrichment in the Milky Way. Neutron star-black hole mergers and fast-merging binary neutron star systems are among the leading alternative candidates. In this paper, we combine gravitational-wave observations from LIGO-Virgo-KAGRA with data from short gamma-ray bursts, Galactic pulsars, and Galactic Eu/Fe versus Fe/H abundance observations to assess the contribution of these mergers to r-process enrichment in the Galactic disk. Our analysis employs a unified, likelihood-based inference framework that consistently propagates uncertainties in merger rates, delay-time distributions, mass- and spin-dependent ejecta yields, and stellar abundance measurements. We find that neither neutron star-black hole mergers nor fast- merging binary neutron star populations can serve as the dominant additional channel without generating strong tension with existing observations and theoretical expectations. These results constrain the viable sources of Galactic r-process enrichment and underscore the necessity of non- merger production channels.

astro-ph.HE

Microwave Imaging of Edge Conductivity in Graphene at Charge Neutrality and Quantum Hall States

We report local conductivity imaging of edge states in monolayer graphene by millikelvin microwave impedance microscopy (MIM). At the charge-neutrality point, as the magnetic field increases, the local conductivity at the edge drops to zero more slowly than in the bulk. This behavior is consistent with the calculated spatial profile of the charge gap in the canted antiferromagnetic phase. For comparison, we also perform microwave imaging of integer quantum Hall states away from neutrality, which host dissipationless chiral edge channels. The evolution of the edge signal as a function of the bulk gap is fundamentally different between the Landau level filling factor $ν= 0$ and $|ν| \ge 1$ integer quantum Hall states, which can be qualitatively explained by numerical simulations and theoretical analysis. Our results provide a comprehensive microscopic picture of the edge and bulk states as the Fermi level moves across the unique Landau-level spectrum of graphene.

cond-mat.str-el

FedRD: Reducing Divergences for Generalized Federated Learning via Heterogeneity-aware Parameter Guidance

Heterogeneous federated learning (HFL) aims to ensure effective and privacy-preserving collaboration among different entities. As newly joined clients require significant adjustments and additional training to align with the existing system, the problem of generalizing federated learning models to unseen clients under heterogeneous data has become progressively crucial. Consequently, we highlight two unsolved challenging issues in federated domain generalization: Optimization Divergence and Performance Divergence. To tackle the above challenges, we propose FedRD, a novel heterogeneity-aware federated learning algorithm that collaboratively utilizes parameter-guided global generalization aggregation and local debiased classification to reduce divergences, aiming to obtain an optimal global model for participating and unseen clients. Extensive experiments on public multi-domain datasets demonstrate that our approach exhibits a substantial performance advantage over competing baselines in addressing this specific problem.

cs.LG

FedCCA: Client-Centric Adaptation against Data Heterogeneity in Federated Learning on IoT Devices

With the rapid development of the Internet of Things (IoT), AI model training on private data such as human sensing data is highly desired. Federated learning (FL) has emerged as a privacy-preserving distributed training framework for this purpuse. However, the data heterogeneity issue among IoT devices can significantly degrade the model performance and convergence speed in FL. Existing approaches limit in fixed client selection and aggregation on cloud server, making the privacy-preserving extraction of client-specific information during local training challenging. To this end, we propose Client-Centric Adaptation federated learning (FedCCA), an algorithm that optimally utilizes client-specific knowledge to learn a unique model for each client through selective adaptation, aiming to alleviate the influence of data heterogeneity. Specifically, FedCCA employs dynamic client selection and adaptive aggregation based on the additional client-specific encoder. To enhance multi-source knowledge transfer, we adopt an attention-based global aggregation strategy. We conducted extensive experiments on diverse datasets to assess the efficacy of FedCCA. The experimental results demonstrate that our approach exhibits a substantial performance advantage over competing baselines in addressing this specific problem.

cs.LG

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: What essential capabilities are still missing? A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation. To move closer to human-level intelligence, models must similarly support multi-turn, multimodal interaction. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges. In this work, we present an initial exploration through the InterMT -- the first preference dataset for multi-turn multimodal interaction, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. InterMT captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances. To further this goal, we introduce InterMT-Bench to assess the ability of MLLMs in assisting judges with multi-turn, multimodal tasks. We demonstrate the utility of \InterMT through applications such as judge moderation and further reveal the multi-turn scaling law of judge model. We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Our project website can be found at https://pku-intermt.github.io .

cs.AI

SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates

Generative image compression has recently shown impressive perceptual quality, but often suffers from semantic inconsistency at ultra-low bitrates (bpp < 0.05), limiting its reliable deployment in bandwidth-constrained scenarios such as 6G semantic communications. This inconsistency stems from incomplete guidance information, which introduces semantic ambiguity into the generation process and may lead to natural-looking but source-inconsistent content. In this work, we propose a Semantic-Disambiguation-Guided Generative Image Compression (SDGIC) framework to constrain diffusion-based reconstruction at ultra-low bitrates. Specifically, SDGIC compresses the source image into three compact and complementary guidance streams: a concise text caption for global semantics, a highly compressed image (HCI) for dense visual evidence, and Reconstruction-Aware Semantic Residual Tokens (RSRTs) for reconstruction-relevant residual semantics that remain ambiguous under the text caption and HCI conditions. The RSRTs are directly optimized toward the downstream denoising objective, enabling them to provide source-specific semantic constraints for disambiguating diffusion-based reconstruction. To inject these three guidance streams into the generation process effectively, we design a Dual-Path Conditioned Diffusion Decoder (DPCD), which uses cross-attention for semantic conditions and ControlNet residuals for dense visual guidance. Extensive experiments demonstrate that SDGIC improves semantic consistency at ultra-low bitrates while maintaining favorable perceptual quality, with a 23.4% reduction in AFINE on the CLIC2020 dataset.

cs.CV

AI Deception: Risks, Dynamics, and Controls

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.

cs.AI

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their performance leaps lie more insidious and destructive safety risks, namely deception. Unlike hallucination, which arises from insufficient capability and leads to mistakes, deception represents a deeper threat in which models deliberately mislead users through complex reasoning and insincere responses. As system capabilities advance, deceptive behaviours have spread from textual to multimodal settings, amplifying their potential harm. First and foremost, how can we monitor these covert multimodal deceptive behaviors? Nevertheless, current research remains almost entirely confined to text, leaving the deceptive risks of multimodal large language models unexplored. In this work, we systematically reveal and quantify multimodal deception risks, introducing MM-DeceptionBench, the first benchmark explicitly designed to evaluate multimodal deception. Covering six categories of deception, MM-DeceptionBench characterizes how models strategically manipulate and mislead through combined visual and textual modalities. On the other hand, multimodal deception evaluation is almost a blind spot in existing methods. Its stealth, compounded by visual-semantic ambiguity and the complexity of cross-modal reasoning, renders action monitoring and chain-of-thought monitoring largely ineffective. To tackle this challenge, we propose debate with images, a novel multi-agent debate monitor framework. By compelling models to ground their claims in visual evidence, this method substantially improves the detectability of deceptive strategies. Experiments show that it consistently increases agreement with human judgements across all tested models, boosting Cohen's kappa by 1.5x and accuracy by 1.25x on GPT-4o.

cs.AI

Language Models Resist Alignment: Evidence From Data Compression

Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts merely superficial? In this work, we make the first exploration of this phenomenon from both theoretical and empirical perspectives. Empirically, we demonstrate the $\mathbf{elasticity}$ of post-alignment models, i.e., the tendency to revert to the behavior distribution formed during the pre-training phase upon further fine-tuning. Leveraging compression theory, we formally deduce that fine-tuning disproportionately undermines alignment relative to pre-training, potentially by orders of magnitude. We validate the presence of elasticity through experiments on models of varying types and scales. Specifically, we find that model performance declines rapidly before reverting to the pre-training distribution, after which the rate of decline drops significantly. Furthermore, we further reveal that elasticity positively correlates with the increased model size and the expansion of pre-training data. Our findings underscore the need to address the inherent elasticity of LLMs to mitigate their resistance to alignment. The model weight and code are available at pku-lm-resist-alignment.github.io.

cs.CL

FedAli: Personalized Federated Learning Alignment with Prototype Layers for Generalized Mobile Services

Personalized Federated Learning (PFL) enables distributed training on edge devices, allowing models to collaboratively learn global patterns while tailoring their parameters to better fit each client's local data, all while preserving data privacy. However, PFL faces two key challenges in mobile systems: client drift, where heterogeneous data cause model divergence, and the overlooked need for client generalization, as the dynamic of mobile sensing demands adaptation beyond local environments. To overcome these limitations, we introduce Federated Alignment (FedAli), a prototype-based regularization technique that enhances inter-client alignment while strengthening the robustness of personalized adaptations. At its core, FedAli introduces the ALignment with Prototypes (ALP) layer, inspired by human memory, to enhance generalization by guiding inference embeddings toward personalized prototypes while reducing client drift through alignment with shared prototypes during training. By leveraging an optimal transport plan to compute prototype-embedding assignments, our approach allows pre-training the prototypes without any class labels to further accelerate convergence and improve performance. Our extensive experiments show that FedAli significantly enhances client generalization while preserving strong personalization in heterogeneous settings.

cs.LG

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

In this study, we introduce the safety human preference dataset, PKU-SafeRLHF, designed to promote research on safety alignment in large language models (LLMs). As a sibling project to SafeRLHF and BeaverTails, we separate annotations of helpfulness and harmlessness for question-answering pairs, providing distinct perspectives on these coupled attributes. Overall, we provide 44.6k refined prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels ranging from minor to severe, with answers generated by Llama-family models. Based on this, we collected 166.8k preference data, including dual-preference (helpfulness and harmlessness decoupled) and single-preference data (trade-off the helpfulness and harmlessness from scratch), respectively. Using the large-scale annotation data, we further train severity-sensitive moderation for the risk control of LLMs and safety-centric RLHF algorithms for the safety alignment of LLMs. We believe this dataset will be a valuable resource for the community, aiding in the safe deployment of LLMs. Data is available at https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.

cs.AI

Reward Generalization in RLHF: A Topological Perspective

Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune language models. However, this shared topology has not been systematically characterized, nor have its alternatives been thoroughly explored, leaving the problems of low data efficiency and unreliable generalization unaddressed. As a solution, we introduce a theory of reward generalization in reinforcement learning from human feedback (RLHF), focusing on the topology of information flow at both macro and micro levels. At the macro level, we portray the RLHF information flow as an autoencoding process over behavior distributions, formalizing the RLHF objective of distributional consistency between human preference and model behavior. At the micro level, we present induced Bayesian networks to model the impact of dataset topologies on reward generalization. Combining analysis on both levels, we propose reward modeling from tree-structured preference information. It is shown to reduce reward uncertainty by up to $Θ(\log n/\log\log n)$ times compared to baselines, where $n$ is the dataset size. Validation on three NLP tasks shows that it achieves an average win rate of 65% against baselines, thus improving reward generalization for free via topology design, while reducing the amount of data requiring annotation.

cs.LG

Mitigating Deceptive Alignment via Self-Monitoring

Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly pursuing misaligned goals. Existing safety pipelines treat deception as a black-box output to be filtered post-hoc, leaving the model free to scheme during its internal reasoning. We ask: Can deception be intercepted while the model is thinking? We answer this question, the first framework that embeds a Self-Monitor inside the CoT process itself, named CoT Monitor+. During generation, the model produces (i) ordinary reasoning steps and (ii) an internal self-evaluation signal trained to flag and suppress misaligned strategies. The signal is used as an auxiliary reward in reinforcement learning, creating a feedback loop that rewards honest reasoning and discourages hidden goals. To study deceptive alignment systematically, we introduce DeceptionBench, a five-category benchmark that probes covert alignment-faking, sycophancy, etc. We evaluate various LLMs and show that unrestricted CoT roughly aggravates the deceptive tendency. In contrast, CoT Monitor+ cuts deceptive behaviors by 43.8% on average while preserving task accuracy. Further, when the self-monitor signal replaces an external weak judge in RL fine-tuning, models exhibit substantially fewer obfuscated thoughts and retain transparency. Our project website can be found at cot-monitor-plus.github.io

cs.AI

Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: $\mathbf{(I)}$ BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); $\mathbf{(II)}$ Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model's overall safety by an average of 40.9%. $\mathbf{(III)}$ Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%.

cs.LG