SearcharxivSearch

arXiv subjects

Rui Mei

Publications and source records attributed to Rui Mei.

11 recordsLinked to original sources

Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.

cs.CR

Quaternion Tensor Modeling for Joint Color-Polarization Demosaicking

Division-of-focal-plane (DoFP) color polarization cameras enable snapshot acquisition of color polarization mosaic images, but the inherently sparse sampling pattern makes color polarization demosaicking severely ill-posed. Existing methods often fail to jointly exploit the correlations among polarization channels and the physical constraints inherent in polarization imaging, resulting in noticeable demosaicking artifacts. To address this issue, a quaternion-tensor-based color polarization demosaicking (CPDM) method incorporating Stokes-domain total variation (TV) regularization is proposed. Correlation analysis shows that the correlations among polarization channels are stronger than those among color channels. Accordingly, the color polarization images acquired at $0^\circ$, $45^\circ$, $90^\circ$, and $135^\circ$ are encoded into the four components of a third-order quaternion tensor, with the color channels organized along its third mode. A low-rank prior is then imposed on the quaternion tensor to exploit the global structural redundancy in the color polarization data. Moreover, spatial gradients are mapped to the Stokes domain through an orthogonal transformation to separate intensity, polarization and residual variations, with adaptive quaternion weights enabling component-specific regularization and preserving the energy consistency of the reconstructed Stokes vectors. An efficient optimization algorithm is derived for the resulting model. Extensive experiments demonstrate the superior demosaicking performance of the proposed method.

cs.CV

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL training environment, which reduces deployment overhead in Docker-level environments by two orders of magnitude. Finally, we deploy AgentDoG 1.5 as a training-free online guardrail for real-time safety moderation. Extensive experimental results indicate that AgentDoG 1.5 achieves state-of-the-art performance in diverse and complex interactive agentic scenarios. All models and datasets are openly released.

cs.AI

Birefringence-Driven Anisotropic $\alpha$-MoO3 Optical Cavities

Many anisotropic layered materials, despite their strong in-plane birefringence, exhibit substantial visible absorption, which severely restricts cavity lengths and hinders the observation of purely birefringence-governed optical phenomena. Here, we realize a birefringence-driven anisotropic optical cavity using $\alpha$-MoO3 flakes, capitalizing on their ultralow optical loss and pronounced in-plane birefringence. Using angle-resolved polarized Raman (ARPR) spectroscopy, we observe a mode-sensitive enhancement of anisotropy, dependent on both flake thickness and Raman shift. A unified model that incorporates the intrinsic Raman tensor, birefringence, and chromatic dispersion accurately reproduces the experimental data, elucidating how cavity resonances at both excitation and scattered wavelengths interact. Within this framework, the intrinsic phonon anisotropy is quantified, providing invaluable insights for accurately predicting ARPR responses and identifying crystallographic orientation. This work provides fundamental insights into birefringence-governed cavities and opens avenues for high-performance birefringent optics and cavity-enhanced anisotropic phenomena.

physics.optics

APT-ClaritySet: A Large-Scale, High-Fidelity Labeled Dataset for APT Malware with Alias Normalization and Graph-Based Deduplication

Large-scale, standardized datasets for Advanced Persistent Threat (APT) research are scarce, and inconsistent actor aliases and redundant samples hinder reproducibility. This paper presents APT-ClaritySet and its construction pipeline that normalizes threat actor aliases (reconciling approximately 11.22\% of inconsistent names) and applies graph-feature deduplication -- reducing the subset of statically analyzable executables by 47.55\% while retaining behaviorally distinct variants. APT-ClaritySet comprises: (i) APT-ClaritySet-Full, the complete pre-deduplication collection with 34{,}363 malware samples attributed to 305 APT groups (2006 - early 2025); (ii) APT-ClaritySet-Unique, the deduplicated release with 25{,}923 unique samples spanning 303 groups and standardized attributions; and (iii) APT-ClaritySet-FuncReuse, a function-level resource that includes 324{,}538 function-reuse clusters (FRCs) enabling measurement of inter-/intra-group sharing, evolution, and tooling lineage. By releasing these components and detailing the alias normalization and scalable deduplication pipeline, this work provides a high-fidelity, reproducible foundation for quantitative studies of APT patterns, evolution, and attribution.

cs.CR

ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models

The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing automated jailbreak generation approaches e.g. AutoDAN, suffer from limited mutation diversity, shallow fitness evaluation, and fragile keyword-based detection. To address these limitations, we propose ForgeDAN, a novel evolutionary framework for generating semantically coherent and highly effective adversarial prompts against aligned LLMs. First, ForgeDAN introduces multi-strategy textual perturbations across \textit{character, word, and sentence-level} operations to enhance attack diversity; then we employ interpretable semantic fitness evaluation based on a text similarity model to guide the evolutionary process toward semantically relevant and harmful outputs; finally, ForgeDAN integrates dual-dimensional jailbreak judgment, leveraging an LLM-based classifier to jointly assess model compliance and output harmfulness, thereby reducing false positives and improving detection effectiveness. Our evaluation demonstrates ForgeDAN achieves high jailbreaking success rates while maintaining naturalness and stealth, outperforming existing SOTA solutions.

cs.CR

Raman Forbidden Layer-Breathing Modes in Layered Semiconductor Materials Activated by Phonon and Optical Cavity Effects

We report Raman forbidden layer-breathing modes (LBMs) in layered semiconductor materials (LSMs). The intensity distribution of all observed LBMs depends on layer number, incident light wavelength and refractive index mismatch between LSM and underlying substrate. These results are understood by a Raman scattering theory via the proposed spatial interference model, where the naturally occurring optical and phonon cavities in LSMs enable spatially coherent photon-phonon coupling mediated by the corresponding one-dimensional periodic electronic states. Our work reveals the spatial coherence of photon and phonon fields on the phonon excitation via photon/phonon cavity engineering.

cond-mat.mtrl-sci

Anomalous Gate-tunable Capacitance in Van der Waals Heterostructures

The ferroelectricity emerging in non-polar graphene/hexagonal boron nitride (hBN) heterostructures has drawn considerable attention because of its fascinating properties and promising high-frequency electrical polarization switching. Yet, the underlying mechanism is still under debate. Here in twisted double bilayer graphene (TDBLG) aligned with its neighboring hBN, we observed two types of hysteresis - delayed hysteresis in top gate induced by the anomalous screening, and advanced hysteresis in back gate caused by the anomalous gate-tunable capacitance. To investigate the role played by moir\'e potential in the anomalous hysteresis, we studied a moir\'eless graphene heterostructure as control experiment. Unexpectedly, we observed exactly the same phenomena in this control device. Our findings suggest that the anomalous ferroelectricity in graphene/hBN heterostructures may originate from the dielectric material hBN, calling for further structural investigations on hBN. The observation of gate-tunable capacitance provides more insights in the mysterious ferroelectricity in graphene/hBN heterostructures, and should enable new design of memory devices such as memcapacitor based on tunable capacitance.

cond-mat.mes-hall

Quantitatively predicting angle-resolved polarized Raman intensity of anisotropic layered materials

Angle-resolved polarized Raman (ARPR) spectroscopy provides insights into optical anisotropy and symmetry-related electron-photon/electron-phonon couplings of anisotropic layered materials (ALMs). However, since their discovery over ten years ago, ARPR responses in ALM flakes has exhibited a puzzling dependence on flake thickness, excitation wavelength, and dielectric environment, complicating their understanding and prediction. By taking black phosphorus (BP) (large than 20 nm) flakes and four-layer Td-WTe2 as examples, this study introduces intrinsic Raman tensors (Rint) and proposes strategies to predict the ARPR intensity profiles of thick and atomically-thin ALM flakes by considering birefringence, linear dichroism and multilayer interference inside multilayered structures with experimentally determined complex refractive indexes along in-plane axes and complex tensor elements of Rint for the corresponding phonon modes. The tensor elements of effective Raman tensors (Reff), which are directly linked to the polarization vectors of incident and scattered light outside the ALM surface, were derived to quantitatively predict ARPR intensity for these ALM flakes, showing intricate dependence on ALM thickness, dielectric substrates, and excitation wavelengths. This framework can be extended to other ALM flakes from atomically-thin layers to bulk limit, facilitating comprehensive prediction of their ARPR intensity regardless of layer-dependent electronic properties.

cond-mat.mes-hall

Tuning the Interlayer Microstructure and Residual Stress of Buffer-Free Direct Bonding GaN/Si Heterostructures

The direct integration of GaN with Si can boost great potential for low-cost, large-scale, and high-power device applications. However, it is still challengeable to directly grow GaN on Si without using thick strain relief buffer layers due to their large lattice and thermal-expansion-coefficient mismatches. In this work, a GaN/Si heterointerface without any buffer layer is successfully fabricated at room temperature via surface activated bonding (SAB). The residual stress states and interfacial microstructures of GaN/Si heterostructures were systematically investigated through micro-Raman spectroscopy and transmission electron microscopy. Compared to the large compressive stress that existed in GaN layers grown-on-Si by MOCVD, a significantly relaxed and uniform small tensile stress was observed in GaN layers bonded-to-Si by SAB; this is mainly ascribed to the amorphous layer formed at the bonding interface. In addition, the interfacial microstructure and stress states of bonded GaN/Si heterointerfaces was found can be significantly tuned by appropriate thermal annealing. This work moves an important step forward directly integrating GaN to the present Si CMOS technology with high quality thin interfaces, and brings great promises for wafer-scale low-cost fabrication of GaN electronics.

cond-mat.mtrl-sci

Control of Raman scattering quantum interference pathways in graphene

Graphene is an ideal platform to study the coherence of quantum interference pathways by tuning doping or laser excitation energy. The latter produces a Raman excitation profile that provides direct insight into the lifetimes of intermediate electronic excitations and, therefore, on quantum interference, which has so far remained elusive. Here, we control the Raman scattering pathways by tuning the laser excitation energy in graphene doped up to 1.05eV, above what achievable with electrostatic doping. The Raman excitation profile of the G mode indicates its position and full width at half maximum are linearly dependent on doping. Doping-enhanced electron-electron interactions dominate the lifetime of Raman scattering pathways, and reduce Raman interference. This paves the way for engineering quantum pathways in doped graphene, nanotubes and topological insulators.

cond-mat.mes-hall