SearcharxivSearch

arXiv subjects

Kai Huang

Publications and source records attributed to Kai Huang.

At least 19 recordsLinked to original sources

Purely Electric-Field Control of Topological Magnetism in Two-Dimensional Magnets

Electrical control of topological magnetism is central to realizing energy-efficient topological spintronics. Yet most electric-field approaches modify the competing magnetic interactions through a nonselective rearrangement of low-energy electronic states, yielding coarse magnetic phase control that often requires external magnetic fields to stabilize topological quasiparticles, while few schemes solely based on electric fields are restricted to specific conducting materials. Here, we establish a general approach for purely electric-field control of topological magnetism, in which an applied electric field electrostatically dopes a selected orbital-angular-momentum-polarized band edge of a two-dimensional (2D) van der Waals (vdW) magnetic semiconductor via proximity to an adjacent nonmagnetic vdW metal. We show that the resulting electrostatic doping predominantly tunes magnetic anisotropy of the 2D magnet, while leaving exchange interaction and Dzyaloshinskii-Moriya interaction nearly unchanged, thereby reversibly driving the system through ferromagnetic, skyrmion, spiral, and bimeron phases. We demonstrate this mechanism for CrBr3/graphene and Cr2Ge2Te6/TaS2 vdW heterostructures hosting electron and hole pockets of different orbital characters. These results establish a broadly applicable strategy for purely electric-field control in topological spintronics.

cond-mat.mes-hall

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long-context modeling, whereas speech generation requires reconstructible features that preserve fine-grained acoustic detail. We introduce FireRedAudio, a general-purpose audio language model with a shared 9B-parameter LLM. To the best of our knowledge, it is the first publicly disclosed unified audio-language model to provide separate continuous input representations for understanding and generation within a single trainable autoregressive LLM. Audio to be recognized or analyzed is processed by a dedicated Audio Encoder, while speech inputs for generation use a RedAE-based pathway. The LLM directly generates text or conditions a flow-matching DiT to produce continuous acoustic latents. Through progressive multitask training, FireRedAudio supports ASR and audio understanding, with the latter extending to recordings of up to one hour, as well as zero-shot TTS, Instruct TTS, and semantic and acoustic speech editing. Its structured organization of long-form audio achieves second-level timestamp accuracy. Across comprehensive evaluations, FireRedAudio achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantial improvements over Ming-UniAudio-Edit in both semantic and acoustic speech editing. These results demonstrate the viability of decoupled continuous input representations for unifying audio understanding and continuous-latent speech generation in a model of moderate scale. Our code is available at https://github.com/FireRedTeam/FireRedAudio.

cs.CL

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.

cs.AI

Beyond Textual Repository Exploration: Dual-Modal Structural Reasoning for Agentic Issue Resolution

Recent advances in agentic program repair have significantly improved issue resolution by enabling iterative repository exploration. However, existing approaches predominantly rely on sequential, text-based code navigation, which fundamentally limits their ability to reason over large-scale long-horizon repositories with complex and long-range dependencies. As issue-resolution agents traverse repositories through fragmented textual observations, structural information such as module organization, call relationships, and dependency chains must be repeatedly reconstructed across interaction steps, often leading to exploration drift and incomplete localization. We present DUALVIEW, a dual-modal structural scaffolding framework that brings visual reasoning into repository exploration for issue-resolution agents. DUALVIEW represents repository structure through four complementary graph views: Module Coupling Graph (MCG), Function Call Graph (FCG), Class Hierarchy Graph (CHG), and Program Dependence Graph (PDG), and exposes them through a queryable interface with visual and textual responses. Rather than reconstructing repository structure from a sequence of textual observations, agents can directly reason over persistent visual representations of code dependencies, enabling more effective exploration and understanding of long-horizon codebases. We evaluate DUALVIEW on SWE-bench Pro and Verified. Results show that DUALVIEW consistently improves issue-resolution performance across different agent architectures and model families. Further ablation studies demonstrate that the gains arise not only from textual structural information but also from visual externalization of repository dependencies, which better supports long-horizon repository exploration.

cs.SE

A11YRepair: Bridging Web Accessibility Barriers via Knowledge-Enhanced Divide-and-Conquer Repair

Web accessibility (A11Y), which ensures web content is perceivable and usable for users with disabilities, is a critical requirement for modern web applications. Yet existing tooling overwhelmingly focuses on detecting A11Y violations rather than repairing them. Automated program repair (APR) techniques appear promising for this setting, but our study shows that state-of-the-art APR systems perform poorly when applied to real-world A11Y violations. Unlike conventional sparse-bug scenarios, web A11Y issues often manifest as multiple structurally related violations per page, requiring coordinated edits across multiple files. Existing repair systems fail to manage this multi-fault scale, as they handle each bug individually without considering their relationships or incorporating domain rules such as the Web Content Accessibility Guidelines (WCAG). We propose A11YRepair, an LLM-based framework for web A11Y repair. A11YRepair introduces a divide-and-conquer workflow that first clusters violations requiring coordinated edits to reduce redundant localization, and then decomposes each cluster by root cause so the LLM can generate focused and consistent patches. The framework further incorporates WCAG-driven knowledge to strengthen domain awareness during both fault localization and patch synthesis. To support systematic evaluation, we construct A11YBench, a benchmark of 60 real-world web projects collected from GitHub. Experimental results show that A11YRepair achieves higher repair effectiveness and lower cost than state-of-the-art baselines, and ablation studies confirm the importance of its divide-and-conquer design and selective domain knowledge integration. Specifically, patches generated by A11YRepair have been merged into open-source projects from Google, Microsoft, Facebook, IBM, K8s, Docker, and Alibaba, demonstrating its practical value.

cs.SE

Blow-Up Constructions and Applications to Segre Classes and Multidegree Formulas

By comparing simultaneous multigraded blow-ups with suitable iterated blow-up construction, we establish a birational correspondence between the associated exceptional divisors. We further investigate blow-ups arising from rational maps to multiprojective spaces and derive intersection-theoretic formulas for the Chern classes of the resulting exceptional divisors and pullback of tautological line bundles. These constructions have two main applications. First, they yield a proof of the product formula for Segre classes without any pure-dimensionality hypothesis. Second, they lead to a degree formula for arbitrary closed subschemes of multiprojective spaces, recovering and extending the classical formula of van der Waerden.

math.AG

Re-acceleration of Energetic Ions via Small-Scale Reconnection in Magnetic Fusion Plasmas

We report the first observation on the EXL-50U spherical torus that energetic particles injected by neutral beam injection (NBI) can be stably accelerated to significantly higher energies - reaching up to 2.5 times the injection energy, occurring without significant large-scale magnetohydrodynamic (MHD) bursts. Simulations based on EXL-50U parameters indicate that small-scale magnetic reconnection, mediated by multiple magnetic islands, fails to accelerate bulk thermal ions but efficiently energizes seed fast ions. Unlike global MHD events, such small-scale reconnection is ubiquitous in magnetic confinement devices and does not degrade core confinement. This mechanism offers a novel and potentially universal channel for auxiliary ion heating in future fusion reactors.

physics.plasm-ph

VisualNeo: Bridging the Gap between Visual Query Interfaces and Graph Query Engines

Visual Graph Query Interfaces (VQIs) empower non-programmers to query graph data by constructing visual queries intuitively. Devising efficient technologies in Graph Query Engines (GQEs) for interactive search and exploration has also been studied for years. However, these two vibrant scientific fields are traditionally independent of each other, causing a vast barrier for users who wish to explore the full-stack operations of graph querying. In this demonstration, we propose a novel VQI system built upon Neo4j called VisualNeo that facilities an efficient subgraph query in large graph databases. VisualNeo inherits several advanced features from recent advanced VQIs, which include the data-driven gui design and canned pattern generation. Additionally, it embodies a database manager module in order that users can connect to generic Neo4j databases. It performs query processing through the Neo4j driver and provides an aesthetic query result exploration.

cs.DB

EviRCOD: Evidence-Guided Probabilistic Decoding for Referring Camouflaged Object Detection

Referring Camouflaged Object Detection (Ref-COD) focuses on segmenting specific camouflaged targets in a query image using category-aligned references. Despite recent advances, existing methods struggle with reference-target semantic alignment, explicit uncertainty modeling, and robust boundary preservation. To address these issues, we propose EviRCOD, an integrated framework consisting of three core components: (1) a Reference-Guided Deformable Encoder (RGDE) that employs hierarchical reference-driven modulation and multi-scale deformable aggregation to inject semantic priors and align cross-scale representations; (2) an Uncertainty-Aware Evidential Decoder (UAED) that incorporates Dirichlet evidence estimation into hierarchical decoding to model uncertainty and propagate confidence across scales; and (3) a Boundary-Aware Refinement Module (BARM) that selectively enhances ambiguous boundaries by exploiting low-level edge cues and prediction confidence. Experiments on the Ref-COD benchmark demonstrate that EviRCOD achieves state-of-the-art detection performance while providing well-calibrated uncertainty estimates. Code is available at: https://github.com/blueecoffee/EviRCOD.

cs.CV

Remarks on Brauer-Manin obstruction for Weil restrictions

Given a finite extension $K/k$ of number fields and a smooth quasi-projective variety $X$ over $K$. If the abelianized fundamental group of $X$ is trivial, we prove that there is a natural identification between Brauer-Manin sets of $X$ and its Weil restriction $R_{K/k}X$. If $X$ is projective and $Pic(X\times_{K}\overline{k})$ is a torsion-free abelian group, we prove that there is a natural identification between algebraic Brauer-Manin sets of $X$ and $R_{K/k}X$.

math.NT

Wavelength-dependent photo-creep in halide perovskite single crystals

Halide perovskites are promising optoelectronic materials, but their time-dependent permanent deformation under illumination (i.e., photo-creep) is poorly understood, limiting their mechanical stability. Here we report wavelength-dependent photo-creep phenomena in CsPbBr3 and FAPbBr3 single crystals, studied by constant-load nanoindentation under controlled light with various wavelengths. Compared with creep in dark, continuous green light (near-bandgap) suppresses creep by 19% in CsPbBr3 and 10% in FAPbBr3, whereas violet (far above-bandgap) light enhances creep by 16% in CsPbBr3 and 8% in FAPbBr3. In contrast, when light is onset during creep, blue light enhances creep most prominently, whereas green light exhibits minimal influence. Such photo-creep behavior in halide perovskites are distinct with photo-plasticity phenomenon in conventional semiconductors. By combining the photoluminescence and photocurrent measurements, we unveil that ion migration promotes dislocation climb and creep, while carrier trapping suppresses dislocation glide and related creep in halide perovskites. Such competition between carrier trapping and ion migration tuned by wavelength governs the photo-creep response. Our findings uncover a photomechanical effect in halide perovskites and highlight how coupled carrier and ion dynamics under illumination affect their device reliability.

cond-mat.mtrl-sci

Exchange Frustration and Topological Magnetism in Electrostatically Doped SrRuO3

Magnetism in transition-metal systems emerges from exchange interactions that depend sensitively on carrier density. Yet leveraging this sensitivity to deliberately engineer exchange frustration and associated topological spin textures remains largely unexplored. Here, combining first-principles calculations with atomistic Monte Carlo simulations, we demonstrate that ferroelectric polarization enables electrostatic control of exchange frustration in the itinerant ferromagnet SrRuO3. We show that electrostatic hole doping renormalizes competing exchange interactions, driving SrRuO3 away from its bulk ferromagnetic ground state toward frustrated regimes, whereas electron doping largely preserves ferromagnetism. At BaTiO3/SrRuO3 interfaces, polarization-induced charge depletion modulates layer dependent exchange couplings, enhancing competition among J1, J2 and J3. The resulting exchange frustration stabilizes a sequence of magnetic phases as a function of thickness and applied magnetic field, including stripe and spiral states, topological meron and bimeron textures, and diverse skyrmionic objects. A minimal spin model identifies exchange frustration as the primary control parameter governing these crossovers, with magnetic anisotropy, Dzyaloshinskii-Moriya interaction, and external field selecting the emergent topology. Our results establish electrostatic doping as a route to engineer frustrated and topological magnetism in itinerant oxide metals.

cond-mat.mtrl-sci

A Differentiable Physical Framework for Goal-Driven Spin-State Engineering in Magnetic Resonance Spectroscopy

Magnetic Resonance Spectroscopy (MRS) offers a unique non-invasive window into metabolic processes, yet its potential remains strictly constrained by severe spectral congestion and intrinsic insensitivity. Traditional pulse sequence design, tethered to human intuition, predominantly targets simple quantum states, thereby overlooking the vast majority of the exponentially scaling operator space which consists of complex spin superpositions. Here, we introduce a spectrum-driven, end-to-end differentiable physical framework that transcends these heuristic limitations. By integrating physical laws with automatic differentiation algorithm, our approach directly navigates the high-dimensional spin dynamics space, bypassing the intractable inverse problem of state preparation. This enables the discovery of non-intuitive, complex mixed states that simultaneously satisfy the dual objectives of selective excitation and interferometric signal enhancement. We validate this paradigm by achieving the robust separation of Glutamate and Glutamine, which is a longstanding neuroimaging challenge, in the human brain at 3T, demonstrating spectral fidelity superior to conventional methods. By unlocking the "dark" informational content of nuclear spin ensembles, our work establishes a generalizable paradigm for goal-driven quantum state engineering in magnetic resonance and beyond.

quant-ph

Monolithic integration of diverse crystalline thin films on diamond for near-junction thermal management

The pursuit of extreme miniaturization and high power in 6G RF front-ends has cast thermal dissipation as the central challenge. Here, we have demonstrated the monolithic integration of functionally distinct single-crystal thin films, including \b{eta}-Ga2O3, Si, GaN, and LiTaO3, onto a single diamond substrate using a multi-step transfer printing technique. Focusing on the critical \b{eta}-Ga2O3/diamond interface, we achieve an exceptional interfacial thermal conductance (ITC) of 149 MW m-2 K-1 through ultra-high vacuum (UHV) annealing, creating an atomically sharp interface featuring covalent bonding. Vibrational electron energy-loss spectroscopy (EELS) analysis combining with molecular dynamics (MD) simulations reveal that distinctive interfacial phonon modes at the \b{eta}-Ga2O3/diamond heterointerface dominate ultrahigh ITC. We experimentally demonstrate that by improving the ITC, the thermal resistance (Rth) of a diamond-based \b{eta}-Ga2O3 MOSFET is driven to a record-low value of 1.58 K mm W-1, underscoring the critical role of interface engineering in near-junction thermal management for diamond-integrated devices. This work demonstrates a scalable, diamond-based monolithic integration platform designed to solve the near-junction thermal challenges in high-power RF front-ends.

cond-mat.mtrl-sci

FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System

We present FireRedASR2S, a state-of-the-art industrial-grade all-in-one automatic speech recognition (ASR) system. It integrates four modules in a unified pipeline: ASR, Voice Activity Detection (VAD), Spoken Language Identification (LID), and Punctuation Prediction (Punc). All modules achieve SOTA performance on the evaluated benchmarks: FireRedASR2: An ASR module with two variants, FireRedASR2-LLM (8B+ parameters) and FireRedASR2-AED (1B+ parameters), supporting speech and singing transcription for Mandarin, Chinese dialects and accents, English, and code-switching. Compared to FireRedASR, FireRedASR2 delivers improved recognition accuracy and broader dialect and accent coverage. FireRedASR2-LLM achieves 2.89% average CER on 4 public Mandarin benchmarks and 11.55% on 19 public Chinese dialects and accents benchmarks, outperforming competitive baselines including Doubao-ASR, Qwen3-ASR, and Fun-ASR. FireRedVAD: An ultra-lightweight module (0.6M parameters) based on the Deep Feedforward Sequential Memory Network (DFSMN), supporting streaming VAD, non-streaming VAD, and multi-label VAD (mVAD). On the FLEURS-VAD-102 benchmark, it achieves 97.57% frame-level F1 and 99.60% AUC-ROC, outperforming Silero-VAD, TEN-VAD, FunASR-VAD, and WebRTC-VAD. FireRedLID: An Encoder-Decoder LID module supporting 100+ languages and 20+ Chinese dialects and accents. On FLEURS (82 languages), it achieves 97.18% utterance-level accuracy, outperforming Whisper and SpeechBrain. FireRedPunc: A BERT-style punctuation prediction module for Chinese and English. On multi-domain benchmarks, it achieves 78.90% average F1, outperforming FunASR-Punc (62.77%). To advance research in speech processing, we release model weights and code at https://github.com/FireRedTeam/FireRedASR2S.

eess.AS

Nature of granular drag in microgravity

The influence of gravity on the drag force acting on a projectile impacting granular media is investigated experimentally via embedded inertial measurement unit (IMU) sensor and numerically through discrete element method (DEM) simulations. As gravity approaches zero, inertial drag dominates, yielding qualitatively different scaling laws and cavity dynamics. Analogous to fluid dynamics, we define a dimensionless granular drag coefficient $C_{\rm gd}$, which is found to stay largely at a constant $\sim 1.2$ in microgravity while an additional term inversely proportional to impact velocity arises in the presence of gravity. The constant term can be understood from momentum transfer along the penetration direction while the additional term suggests the influence of internal stress built-up due to gravity. Similar discrepancy is also found for the initial peak of the drag force. This analogy provides novel insights into the nature of granular drag in microgravity and sheds light on future space missions.

cond-mat.soft

Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks

With the rapid advancement of Artificial Intelligence, the Graphics Processing Unit (GPU) has become increasingly essential across a growing number of safety-critical application domains. Applying a GPU is indispensable for parallel computing; however, the complex data dependencies and resource contention across kernels within a GPU task may unpredictably delay its execution time. To address these problems, this paper presents a scheduling and analysis method for Directed Acyclic Graph (DAG)-structured GPU tasks. Given a DAG representation, the proposed scheduling scales the kernel-level parallelism and establishes inter-kernel dependencies to provide a reduced and predictable DAG response time. The corresponding timing analysis yields a safe yet nonpessimistic makespan bound without any assumption on kernel priorities. The proposed method is implemented using the standard CUDA API, requiring no additional software or hardware support. Experimental results under synthetic and real-world benchmarks demonstrate that the proposed approach effectively reduces the worst-case makespan and measured task execution time compared to the existing methods up to 32.8% and 21.3%, respectively.

cs.OS