SearcharxivSearch

arXiv subjects

Qiao Li

Publications and source records attributed to Qiao Li.

At least 19 recordsLinked to original sources

Quadrupole White-light Sources in an X1.2 Flare Observed by ASO-S/LST/WST and SDO/HMI

We present observations of an X1.2 white-light flare on 2023 January 6, which exhibits a rare quadrupolar white-light source configuration. This event was observed by the White-light Solar Telescope (WST; 3600 \AA) aboard the Advanced Space-based Solar Observatory and the Helioseismic and Magnetic Imager (HMI; 6173 \AA) aboard the Solar Dynamics Observatory. Four flare-related footpoints (labeled as FP1--FP4) were nearly simultaneously identified in both WST 3600 \AA\ and HMI 6173 \AA\ continua, associated with a quadrupolar magnetic configuration and a failed filament eruption. The inner sources of FP1 and FP2 showed a similar enhancement of $\sim$65%/10% in the WST/HMI continuum, while the outer sources of FP3 and FP4 exhibited weaker responses. The inner footpoints had earlier responses in UV and EUV bands and were spatially coincident with the hard X-ray (HXR) footpoint sources. The two southern footpoints (FP2 and FP4) showed stronger HXR and white-light emissions than their northern counterparts (FP1 and FP3), with FP4 uniquely exhibiting a distinct HXR emission above 60 keV, in contrast to the absence of such an emission at FP3. Notably, faint WST 3600 \AA\ enhancements at FP4 were observed during the gradual phase, temporally consistent with the fallback of filament material. This X1.2 flare presents a novel quadrupolar white-light structure, enriching our understanding of the generation and evolution of white-light flares.

astro-ph.SR

Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors

The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hindering real-world deployment. In this paper, we propose a controllable method operating on the model's continuous textual representation to erase target characters during generation. We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effectiveness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance.

cs.CV

Semantic Steering for Controllable Generation: Tuning-Free Concept Erasure in Multimodal Diffusion Transformers

Multimodal Diffusion Transformers (MM-DiTs) have demonstrated remarkable text-to-image generation performance, surpassing traditional U-Net-based diffusion models. Nevertheless, their powerful generative capabilities also raise significant safety concerns, as they may generate sensitive or inappropriate content. While existing concept erasure methods aim to mitigate such risks, most require modifying model parameters, which are often architecture-specific and impractical for deployed larger models. Several tuning-free approaches face challenges when applied to advanced large-scale MM-DiTs due to their deeply embedded knowledge, broad semantic space, and context-dependent text encoders. To address these challenges, we propose to erase concepts by directly manipulating the model's internal representations. Our key insight, derived from an in-depth analysis of MM-DiT's block-wise generative roles, is that text-conditioned semantic representations are most salient in the middle blocks of MM-DiTs. Based on this, we extract representations of an unwanted concept and a desirable safe one from the middle block, construct a steering vector from their difference, and inject this single vector into consecutive early and middle blocks. By operating exclusively on the sparse text-branch tokens and leveraging the straight sampling trajectory of rectified flow, our method achieves effective concept erasure with negligible overhead and without any training. Extensive experiments across MM-DiT models demonstrate that our method achieves state-of-the-art performance in erasing diverse concepts, enables effective control over the final output, and remains robust to adversarial attacks.

cs.CV

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remediation systems are designed to generate actions rather than to decide whether intervention is warranted, leaving safety as an afterthought enforced by manual approval. This paper makes three contributions to close this gap: (i) we reformulate safe remediation as a risk-constrained intervention decision problem and cast it as a Constrained Markov Decision Process (CMDP), in which the agent maximizes repair success subject to a bounded false remediation rate (FRR); (ii) we introduce a three-dimensional risk decomposition comprising blast radius, reversibility, and epistemic uncertainty, providing operators with an interpretable per-action safety interface; and (iii) we design a context-adaptive human-in-the-loop (HITL) gate that turns escalation from a binary failsafe into a bandwidth-aware control layer responsive to on-call load and business criticality. The full policy is learned offline from historical incident logs, enabling explicit control of the expected FRR. Experiments on the Train Ticket microservice benchmark with Chaos Mesh fault injection and an RCAEval-aligned fault taxonomy show that our framework reduces FRR by 39% while improving repair success by 2.5 points over a strong runbook baseline, and reduces on-call escalation load by 17% relative to a fixed-threshold variant.

cs.AI

Tunable Extended Magnetic Non-Fermi Liquid in Graphene Moir\'e Heterostructures

Exploring exotic quantum metallic states beyond Landau's Fermi liquid theory remains a central focus in condensed matter physics. Such non-Fermi liquid behavior is mostly observed near quantum criticality, yet growing attention is directed toward extended NFL phases with intrinsic quantum fluctuations rooted in the extended ground state. While these extended NFL states have been previously reported only in a limited set of d- and f-electron systems, realizing a single, highly tunable platform capable of exhibiting multiple resistance exponent values is essential for uncovering the connection between the resistance exponent and the dominant quantum fluctuations coupled to quasiparticles. However, corresponding experimental progress remains elusive. Here, we report the observation of tunable extended non-Fermi liquid behavior in twisted double bilayer graphene encapsulated by aligned hBN layers. This NFL phase spans a broad range of carrier densities and exhibiting a carrier density dependent resistance exponent. Combined with temperature dependent resistance, magnetotransport and differential resistance measurements, these findings support a scenario where strong quantum fluctuations emerge from the interplay between localized and itinerant carriers. Our work establishes a highly tunable platform beyond conventional frameworks to investigate the organizing principles of non-Fermi liquid physics manifested in diverse behaviors.

cond-mat.mes-hall

Layer-tunable Hubbard bands probed via moir\'e excitons in MoSe$_2$/WS$_2$ heterostructures

Moir\'e superlattices in transition metal dichalcogenide heterostructures provide a highly tunable platform for engineering strongly interacting states at the nanoscale. However, quantitatively determining and in-situ tuning of the underlying Hubbard parameters remains experimentally challenging. Here, we report electric-field-driven reordering of layer-specific Hubbard bands by performing optical spectroscopy on a dual-gated, 60{\deg}-aligned MoSe$_2$/WS$_2$ heterobilayer. Using two spatially distinct moir\'e excitons as local optical probes and tracking them as a function of carrier filling and vertical electric field, we quantitatively extract the layer-dependent on-site Coulomb repulsions, U$_M$~60 meV in MoSe$_2$ and U$_W$~30 meV in WS$_2$. Furthermore, we stabilize generalized Wigner crystal and stripe phases by electrostatically tuning the system to a type-II band alignment, shifting the ground state into the WS$_2$ layer where reduced on-site repulsion allows inter-site Coulomb interactions to dominate. Our results establish vertical electric fields as a deterministic tuning knob for layer-selective Hubbard physics, enabling device-level control of complex many-body phases.

cond-mat.mes-hall

Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures

Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet after several planning-related functions are absorbed into a unified VLA model, part of the original CPU budget becomes underutilized, while the visual encoder and the main reasoning path still concentrate most computation and memory demand on the GPU. As a result, directly deploying VLA together with the rest of the onboard system can be hard under realistic GPU memory constraints. To address this issue, we present a hybrid CPU--GPU inference framework with flexible resource scheduling for autonomous driving. Our design partitions the VLA backbone at the block-layer granularity, executes the visual encoder and LLM prefix on the GPU, and offloads the LLM suffix to the CPU through a cross-frame asynchronous pipeline, thereby exposing a schedulable boundary for redistributing compute and memory pressure across heterogeneous processors. We evaluate the proposed framework on two representative driving VLA models, Orion and MindDrive. On Bench2Drive, our method reduces average latency from 521ms to 408.0ms for Orion and from 443ms to 306.2ms for MindDrive, corresponding to 21.7% and 30.9% reduction, respectively. For Orion, the estimated peak GPU memory is further reduced from 45GB to 29GB. In real-vehicle deployment under coexistence with Autoware.Universe, native Orion cannot run because the onboard GPU memory budget is insufficient, whereas the hybrid version runs successfully together with the full vehicle stack.

cs.DC

Stokes phenomenon and quantum supergroup $U_q(\mathfrak{gl}(m|n))$

In this paper we study the Stokes phenomenon of the quantum confluent hypergeometric supersystem, certain meromorphic linear system of ordinary differential equation with a second order pole, associated to the Lie superalgebra $\mathfrak{gl}_{m|n}$. We prove that its Stokes supermatrices satisfy the Yang-Baxter equation, and thus give rise to the quantum supergroup $U_q(\mathfrak{gl}(m|n))$.

math.QA

GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference

Deploying large language models (LLMs) as cloud services raises privacy concerns as inference may leak sensitive data. Fully Homomorphic Encryption (FHE) allows computation on encrypted data, but current FHE methods struggle with efficient and precise nonlinear function evaluation. Specifically, CKKS-based approaches require high-degree polynomial approximations, which are costly when target precision increases. Alternatively, TFHE's Programmable Bootstrapping (PBS) outperforms CKKS by offering exact lookup-table evaluation. But it lacks high-precision implementations of LLM nonlinear layers and underutilizes GPU resources. We propose \emph{TIGER}, the first GPU-accelerated framework for high-precision TFHE-based nonlinear LLM layer evaluation. TIGER offers: (1) GPU-optimized WoP-PBS method combined with numerical algorithms to surpass native lookup-table precision limits on nonlinear functions; (2) high-precision and efficient implementations of key nonlinear layers, enabling practical encrypted inference; (3) batch-driven design exploiting inter-input parallelism to boost GPU efficiency. TIGER achieves 7.17$\times$, 16.68$\times$, and 17.05$\times$ speedups over a CPU baseline for GELU, Softmax, and LayerNorm, respectively.

cs.CR

GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning and inherits errors from 2D predictions. To address these limitations, we propose GeoGuide, a novel framework that leverages pretrained 3D models to integrate hierarchical geometry-semantic consistency for open-vocabulary 3D segmentation. Specifically, we introduce an Uncertainty-based Superpoint Distillation module to fuse geometric and semantic features for estimating per-point uncertainty, adaptively weighting 2D features within superpoints to suppress noise while preserving discriminative information to enhance local semantic consistency. Furthermore, our Instance-level Mask Reconstruction module leverages geometric priors to enforce semantic consistency within instances by reconstructing complete instance masks. Additionally, our Inter-Instance Relation Consistency module aligns geometric and semantic similarity matrices to calibrate cross-instance consistency for same-category objects, mitigating viewpoint-induced semantic drift. Extensive experiments on ScanNet v2, Matterport3D, and nuScenes demonstrate the superior performance of GeoGuide.

cs.CV

The Lyman-alpha Emission in Solar Flares. II. A Statistical Study on Its Relationship with the White-light plus Soft X-ray Emission

The hydrogen \lya\ line and the white-light (WL) continuum are two key diagnostics of energy transport in the lower atmosphere during solar flares, yet their relationship remains poorly understood. Here we present a statistical analysis of 69 white-light flares (WLFs) to investigate the relationships among the \lya, soft X-ray (SXR), and WL continuum emissions using the data from GOES and the Helioseismic and Magnetic Imager (HMI) on the Solar Dynamics Observatory. We find that the \lya\ contrast in these WLFs ranges 0.8--28.5\% with a mean value of 7.0\%. Positive power-law relationships exist among peak enhancements in SXR, \lya, and WL. For most events, the \lya\ peak is nearly co-temporal with the peak of SXR time derivative, whereas the WL peak is either co-temporal with or lags those of \lya\ and SXR derivative. The \lya\ and WL rise times are similar ($\sim$3--4 min) and correlated. We also find that the radiated energy in \lya\ and HMI narrow-band WL has a positive power-law relationship with duration. In particular, the power-law index for the narrow-band WL is very close to 1/3 as predicted by magnetic reconnection theory. On average, the radiated energies in GOES \lya\ and SXR bands are approximately three orders of magnitude greater than the energy emitted in the continuum near 6173 \AA\ with a bandwidth of 1 \AA. Our findings provide new constraints on lower-atmosphere energy transport in solar flares and can serve as valuable references for modelling and interpreting the flares on solar-type stars.

astro-ph.SR

ClawMobile: Rethinking Smartphone-Native Agentic Systems

Smartphones represent a uniquely challenging environment for agentic systems. Unlike cloud or desktop settings, mobile devices combine constrained execution contexts, fragmented control interfaces, and rapidly changing application states. As large language models (LLMs) evolve from conversational assistants to action-oriented agents, achieving reliable smartphone-native autonomy requires rethinking how reasoning and control are composed. We introduce ClawMobile as a concrete exploration of this design space. ClawMobile adopts a hierarchical architecture that separates high-level language reasoning from structured, deterministic control pathways, improving execution stability and reproducibility on real devices. Using ClawMobile as a case study, we distill the design principles for mobile LLM runtimes and identify key challenges in efficiency, adaptability, and stability. We argue that building robust smartphone-native agentic systems demands principled coordination between probabilistic planning and deterministic system interfaces. The implementation is open-sourced~\footnote{https://github.com/ClawMobile/ClawMobile} to facilitate future exploration.

cs.MA

Stack of correlated insulating states in bilayer graphene kagome superlattice

Graphene-based systems have emerged as a rich platform for exploring emergent quantum phenomena-including superconductivity, magnetism, and correlated insulating behavior-arising from flat electronic bands that enhance many-body interactions. Realizing such flat bands has thus far relied primarily on moiré graphene superlattices or rhombohedral stacking graphene systems, both of which face challenges in reproducibility and tunability. Here, we introduce an artificial Kagome superlattice in bilayer graphene, engineered via nanopatterning of the dielectric substrate to create a precisely defined and electrostatically tunable periodic potential. Magnetotransport measurements reveal the emergence of a stack of correlated insulating states at moderate superlattice potentials, characteristic of strong electron-electron interactions within Kagome-induced flat bands. As temperature increases, these correlated gaps collapse, signaling the thermal suppression of interaction-driven states. Continuum-model calculations confirm the formation of multiple flat minibands and reproduce the observed evolution of band reconstruction. Our results establish dielectric-patterned graphene superlattices as a robust and controllable architecture for realizing flat-band-induced correlated phenomena beyond moiré systems.

cond-mat.mes-hall

Characterize LSM-tree Compaction Performance via On-Device LLM Inference

Modern key-value storage engines built on Log-Structured Merge-trees (LSM-trees), such as RocksDB and LevelDB, rely heavily on the performance of their compaction operations, which are impacted by a complex set of interdependent configuration parameters. Manually tuning these parameters for optimal performance demands considerable expertise, while traditional auto-tuning approaches struggle with the enormous search space and low sample efficiency inherent to this domain. In recent years, Large Language Models (LLMs) have demonstrated strong capabilities in code generation and logical reasoning, offering new possibilities for system optimization. However, applying LLMs to real-time compaction tuning in such latency-sensitive environments is a double-edged sword. While large-scale LLMs can offer superior reasoning for strategy generation, their high inference latency and computational cost make them impractical for interactive, low-latency tuning. In contrast, small-scale LLMs achieve low latency but often at the expense of reasoning accuracy and tuning effectiveness. In this paper, we first evaluate this trade-off by analyzing the compaction-tuning performance and inference latency of LLMs at different scales in an LSM-tree-based tuning case. We then characterize the performance of LSM-tree on RocksDB v8.8.1, with a focus on adjusting the key compaction-related parameters under db_bench workloads. Our experimental results show a clear positive correlation between model capability and tuning effectiveness.

cs.PF

Device-Level Optimization Techniques for Solid-State Drives: A Survey

Solid-state drives (SSDs) have revolutionized data storage with their high performance, energy efficiency, and reliability. However, as storage demands grow, SSDs face critical challenges in scalability, endurance, latency, and security. This survey provides a comprehensive analysis of SSD architecture, key challenges, and device-level optimization techniques. We first examine the fundamental components of SSDs, including NAND flash memory structures, SSD controller functionalities (e.g., address mapping, garbage collection, wear leveling), and host interface protocols. Next, we discuss major challenges such as reliability degradation, endurance limitations, latency variations, and security threats. We then explore advanced optimization techniques, including error correction mechanisms, flash translation layer (FTL) enhancements, and emerging architectures like zoned namespace (ZNS) SSDs and flexible data placement (FDP). Finally, we highlight open research challenges, such as QLC/PLC NAND scalability, performance-reliability trade-offs, and SSD optimizations for AI/LLM workloads. This survey aims to guide future research in the development of next-generation SSDs that balance performance, endurance, and security in evolving storage ecosystems.

cs.AR

ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management

Efficiently serving Large Language Models (LLMs) with persistent Prefix Key-Value (KV) Cache is critical for applications like conversational search and multi-turn dialogue. Serving a request requires loading the pre-computed prefix KV cache and generating the first token, defined as the Re-Prefill Phase. Offloading this shared prefix cache to secondary storage is essential for memory scalability. Re-Prefill with offloading suffers from severe I/O bottlenecks in two aspects. First, semantic-aware KV cache pruning algorithms select important tokens in fine granularity, while systems manage I/O in coarse, fixed-size blocks, causing severe read amplification. Second, the sequential dependency between identifying important tokens and loading KV cache creates idle I/O and compute bubbles, under-utilizing system resources. This paper proposes \textit{ContiguousKV}, a high-performance prefix KV cache offloading system that bridges algorithmic semantics with I/O efficiency to accelerate the Re-Prefill phase. We first introduce \textit{ContiguousChunk}, a unified data management granularity that aligns KV cache pruning with I/O operations. All the mechanisms critical for I/O performance are performed at the granularity of ContiguousChunk, thereby eliminating read amplification. By exploiting the high similarity in important ContiguousChunk indices across layers, we propose intra- and inter-period asynchronous prefetching to break the sequential dependency between I/O and compute, effectively eliminating idle bubbles. Finally, we propose attention-guided cache management to retain semantically critical prefix data in memory. Evaluations on Qwen2.5 series models show that ContiguousKV achieves a 3.85x speedup in the Re-Prefill phase over the state-of-the-art offloading system IMPRESS, while maintaining high output quality.

cs.OS

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to extreme viewpoint discrepancies, occlusions, and domain gaps between aerial and ground imagery. While prior works have made progress by learning cross-view representations, they remain limited in handling severe pose variations and spatial misalignment. To address these issues, we propose a Geometric and Semantic Alignment Network (GSAlign) tailored for AG-ReID. GSAlign introduces two key components to jointly tackle geometric distortion and semantic misalignment in aerial-ground matching: a Learnable Thin Plate Spline (LTPS) Module and a Dynamic Alignment Module (DAM). The LTPS module adaptively warps pedestrian features based on a set of learned keypoints, effectively compensating for geometric variations caused by extreme viewpoint changes. In parallel, the DAM estimates visibility-aware representation masks that highlight visible body regions at the semantic level, thereby alleviating the negative impact of occlusions and partial observations in cross-view correspondence. A comprehensive evaluation on CARGO with four matching protocols demonstrates the effectiveness of GSAlign, achieving significant improvements of +18.8\% in mAP and +16.8\% in Rank-1 accuracy over previous state-of-the-art methods on the aerial-ground setting.

cs.CV

On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices

Sparsity is essential for deploying large models on resource constrained edge platforms. However, optimizing sparsity patterns for individual tasks in isolation ignores the significant I/O overhead incurred during frequent task switching. We introduce an on-demand multi-task sparsity framework specifically designed to minimize switching costs by maximizing parameter reuse. Unlike monolithic approaches, we decompose weights into reusable block-granular units and align sparse structures across tasks to maximize overlap. By dynamically loading only the small differential set of blocks required for the next task, our method effectively mitigates the cold-start latency inherent in traditional monolithic approaches.Experiments on a real-world autonomous driving platform demonstrate that our framework achieves superior switching efficiency, accelerating task switching by over 6.6X on average compared to existing sparsity methods.

cs.LG