SearcharxivSearch

arXiv subjects

Hao Li

Publications and source records attributed to Hao Li.

At least 19 recordsLinked to original sources

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18$\rightarrow$14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut

cs.SD

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an $\mathcal{O}(n^{2/3})$ leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.

cs.LG

Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering

Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.

cs.SD

TacClip: a clip-on sensor measures dynamic contact forces without covering the fingerpads

TacClip is a minimally encumbering wearable device for recording fingertip deformation caused by contact forces and vibrations. It can be combined with vision- or glove-based hand tracking systems that leave the fingertips uncovered and provides a measure of dynamic contact interactions, while leaving the finger pads exposed so that the user retains natural sensitivity to texture, friction, temperature, and fine surface features. The signal is produced by a Fiber Bragg Grating (FBG) embedded on a small plastic clip mounted over the fingernail. Optionally, for use with vision-based tracking, additional FBGs on polyimide strips can complement camera-based pose estimation. In finger pressing tests, TacClip estimates the force magnitude with typical errors below $0.5~\mathrm{N}$ over a $0$--$8~\mathrm{N}$ range. In tests of cloth handling and tape edge finding, we show that it captures the vibrations and dynamic events generated during exploratory sliding. With no electronics, TacClip can also be used submerged in water, while preserving bare finger contact.

cs.RO

Edge-Dominated Twist Mechanics at van der Waals Interfaces

Despite the pivotal role of twist in modulating physical properties at van der Waals (vdW) interfaces, the mechanics governing torsional response remain poorly understood. Here, we probe twist mechanics at homo- and heterogeneous vdW interfaces, together with their sliding behaviors within a unified experimental framework. For both systems, the peak torque scales nearly linearly with contact area, in contrast to predictions from linear elastic and rigid models. Remarkably, while the sliding friction of the two interfaces diverges by over three orders of magnitude owing to different scaling laws, the corresponding torque follows the same linear scaling and differs by only about twenty-fold. Large-scale atomistic simulations reveal an edge-dominated yielding mechanism for torsional motion, wherein elastic reconstruction shifts the effective load-bearing region toward the edges, eliminating torque from the contact interior. This mechanism contrasts with the bulk-mediated stress transmission governing translational sliding, a distinction rooted in the different loading geometries inherent to the two motion modes, where torsional loading necessitates perimeter actuation, whereas sliding enables center-driven loading. This symmetry-imposed divergence demonstrates that translational and torsional properties cannot be predicted from one another at vdW interfaces, providing critical insights for the design of dynamically reconfigurable micro- and nanoelectromechanical devices.

cond-mat.mtrl-sci

Collective Excitonic Structure Governs Anomalously Weak Thermal Optical Dephasing in Conjugated Polymers

Conjugated polymer aggregates exhibit optical decoherence in the presence of strong vibronic coupling and substantial diversity in chemical structure, solid-state organization, and excitonic character. Here, we use coherence-detected and population-detected two-dimensional electronic spectroscopies to examine the temperature dependence of the homogeneous optical linewidth across a series of semiconducting polymers. Despite substantial differences in molecular architecture and solid-state organization, all polymers studied exhibit remarkably weak thermal linewidth scaling over the measured temperature range. Comparison between complementary detection modalities further shows that, while this weak temperature dependence is obust, the absolute homogeneous linewidth depends on the measured observable. This behavior reflects the different ways in which coherence- and population-detected measurements project population relaxation and pure dephasing onto the spectroscopic response. These results establish the weak thermal scaling of optical decoherence across a diverse series of conjugated polymers and show that its experimental manifestation must be interpreted in the context of the detection observable.

cond-mat.mtrl-sci

Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent decay gates on geometric increments yields 28.9% success, compared with 36.7% when those increments are shuffled during training and 24.4% without explicit object/goal geometry. Both geometry policies receive correct inputs at evaluation. A token adapter using the same increments scores 27.8%; differences vary across seeds and remain inconclusive. Token-clock conditioning scores 11.1%, including one seed that fails to converge. In separate robustness tests, a state-only relative-coordinate policy retains 7/10 success under frame relabeling, whereas all four tested visual policies fall to at most 3/20 after a 5 cm object displacement. These results show no reliable advantage from training-time geometric alignment under this recipe and illustrate the gap between coordinate invariance and physical-layout generalization. Episode records, seed-level analyses, and figure-generation code accompany the paper.

cs.RO

Twin-photon generation in a silicon nitride microresonator

Photonic chips with silicon nitride ($\mathrm{Si_3N_4}$) microring resonators are well established as heralded single-photon sources, but their operation as frequency-degenerate twin-photon sources has not previously been demonstrated. Here, we realise a twin-photon source at telecommunication wavelengths in a $\mathrm{Si_3N_4}$ ring microresonator via an inverse four-wave mixing (FWM) process, in which two photons from spectrally distinct pumps are converted into a pair of identical twin photons. The measurements show a maximum coincidence-to-accidental ratio (CAR) of $5.4\pm0.6$. In addition, the microresonator functions as a heralded single-photon source through pump-degenerate spontaneous four-wave mixing (SFWM), exhibiting a spectral purity of $P=0.67\pm0.05$ and a heralded anti-bunching of $g^{(2)}_h(0)=0.0042\pm0.0015$. Together, these results demonstrate both photon-generation schemes on a single integrated $\mathrm{Si_3N_4}$ platform, highlighting its potential for scalable, tailored quantum light generation.

physics.optics

SchedBlame: Who Ran While You Waited? Culprit-Attributed CPU Contention for Containers on Stock Kernels

Containers that share a machine compete for CPU. When one slows down, the operator needs to know which co-tenant is responsible, and no deployed signal can say. Pressure stall information, per-cgroup wait counters, and run-queue latency histograms are all victim-side: they report that a container waited, never who it waited for. Recovering the culprit means a kernel patch, full scheduler tracing, or statistical inference: unportable, too costly to leave on, or unreliable when victims coexist. SchedBlame is an eBPF tracer that attributes CPU contention to the cgroups that caused it, on stock kernels, continuously. It inverts the accounting: instead of measuring how long a victim waited, it measures the CPU time every other cgroup consumed while that victim was runnable but not running on the same CPU. The mechanism is a per-CPU bitmap of which measured cgroups are waiting, maintained from the kernel's own runnable counts at four scheduler hooks. Every run slice carries that bitmap, so one 16-byte record charges CPU time to a full row of a competitor x victim blame matrix; the kernel stores no per-pair state. Three properties follow. Slices are self-describing, so userspace holds no waiting state and a lost record costs measurements, not correctness. The measured set is reconfigured by publishing an epoch, invalidating every cache and per-CPU bitmap in constant time while the hooks keep running. Sampling never touches waiting state, so rescaling by the inverse keep probability keeps the estimator unbiased. SchedBlame splits each container's per-second CPU demand into runtime, internal contention, external contention, and throttling, flags anomalies against a rolling 99th-percentile baseline, and names the competitors responsible. In production on unmodified 4.18 and 5.10 kernels, tracking 84 containers on a 96-core host, it costs about 1% of Redis throughput and 6% of one core.

cs.OS

Observation of Hong-Ou-Mandel interference between photon and polariton

Light-matter interactions underlie many quantum technologies, yet whether quasiparticles formed from such interactions preserve the full quantum state of light remains unresolved. Surface plasmon polaritons (SPPs), a class of polaritons formed by interacting photons with free-electron oscillations at metal-dielectric interfaces, are prime candidates to explore this question. Here we demonstrate quantum interference between single photons and SPPs using an Au-SiN$_{\mathrm{x}}$ integrated photonic-plasmonic device. Our results reveal that SPPs retain the indistinguishability of their excitation photons, establishing SPP as a viable quantum information carrier and opening a potential route toward photonic-plasmonic quantum circuitry.

quant-ph

Converse and Collision-Based Achievability for Node Localization with Hybrid Distance-Spectral Graph Positional Encodings

Graph positional encodings are widely used in graph neural networks and graph Transformers, yet it remains unclear when the code itself can identify nodes. We study a hybrid distance-spectral encoding that combines anchor-distance profiles with quantized low-frequency Laplacian-energy coordinates. Treating the encoding as an observation map yields a simplex-refined converse, an exact collision factorization \(\kappa_H=\kappa_D\kappa_{S|D}\), and the collision information \(I_H=-\log\kappa_D-\log\kappa_{S|D}\). On random regular graphs, the criterion is made explicit through a bounded-correlation Gaussian-wave surrogate; for actual Laplacian-energy coordinates, we give the distance-conditioned spectral collision condition sufficient for conditional actual-coordinate achievability. Experiments show that \(I_H/\log n\) calibrates localization success, and PE-only structural task probes on Universal Dependencies trees show that hybrid encodings better recover syntactic-tree geometry than distance-only or spectral-only baselines.

cs.LG

Quantum-interference metrology of dissipative Kerr solitons

Dissipative Kerr solitons in optical microresonators underpin chip-scale frequency combs with applications ranging from coherent telecommunications to precision spectroscopy. Yet the characterization of their intrinsic femtosecond temporal structure remains challenging, as the low pulse energy and broad spectral bandwidth necessitate optical amplification and careful dispersion compensation in conventional ultrafast diagnostics, both of which can significantly distort the waveform. Here we demonstrate a quantum-interference metrology of microcomb solitons based on Hong-Ou-Mandel interference. By attenuating the soliton stream to the single-photon level and measuring fourth-order interference, we directly retrieve near transform-limited pulse durations without amplification or dispersion management, remaining accurate even after propagation through 25 km of standard fiber. The same interferogram also provides direct access to the temporal separations in multi-soliton states by converting inter-soliton separations into additional interference dips at corresponding delays, enabling sub-picosecond characterization of their intracavity temporal structure. This quantum-inspired paradigm introduces a fundamentally new metrological approach that is immune to amplification and dispersion distortions, offering a powerful tool for the characterization of complex soliton physics.

quant-ph

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly extended through plugin marketplaces, yet the structure, maintenance, and co-evolution dynamics of these emerging repositories remain empirically unexplored. Unlike traditional software packages that deliver functionality through source code, agent plugins deliver functionality through a combination of natural-language instruction files, scripts, and configuration files, raising the question of whether these plugins are maintained artifacts that co-evolve across components, or one-off artifacts that developers write once and do not need to revisit. To study the maintenance and co-evolution of agent plugins, we conduct an empirical study of 1,926 repositories hosting Claude Code plugin marketplaces, analyzing 8,351 plugins and 77,773 commits across 2,018 marketplaces. We find that the marketplace is expanding rapidly, plugin-touching commit activity growing 8.8x over six months after the October 2025 launch, and plugins targeting Software Engineering tasks accounting for 61.3% of all plugins. Plugin development is predominantly feature-driven, with feature commits occurring at more than twice the rate of conventional open-source software (OSS) (39.6% vs. 17.2%). Claude co-authors 34.9% of all commits, and four commit types (docs, perf, style, and refactor) carry substantially different meanings in plugin repositories than in traditional software. Most component types evolve independently, but within skills directories, natural-language instruction files and implementation scripts co-evolve at above-chance rates, with 78% of co-changes being functionally coupled, representing a new class of maintenance dependency not observed in traditional software engineering.

cs.SE

Fiber Bragg Grating Whiskers for Bioinspired Hydrodynamic Perception on Underwater Robots

Harbor seals track hydrodynamic trails with their vibrissae, enabling passive perception of moving targets in dark or turbid water. Inspired by this capability, we present compact fiber Bragg grating (FBG) whiskers for underwater robots. Like seal whiskers, they have a non-uniform taper and elliptical cross-section. Controlled towing experiments show a monotonic relative-flow response from 0.1 to 0.6 m/s, a strong reduction of self-induced oscillation relative to a cylindrical baseline, and a pronounced dependence on angle of attack. Experiments with a pitching foil show that the whiskers can detect the characteristic vortices shed by a stationary or moving source, detectable several seconds after the source has passed. Using this information, a single front-mounted whisker enabled a small underwater robot to distinguish between continuing straight and executing a turn, selecting the correct branch in 17 of 20 trials (85.0%) from whisker signals alone. These results connect bioinspired hydrodynamic sensing to robot action and suggest the utility of whiskers for tracking underwater objects.

cs.RO

LoViF 2026 The First Challenge on Unified Removal of Raindrops and Reflections: Methods and Results

This workshop paper comprehensively reviews the First Challenge on Unified Removal of Raindrops and Reflections. The challenge aims to address a frequently encountered practical problem in the field of autonomous driving, i.e., raindrop-reflection composite degradation on rainy days. This competition attracted 149 registered participants and received 12 valid final submissions with corresponding fact sheets, significantly contributing to the progress of unified removal of raindrops and reflections. All the methods are developed and evaluated on our real-shot RainDrop and ReFlection (RDRF) dataset. A detailed analysis of the submitted methods and corresponding results is provided in this report, which highlights effective approaches and provides interesting insights for future research.

cs.CV

A Novel Decoupling Method for Investigating Distinct Domain Evolution in Ferroelectric Film

Conventional polarization characterization techniques provide only the averaged response of ferroelectric films, limiting the investigation of domain-dependent reliability mechanisms in HfO2-based ferroelectrics. In this work, a new domain decoupling method is proposed to separately analyze ferroelectric domains with distinct switching behaviors. A three-domain model consisting of upward non-switchable domains, downward non-switchable domains, and switchable domains (Pd) is first introduced to describe heterogeneous domain populations during electrical cycling. By combining complementary switching-current measurements, the responses of different domain populations can be selectively extracted and reconstructed. The proposed method enables quantitative tracking of individual domain populations during cycling. As a demonstration, the method is further applied to analyze the time-dependent evolution of domain populations after electrical cycling. This approach provides a new route for investigating imprint, fatigue, and other reliability-related phenomena in ferroelectric devices.

cond-mat.mtrl-sci

Geometric Structures on Graphs: a Holonomy-Based Discretization of Curvature

We propose a holonomy-based framework for discretizing curvature on graphs equipped with local symmetric positive-definite metrics. Each vertex carries a fibre metric \(g_i\), and each directed edge carries a reversible metric-compatible transport \(F_{ij}\). The ordered product around an oriented triangular loop \(\mathcal C\) gives a holonomy \(H_{\mathcal C}\), whose normalized logarithm \(\Omega_{\mathcal C}=-s_{\mathcal C}^{-1}\operatorname{Log}(H_{\mathcal C})\) is used as a finite-loop curvature observation. Thus the construction discretizes the geometric principle that infinitesimal holonomy is controlled by curvature, rather than treating holonomy as a heuristic feature. Since \(\Omega_{\mathcal C}\) lies in the \(g_i\)-orthogonal Lie algebra, it is not itself a velocity of an SPD metric. We therefore introduce two aggregation mechanisms: a commutator with a symmetric response matrix, producing symmetric Ricci-type metric responses, and an incidence-aware covariant divergence of curvature-induced edge fluxes, reflecting the relation between trace and covariant divergence. The resulting responses are locally orthogonal-gauge equivariant and can drive exponential updates that preserve positive definiteness. We also give a reversible metric-compatible parametrization of edge transports, allowing orthogonal edge factors, loop scales, weights, and response matrices to be learned while respecting the graph geometry. Known-geometry calibrations on the unit sphere test the holonomy--curvature relation, curvature preservation under nontrivial local metric representations, and the empirical recovery of edge transports from local observations.

cs.LG

Generation of TeV Photons by PeV Neutrinos in Dense Astrophysical Environments

Recent observations by IceCube and KM3Net of PeV-scale ultra-high-energy (UHE) neutrinos, together with detections of TeV-PeV photons from various sources such as the Crab Nebula, the Galactic Center, and gamma-ray burst by ground-based observatories including Tibet AS$\gamma$, MAGIC, Carpet-3, and LHAASO, point to the existence of extreme astrophysical environments capable of accelerating particles to ultra-high energies. These findings motivate investigations of possible connections between UHE neutrinos and photons in such environments. Theoretically, dense regions surrounding compact objects can efficiently produce UHE neutrinos. In this work, we calculate the production of UHE photons from neutrino-nucleon interactions, and note that if these interactions occur in the outer, optically thin regions of dense environments, the resulting photons could potentially be observed. In our model, an incident neutrino scatters off a nucleon, generating secondary partons that hadronize into pions and subsequently decay into UHE photons. We calculate the resulting photon energy spectra and find that for incident (anti)neutrinos with energies above 1 PeV, the probability of producing photons with energies exceeding 1 TeV is greater than 13%. As a concrete application, we show that this mechanism can quantitatively account for the preburst TeV photons observed in GRB 221009A, providing a natural explanation for both their energies and lead times. These findings establish a plausible mechanism linking UHE neutrino events to gamma-ray observations, providing new insights into hadronic processes in extreme astrophysical environments and supporting multi-messenger astronomy studies.

astro-ph.HE