SearcharxivSearch

arXiv subjects

Sheng Jiang

Publications and source records attributed to Sheng Jiang.

At least 19 recordsLinked to original sources

Post-Corrected Raw-Score Martingale Posterior Sampling for von Mises-Fisher Models

We develop a finite-horizon calibration method for raw-score martingale posteriors, with von Mises--Fisher models as the main worked example. Starting from the maximum likelihood estimator, predictive paths are generated by simulating future observations from the current fitted model and updating the natural parameter by unpreconditioned score increments. The main methodological step is to separate predictive simulation from covariance calibration. Raw-score increments have Fisher-information covariance, whereas Bernstein--von Mises calibration requires inverse-information covariance. We therefore apply a terminal linear correction based on a local information estimate. For more efficient implementation, we also introduce a hybrid version that replaces the omitted tail of the infinite predictive continuation by a Gaussian approximation with matching leading-order quadratic variation. We prove fixed-n convergence, a finite-horizon approximation bound for the Gaussian tail, and a Bernstein--von Mises limit for the hybrid post-corrected sampler under local regularity and consistent terminal calibration. Simulations show that tail correction reduces truncation-induced underdispersion, and an OSCAR ocean-current example illustrates local directional uncertainty summaries.

stat.ME

Martingale Posteriors for Discretely Observed Diffusions

In this paper we consider parameter estimation for discretely observed diffusion processes. In particular, we focus on data that are observed at low frequency and methodology that can estimate parameters with uncertainty quantification. Most statistical work in this domain develops advanced Markov chain Monte Carlo (MCMC) algorithms for sampling from the posterior of the parameters, a task which is often complicated by the fact that one seldom has access to the transition density of the diffusion process; one has to combine sophisticated MCMC methods which are robust to the required time discretization of the diffusion, which can yield expensive algorithms. We focus on developing the martingale posterior method for the context of interest, when one can only numerically approximate the transition density of the diffusion. Based on using types of diffusion bridges we introduce a new martingale posterior method for parameter estimation for discretely observed diffusion processes. We prove that this algorithm approximates, in some sense, the martingale posterior which has no time-discretization bias up-to $\mathcal{O}(\Delta)$ if $\Delta$ is the time discretization step. Our approach is illustrated on several examples, showing orders of magnitude speed up versus state-of-the-art MCMC algorithms.

stat.CO

MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark

Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly.

cs.CL

High Pressure Superconducting transition in Dihydride BiH$_2$ with Bismuth Open-Channel Framework

Metal hydrides MHx with low hydrogen content are not expected to show high-Tc superconductivity owing to the low hydrogen-derived electronic density of states at Fermi level and the limited hydrogen contribution to electron-phonon coupling strength. In this work, we report on the successful synthesis of a novel bismuth dihydride superconductor, Cmcm-BiH$_2$, at approximately 150 GPa, and the discovery of superconductivity with Tc about 62 K at 163 GPa, marking the first instance of superconductor among the MH$_2$-type metal dihydrides. Cmcm-BiH$_2$ adopts a unique host-guest type structure, in which the Bi atoms via weak Bi-Bi covalent bonds form a three-dimensional open-channel framework that encapsulates H$_2$-like molecules as guests, thereby broadening the structural diversity of hydrides under high pressures. The occurrence of superconductivity is evidenced by a sharp drop of resistivity to zero and the characteristic downward shift of Tc under applied magnetic fields. Notably, Cmcm-BiH$_2$ remains stable down to at least 97 GPa during decompression, with the calculated lowest pressure for dynamic stability of 10 GPa. In-depth analysis reveals that the covalent bismuth open-channel structure forms metallic conduction channels, dominates the electronic states near the Fermi level, and contributes approximately 51% of the total $lambda$ in Cmcm-BiH$_2$, distinguishing it from known high-pressure hydride superconductors. These findings highlight the critical role of non-hydrogen elements in producing superconductivity and open new avenues for the design and optimization of high-Tc hydride superconductors.

cond-mat.supr-con

DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative framework for audio-driven talking-portrait video synthesis that delivers disentangled, high-fidelity control of lip motion, head pose, and eye gaze. The core contribution is a motion auto-encoder that builds a structured latent space in which motion factors are independently represented and approximately orthogonalized. On this disentangled motion space, we apply optimal-transport-based flow matching with a transformer predictor to generate temporally smooth motion trajectories conditioned on audio. Extensive experiments across multiple benchmarks show that DEMO outperforms prior methods in video realism, lip-audio synchronization, and motion fidelity. These results demonstrate that combining fine-grained motion disentanglement with flow-based generative modeling provides a powerful new paradigm for controllable talking-head video synthesis.

cs.CV

LVLMs as inspectors: an agentic framework for category-level structural defect annotation

Automated structural defect annotation is essential for ensuring infrastructure safety while minimizing the high costs and inefficiencies of manual labeling. A novel agentic annotation framework, Agent-based Defect Pattern Tagger (ADPT), is introduced that integrates Large Vision-Language Models (LVLMs) with a semantic pattern matching module and an iterative self-questioning refinement mechanism. By leveraging optimized domain-specific prompting and a recursive verification process, ADPT transforms raw visual data into high-quality, semantically labeled defect datasets without any manual supervision. Experimental results demonstrate that ADPT achieves up to 98% accuracy in distinguishing defective from non-defective images, and 85%-98% annotation accuracy across four defect categories under class-balanced settings, with 80%-92% accuracy on class-imbalanced datasets. The framework offers a scalable and cost-effective solution for high-fidelity dataset construction, providing strong support for downstream tasks such as transfer learning and domain adaptation in structural damage assessment.

cs.CV

Femtosecond Engineering of magnetic Domain Walls via Nonequilibrium Spin Textures

Ultrafast optical control of magnetic textures offers new opportunities for energy-efficient, high-speed spintronic devices. While uniform magnetization reversal via all-optical switching is well established, the formation dynamics of non-uniform domain walls (DWs) under ultrafast excitation remain poorly understood. Here, we use Lorentz ultrafast electron microscopy combined with transient optical grating excitation to directly image the real-time formation of DWs in a ferrimagnetic GdFeCo film. We observe a rapid evolution from disordered spin contrast to ordered DW arrays within 10 ps, including a transient, strongly asymmetric DW state. In a narrow fluence window, short-lived DWs form and spontaneously vanish within picoseconds. Multiscale simulations combining atomistic spin dynamics and micromagnetics reveal a nonlinear nucleation pathway involving a hybrid transition state where localized, unstable spin textures coalesce into metastable DWs. This nonequilibrium mechanism explains the observed asymmetry and spatial ordering, and establishes a framework for controlling spin textures in magnetic materials on femtosecond timescales.

physics.app-ph

Cross-Layer Encrypted Semantic Communication Framework for Panoramic Video Transmission

In this paper, we propose a cross-layer encrypted semantic communication (CLESC) framework for panoramic video transmission, incorporating feature extraction, encoding, encryption, cyclic redundancy check (CRC), and retransmission processes to achieve compatibility between semantic communication and traditional communication systems. Additionally, we propose an adaptive cross-layer transmission mechanism that dynamically adjusts CRC, channel coding, and retransmission schemes based on the importance of semantic information. This ensures that important information is prioritized under poor transmission conditions. To verify the aforementioned framework, we also design an end-to-end adaptive panoramic video semantic transmission (APVST) network that leverages a deep joint source-channel coding (Deep JSCC) structure and attention mechanism, integrated with a latitude adaptive module that facilitates adaptive semantic feature extraction and variable-length encoding of panoramic videos. The proposed CLESC is also applicable to the transmission of other modal data. Simulation results demonstrate that the proposed CLESC effectively achieves compatibility and adaptation between semantic communication and traditional communication systems, improving both transmission efficiency and channel adaptability. Compared to traditional cross-layer transmission schemes, the CLESC framework can reduce bandwidth consumption by 85% while showing significant advantages under low signal-to-noise ratio (SNR) conditions.

eess.IV

Integration of Large Vision Language Models for Efficient Post-disaster Damage Assessment and Reporting

Traditional natural disaster response involves significant coordinated teamwork where speed and efficiency are key. Nonetheless, human limitations can delay critical actions and inadvertently increase human and economic losses. Agentic Large Vision Language Models (LVLMs) offer a new avenue to address this challenge, with the potential for substantial socio-economic impact, particularly by improving resilience and resource access in underdeveloped regions. We introduce DisasTeller, the first multi-LVLM-powered framework designed to automate tasks in post-disaster management, including on-site assessment, emergency alerts, resource allocation, and recovery planning. By coordinating four specialised LVLM agents with GPT-4 as the core model, DisasTeller autonomously implements disaster response activities, reducing human execution time and optimising resource distribution. Our evaluations through both LVLMs and humans demonstrate DisasTeller's effectiveness in streamlining disaster response. This framework not only supports expert teams but also simplifies access to disaster management processes for non-experts, bridging the gap between traditional response methods and LVLM-driven efficiency.

cs.MA

Vision Mamba-based autonomous crack segmentation on concrete, asphalt, and masonry surfaces

Convolutional neural networks (CNNs) and Transformers have shown advanced accuracy in crack detection under certain conditions. Yet, the fixed local attention can compromise the generalisation of CNNs, and the quadratic complexity of the global self-attention restricts the practical deployment of Transformers. Given the emergence of the new-generation architecture of Mamba, this paper proposes a Vision Mamba (VMamba)-based framework for crack segmentation on concrete, asphalt, and masonry surfaces, with high accuracy, generalisation, and less computational complexity. Having 15.6% - 74.5% fewer parameters, the encoder-decoder network integrated with VMamba could obtain up to 2.8% higher mDS than representative CNN-based models while showing about the same performance as Transformer-based models. Moreover, the VMamba-based encoder-decoder network could process high-resolution image input with up to 90.6% lower floating-point operations.

cs.CV

Robust feature knowledge distillation for enhanced performance of lightweight crack segmentation models

Vision-based crack detection faces deployment challenges due to the size of robust models and edge device limitations. These can be addressed with lightweight models trained with knowledge distillation (KD). However, state-of-the-art (SOTA) KD methods compromise anti-noise robustness. This paper develops Robust Feature Knowledge Distillation (RFKD), a framework to improve robustness while retaining the precision of light models for crack segmentation. RFKD distils knowledge from a teacher model's logit layers and intermediate feature maps while leveraging mixed clean and noisy images to transfer robust patterns to the student model, improving its precision, generalisation, and anti-noise performance. To validate the proposed RFKD, a lightweight crack segmentation model, PoolingCrack Tiny (PCT), with only 0.5 M parameters, is also designed and used as the student to run the framework. The results show a significant enhancement in noisy images, with RFKD reaching a 62% enhanced mean Dice score (mDS) compared to SOTA KD methods.

cs.CV

A BiRGAT Model for Multi-intent Spoken Language Understanding with Hierarchical Semantic Frames

Previous work on spoken language understanding (SLU) mainly focuses on single-intent settings, where each input utterance merely contains one user intent. This configuration significantly limits the surface form of user utterances and the capacity of output semantics. In this work, we first propose a Multi-Intent dataset which is collected from a realistic in-Vehicle dialogue System, called MIVS. The target semantic frame is organized in a 3-layer hierarchical structure to tackle the alignment and assignment problems in multi-intent cases. Accordingly, we devise a BiRGAT model to encode the hierarchy of ontology items, the backbone of which is a dual relational graph attention network. Coupled with the 3-way pointer-generator decoder, our method outperforms traditional sequence labeling and classification-based schemes by a large margin.

cs.CL

Magnetic droplet solitons

Magnetic droplets are nanoscale, non-topological, dynamical solitons that can be nucleated in different spintronic devices, such as spin torque nano-oscillators (STNOs) and spin Hall nano-oscillators (SHNOs). This chapter first briefly discusses the theory of spin current driven dissipative magnetic droplets in ferromagnetic thin films with uniaxial anisotropy. We then thoroughly review the research literature on magnetic droplets and their salient features, as measured using electrical, microwave, and synchrotron techniques, and as envisaged by micromagnetic simulations. We also touch upon a closely related soliton, the dynamical skyrmion. Finally, we present an outlook of new routes in droplet science.

cond-mat.mes-hall

BPF-oF: Storage Function Pushdown Over the Network

Storage disaggregation, wherein storage is accessed over the network, is popular because it allows applications to independently scale storage capacity and bandwidth based on dynamic application demand. However, the added network processing introduced by disaggregation can consume significant CPU resources. In many storage systems, logical storage operations (e.g., lookups, aggregations) involve a series of simple but dependent I/O access patterns. Therefore, one way to reduce the network processing overhead is to execute dependent series of I/O accesses at the remote storage server, reducing the back-and-forth communication between the storage layer and the application. We refer to this approach as \emph{remote-storage pushdown}. We present BPF-oF, a new remote-storage pushdown protocol built on top of NVMe-oF, which enables applications to safely push custom eBPF storage functions to a remote storage server. The main challenge in integrating BPF-oF with storage systems is preserving the benefits of their client-based in-memory caches. We address this challenge by designing novel caching techniques for storage pushdown, including splitting queries into separate in-memory and remote-storage phases and periodically refreshing the client cache with sampled accesses from the remote storage device. We demonstrate the utility of BPF-oF by integrating it with three storage systems, including RocksDB, a popular persistent key-value store that has no existing storage pushdown capability. We show BPF-oF provides significant speedups in all three systems when accessed over the network, for example improving RocksDB's throughput by up to 2.8$\times$ and tail latency by up to 2.6$\times$.

cs.OS

Coexisting and interacting spin torque driven free and reference layer magnetic droplet solitons

Magnetic droplets are nanoscale, non-topological, magnetodynamical solitons that can be nucleated in spin torque nano-oscillators (STNOs) or spin Hall nano-oscillators (SHNOs). All theoretical, numerical, and experimental droplet studies have so far focused on the free layer (FL), and any additional dynamics in the reference layer (RL) have been entirely ignored. Here we show, using all-perpendicular STNOs, that there is not only significant magnetodynamics in the RL, but the reference layer itself can host a droplet coexisting with the FL droplet. Both droplets are observed experimentally as stepwise changes and sharp peaks in the dc and differential resistance, respectively. Whereas the single FL droplet is highly stable, the coexistence state exhibits high-power broadband microwave noise. Micromagnetic simulations corroborate the experimental results and reveal a strong interaction between the droplets. Our demonstration of strongly interacting and closely spaced droplets offers a unique platform for fundamental studies of highly non-linear soliton pair dynamics.

cond-mat.mes-hall

Pressure-Induced Superconductivity and Its Scaling with Doping-Induced Superconductivity in the Iron Pnictide with Skutterudite Intermediary Layers

The Ca10(PtnAs8)(Fe2As2)5 (n=3,4) compounds are a new type of iron pnictide superconductors whose structures consist of stacking Ca-PtnAs8-Ca-Fe2As2 layers in a unit cell. When n=3 (the 10-3-8 phase), the undoped compound is an antiferromagnetic (AFM) semiconductor, while, when n=4 (the 10-4-8 phase), the undoped compound is a superconductor with the transition temperature of 26K. Here we report the results of high-pressure studies on the 10-3-8 compound obtained through a combination of in-situ resistance, magnetic susceptibility, and Hall coefficient measurements. We find that its AFM order can be suppressed completely at 3.5 GPa and then superconducting state appears in the pressure range of 3.5-7 GPa. The pressure dependence of superconducting transition temperature displays a dome-like shape.

cond-mat.supr-con

Heavy-Tailed Density Estimation

A novel statistical method is proposed and investigated for estimating a heavy tailed density under mild smoothness assumptions. Statistical analyses of heavy-tailed distributions are susceptible to the problem of sparse information in the tail of the distribution getting washed away by unrelated features of a hefty bulk. The proposed Bayesian method avoids this problem by incorporating smoothness and tail regularization through a carefully specified semiparametric prior distribution, and is able to consistently estimate both the density function and its tail index at near minimax optimal rates of contraction. A joint, likelihood driven estimation of the bulk and the tail is shown to help improve uncertainty assessment in estimating the tail index parameter and offer more accurate and reliable estimates of the high tail quantiles compared to thresholding methods.

stat.ME

Field-free high-frequency exchange-spring spin-torque nano-oscillators

Spin-torque nano-oscillators (STNOs) are a type of nanoscale microwave auto-oscillators utilizing spin-torque to generate magnetodynamics with great promise for applications in microwaves, magnetic memory, and neuromorphic computing. Here, we report the first demonstration of exchange-spring STNOs, with an exchange-spring ([Co/Pd]-Co) reference layer and a perpendicular ([Co/Ni]) free layer. This magnetic configuration results in high-frequency (>10 GHz) microwave emission at a zero magnetic field and exchange-spring dynamics in the reference layer and the observation of magnetic droplet solitons in the free layer at different current polarities. Our demonstration of bipolar and field-free exchange-spring-based STNOs operating over a 20 GHz frequency range greatly extends the design freedom and functionality of the current STNO technology for energy-efficient high-frequency spintronic and neuromorphic applications.

physics.app-ph