SearcharxivSearch

arXiv subjects

Jiaming Wang

Publications and source records attributed to Jiaming Wang.

At least 55 records · Page 3Linked to original sources

$1/f$ Noise in the Heliosphere: A Target for PUNCH Science

We present a broad review of 1/f noise observations in the heliosphere, and discuss and complement the theoretical background of generic 1/f models as relevant to NASA's PUNCH mission. First observed in the voltage fluctuations of vacuum tubes, the scale-invariant 1/f spectrum has since been identified across a wide array of natural and artificial systems, including heart rate fluctuations and loudness patterns in musical compositions. In the solar wind, the interplanetary magnetic field trace spectrum exhibits 1/f scaling within the frequency range from around 2e-6 Hz to around 1e-3 Hz at 1 au. One compelling mechanism for the generation of 1/f noise is the superposition principle, where a composite 1/f spectrum arises from the superposition of a collection of individual power-law spectra characterized by a scale-invariant distribution of correlation times. In the context of the solar wind, such a superposition could originate from scale-invariant reconnection processes in the corona. Further observations have detected 1/f signatures in the photosphere and corona at frequency ranges compatible with those observed at 1 au, suggesting an even lower altitude origin of 1/f spectrum in the solar dynamo itself. This hypothesis is bolstered by dynamo experiments and simulations that indicate inverse cascade activities, which can be linked to successive flux tube reconnections beneath the corona, and are known to generate 1/f noise possibly through nonlocal interactions at the largest scales. Conversely, models positing in situ generation of $1/f$ signals face causality issues in explaining the low-frequency portion of the 1/f spectrum. Understanding 1/f noise in the solar wind may inform central problems in heliospheric physics, such as the solar dynamo, coronal heating, the origin of the solar wind, and the nature of interplanetary turbulence.

astro-ph.SR

ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis

Deep learning has excelled in medical image classification, but its clinical application is limited by poor interpretability. Capsule networks, known for encoding hierarchical relationships and spatial features, show potential in addressing this issue. Nevertheless, traditional capsule networks often underperform due to their shallow structures, and deeper variants lack hierarchical architectures, thereby compromising interpretability. This paper introduces a novel capsule network, ParseCaps, which utilizes the sparse axial attention routing and parse convolutional capsule layer to form a parse-tree-like structure, enhancing both depth and interpretability. Firstly, sparse axial attention routing optimizes connections between child and parent capsules, as well as emphasizes the weight distribution across instantiation parameters of parent capsules. Secondly, the parse convolutional capsule layer generates capsule predictions aligning with the parse tree. Finally, based on the loss design that is effective whether concept ground truth exists or not, ParseCaps advances interpretability by associating each dimension of the global capsule with a comprehensible concept, thereby facilitating clinician trust and understanding of the model's classification results. Experimental results on CE-MRI, PH$^2$, and Derm7pt datasets show that ParseCaps not only outperforms other capsule network variants in classification accuracy, redundancy reduction and robustness, but also provides interpretable explanations, regardless of the availability of concept labels.

cs.CV

Observed Fluctuation Enhancement and Departure from WKB Theory in Sub-Alfvénic Solar Wind

Using Parker Solar Probe data from orbits 8 through 17, we examine fluctuation amplitudes throughout the critical region where the solar wind flow speed approaches and then exceeds the Alfvén wave speed, taking account of various exigencies of the plasma data. In contrast to WKB theory for non-interacting Alfvén waves streaming away from the Sun, the magnetic and kinetic fluctuation energies per unit volume are not monotonically decreasing. Instead, there is clear violation of conservation of standard WKB wave action, which is consistent with previous indications of strong in-situ fluctuation energy input in the solar wind near the Alfvén critical region. This points to strong violations of WKB theory due to nonlinearity (turbulence) and major energy input near the critical region, which we interpret as likely due to driving by large-scale coronal shear flows.

astro-ph.SR

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLMs). Previous mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text translation, and speech enhancement over models using continuous speech features. In this paper, we propose LauraGPT, a novel unified audio-and-text GPT-based LLM for audio recognition, understanding, and generation. LauraGPT is a versatile LLM that can process both audio and text inputs and generate outputs in either modalities. We propose a novel data representation that combines continuous and discrete features for audio: LauraGPT encodes input audio into continuous representations using an audio encoder and generates output audio from discrete codec codes. We propose a one-step codec vocoder to overcome the prediction challenge caused by the multimodal distribution of codec tokens. We fine-tune LauraGPT using supervised multi-task learning. Extensive experiments show that LauraGPT consistently achieves comparable to superior performance compared to strong baselines on a wide range of audio tasks related to content, semantics, paralinguistics, and audio-signal analysis, such as automatic speech recognition, speech-to-text translation, text-to-speech synthesis, speech enhancement, automated audio captioning, speech emotion recognition, and spoken language understanding.

cs.SD

The Alfvén Transition Zone observed by the Parker Solar Probe in Young Solar Wind -- Global Properties and Model Comparisons

The transition from subAlfvénic to superAlfvénic flow in the solar atmosphere is examined by means of Parker Solar Probe (PSP) measurements during solar encounters 8 to 14. Around 220 subAlfvénic periods with a duration $\ge$ 10 minutes are identified. The distribution of their durations, heliocentric distances, and Alfvén Mach number are analyzed and compared with a global magnetohydrodynamic model of the solar corona and wind, which includes turbulence effects. The results are consistent with a patchy and fragmented morphology, and suggestive of a turbulent Alfvén zone within which the transition from subAlfvénic to superAlfvénic flow occurs over an extended range of helioradii. These results inform and establish context for detailed analyses of subAlfvénic coronal plasma that are expected to emerge from PSP's final mission phase, as well as for NASA's planned PUNCH mission.

astro-ph.SR

Anisotropy of Density Fluctuations in the Solar Wind at 1 au

A well-known property of solar wind plasma turbulence is the observed anisotropy of the autocorrelations, or equivalently the spectra, of velocity and magnetic field fluctuations. Here we explore the related but apparently not well-studied issue of the anisotropy of plasma density fluctuations in the energy-containing and inertial ranges of solar wind turbulence. Using 10 years (1998-2008) of in situ data from the Advanced Composition Explorer (ACE) mission, we find that for all but the fastest wind category, the density correlation scale is slightly larger in directions quasi-parallel to the large-scale mean magnetic field as compared to quasi-perpendicular directions. The correlation scale in fast wind is consistent with isotropic. The anisotropy as a function of the level of correlation is also explored. We find at small correlation levels, i.e., at energy-containing scales and larger, the density fluctuations are close to isotropy for fast wind, and slightly favor more rapid decorrelation in perpendicular directions for slow and medium winds. At relatively smaller (inertial range) scales where the correlation values are larger, the sense of anisotropy is reversed in all speed ranges, implying a more "slab-like" structure, especially prominent in the fast wind samples. We contrast this finding with published results on velocity and magnetic field correlations.

physics.space-ph

Nonlinear Kalman Filtering based on Self-Attention Mechanism and Lattice Trajectory Piecewise Linear Approximation

The traditional Kalman filter (KF) is widely applied in control systems, but it relies heavily on the accuracy of the system model and noise parameters, leading to potential performance degradation when facing inaccuracies. To address this issue, introducing neural networks into the KF framework offers a data-driven solution to compensate for these inaccuracies, improving the filter's performance while maintaining interpretability. Nevertheless, existing studies mostly employ recurrent neural network (RNN), which fails to fully capture the dependencies among state sequences and lead to an unstable training process. In this paper, we propose a novel Kalman filtering algorithm named the attention Kalman filter (AtKF), which incorporates a self-attention network to capture the dependencies among state sequences. To address the instability in the recursive training process, a parallel pre-training strategy is devised. Specifically, this strategy involves piecewise linearizing the system via lattice trajectory piecewise linear (LTPWL) expression, and generating pre-training data through a batch estimation algorithm, which exploits the self-attention mechanism's parallel processing ability. Experimental results on a two-dimensional nonlinear system demonstrate that AtKF outperforms other filters under noise disturbances and model mismatches.

eess.SY

OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning

Redundancy is a persistent challenge in Capsule Networks (CapsNet),leading to high computational costs and parameter counts. Although previous works have introduced pruning after the initial capsule layer, dynamic routing's fully connected nature and non-orthogonal weight matrices reintroduce redundancy in deeper layers. Besides, dynamic routing requires iterating to converge, further increasing computational demands. In this paper, we propose an Orthogonal Capsule Network (OrthCaps) to reduce redundancy, improve routing performance and decrease parameter counts. Firstly, an efficient pruned capsule layer is introduced to discard redundant capsules. Secondly, dynamic routing is replaced with orthogonal sparse attention routing, eliminating the need for iterations and fully connected structures. Lastly, weight matrices during routing are orthogonalized to sustain low capsule similarity, which is the first approach to introduce orthogonality into CapsNet as far as we know. Our experiments on baseline datasets affirm the efficiency and robustness of OrthCaps in classification tasks, in which ablation studies validate the criticality of each component. Remarkably, OrthCaps-Shallow outperforms other Capsule Network benchmarks on four datasets, utilizing only 110k parameters, which is a mere 1.25% of a standard Capsule Network's total. To the best of our knowledge, it achieves the smallest parameter count among existing Capsule Networks. Similarly, OrthCaps-Deep demonstrates competitive performance across four datasets, utilizing only 1.2% of the parameters required by its counterparts.

cs.CV

An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task. To be more specific, we benchmark and explore various combinations of LLMs and speech encoders, leading to the optimal LLM-based ASR system, which we call SLAM-ASR. The proposed SLAM-ASR provides a clean setup and little task-specific design, where only the linear projector is trained. To the best of our knowledge, SLAM-ASR achieves the best performance on the Librispeech benchmark among LLM-based ASR models and even outperforms the latest LLM-based audio-universal model trained on massive pair data. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community.

cs.CL

TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

Recently, end-to-end neural diarization (EEND) is introduced and achieves promising results in speaker-overlapped scenarios. In EEND, speaker diarization is formulated as a multi-label prediction problem, where speaker activities are estimated independently and their dependency are not well considered. To overcome these disadvantages, we employ the power set encoding to reformulate speaker diarization as a single-label classification problem and propose the overlap-aware EEND (EEND-OLA) model, in which speaker overlaps and dependency can be modeled explicitly. Inspired by the success of two-stage hybrid systems, we further propose a novel Two-stage OverLap-aware Diarization framework (TOLD) by involving a speaker overlap-aware post-processing (SOAP) model to iteratively refine the diarization results of EEND-OLA. Experimental results show that, compared with the original EEND, the proposed EEND-OLA achieves a 14.39% relative improvement in terms of diarization error rates (DER), and utilizing SOAP provides another 19.33% relative improvement. As a result, our method TOLD achieves a DER of 10.14% on the CALLHOME dataset, which is a new state-of-the-art result on this benchmark to the best of our knowledge.

cs.SD

Local probe investigation of the spin dynamics in the kagome and inter-layers of orthorhombic barlowite Cu$_4$(OD)$_6$FBr: $^{79}$Br and $^{63}$Cu NQR study

We report $^{79}$Br and $^{63}$Cu nuclear quadrupole resonance (NQR) in the paramagnetic state above $T_\text{N} = 15$ K of the antiferromagnetic orthorhombic phase of barlowite Cu$_4$(OD)$_6$FBr consisting of a layered kagome structure. The divergent behavior of the longitudinal $^{79}(1/T_{1})$ and transverse $^{79}(1/T_{2})$ relaxation rates observed at $^{79}$Br sites evidences that critical slowing down of Cu spin fluctuations sets in below $\sim20$ K. This means that one or more Cu sites, most likely at the interlayer Cu(3,4,5) sites between the kagome planes, undergo the antiferromagnetic phase transition in a fairly conventional way. On the other hand, the $^{63}$Cu NQR signal intensity is gradually wiped out below $\sim30$ K, pointing toward gradual spin freezing of the kagome layers instead. These contrasting findings suggest significant roles played by magnetic frustration effects within the kagome layers.

cond-mat.str-el

Probable Object Location (POLo) Score Estimation for Efficient Object Goal Navigation

To advance the field of autonomous robotics, particularly in object search tasks within unexplored environments, we introduce a novel framework centered around the Probable Object Location (POLo) score. Utilizing a 3D object probability map, the POLo score allows the agent to make data-driven decisions for efficient object search. We further enhance the framework's practicality by introducing POLoNet, a neural network trained to approximate the computationally intensive POLo score. Our approach addresses critical limitations of both end-to-end reinforcement learning methods, which suffer from memory decay over long-horizon tasks, and traditional map-based methods that neglect visibility constraints. Our experiments, involving the first phase of the OVMM 2023 challenge, demonstrate that an agent equipped with POLoNet significantly outperforms a range of baseline methods, including end-to-end RL techniques and prior map-based strategies. To provide a comprehensive evaluation, we introduce new performance metrics that offer insights into the efficiency and effectiveness of various agents in object goal navigation.

cs.RO

Deep Reinforcement Learning Based Framework for Mobile Energy Disseminator Dispatching to Charge On-the-Road Electric Vehicles

The exponential growth of electric vehicles (EVs) presents novel challenges in preserving battery health and in addressing the persistent problem of vehicle range anxiety. To address these concerns, wireless charging, particularly, Mobile Energy Disseminators (MEDs) have emerged as a promising solution. The MED is mounted behind a large vehicle and charges all participating EVs within a radius upstream of it. Unfortuantely, during such V2V charging, the MED and EVs inadvertently form platoons, thereby occupying multiple lanes and impairing overall corridor travel efficiency. In addition, constrained budgets for MED deployment necessitate the development of an effective dispatching strategy to determine optimal timing and locations for introducing the MEDs into traffic. This paper proposes a deep reinforcement learning (DRL) based methodology to develop a vehicle dispatching framework. In the first component of the framework, we develop a realistic reinforcement learning environment termed "ChargingEnv" which incorporates a reliable charging simulation system that accounts for common practical issues in wireless charging deployment, specifically, the charging panel misalignment. The second component, the Proximal-Policy Optimization (PPO) agent, is trained to control MED dispatching through continuous interactions with ChargingEnv. Numerical experiments were carried out to demonstrate the demonstrate the efficacy of the proposed MED deployment decision processor. The experiment results suggest that the proposed model can significantly enhance EV travel range while efficiently deploying a optimal number of MEDs. The proposed model is found to be not only practical in its applicability but also has promises of real-world effectiveness. The proposed model can help travelers to maximize EV range and help road agencies or private-sector vendors to manage the deployment of MEDs efficiently.

cs.RO

kTrans: Knowledge-Aware Transformer for Binary Code Embedding

Binary Code Embedding (BCE) has important applications in various reverse engineering tasks such as binary code similarity detection, type recovery, control-flow recovery and data-flow analysis. Recent studies have shown that the Transformer model can comprehend the semantics of binary code to support downstream tasks. However, existing models overlooked the prior knowledge of assembly language. In this paper, we propose a novel Transformer-based approach, namely kTrans, to generate knowledge-aware binary code embedding. By feeding explicit knowledge as additional inputs to the Transformer, and fusing implicit knowledge with a novel pre-training task, kTrans provides a new perspective to incorporating domain knowledge into a Transformer framework. We inspect the generated embeddings with outlier detection and visualization, and also apply kTrans to 3 downstream tasks: Binary Code Similarity Detection (BCSD), Function Type Recovery (FTR) and Indirect Call Recognition (ICR). Evaluation results show that kTrans can generate high-quality binary code embeddings, and outperforms state-of-the-art (SOTA) approaches on downstream tasks by 5.2%, 6.8%, and 12.6% respectively. kTrans is publicly available at: https://github.com/Learner0x5a/kTrans-release

cs.SE

FunASR: A Fundamental End-to-End Speech Recognition Toolkit

This paper introduces FunASR, an open-source speech recognition toolkit designed to bridge the gap between academic research and industrial applications. FunASR offers models trained on large-scale industrial corpora and the ability to deploy them in applications. The toolkit's flagship model, Paraformer, is a non-autoregressive end-to-end speech recognition model that has been trained on a manually annotated Mandarin speech recognition dataset that contains 60,000 hours of speech. To improve the performance of Paraformer, we have added timestamp prediction and hotword customization capabilities to the standard Paraformer backbone. In addition, to facilitate model deployment, we have open-sourced a voice activity detection model based on the Feedforward Sequential Memory Network (FSMN-VAD) and a text post-processing punctuation model based on the controllable time-delay Transformer (CT-Transformer), both of which were trained on industrial corpora. These functional modules provide a solid foundation for building high-precision long audio speech recognition services. Compared to other models trained on open datasets, Paraformer demonstrates superior performance.

cs.SD

Validation of the plasma-wall self-organization model for density limit in ECRH-assisted start-up of Ohmic discharges on J-TEXT

A recently developed plasma-wall self-organization (PWSO) model predicts a significantly enhanced density limit, which may be attainable in tokamaks with ECRH-assisted ohmic startup and sufficiently high initial neutral density. Experiments have been conducted on J-TEXT to validate such a density limit scenario based on this model. Experimental results demonstrate that increasing the pre-filled gas pressure or ECRH power during the startup phase can effectively enhance plasma purity and raise the density limit at the flat-top. Despite the dominant carbon fraction in the wall material, some discharges approach the edge of the density-free regime of the 1D model of PWSO.

physics.plasm-ph

MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition

In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlabeled speech and text data. The main difficulty in speech-text joint pre-training comes from the significant difference between speech and text modalities, especially for Mandarin speech and text. Unlike English and other languages with an alphabetic writing system, Mandarin uses an ideographic writing system where character and sound are not tightly mapped to one another. Therefore, we propose to introduce the phoneme modality into pre-training, which can help capture modality-invariant information between Mandarin speech and text. Specifically, we employ a multi-task learning framework including five self-supervised and supervised tasks with speech and text data. For end-to-end pre-training, we introduce self-supervised speech-to-pseudo-codes (S2C) and phoneme-to-text (P2T) tasks utilizing unlabeled speech and text data, where speech-pseudo-codes pairs and phoneme-text pairs are a supplement to the supervised speech-text pairs. To train the encoder to learn better speech representation, we introduce self-supervised masked speech prediction (MSP) and supervised phoneme prediction (PP) tasks to learn to map speech into phonemes. Besides, we directly add the downstream supervised speech-to-text (S2T) task into the pre-training process, which can further improve the pre-training performance and achieve better recognition results even without fine-tuning. Experiments on AISHELL-1 show that our proposed method achieves state-of-the-art performance, with a more than 40% relative improvement compared with other pre-training methods.

cs.MM

Emergence of the spin polarized domains in the kagome lattice Heisenberg antiferromagnet Zn-barlowite (Zn$_{0.95}$Cu$_{0.05}$)Cu$_{3}$(OD)$_{6}$FBr

Kagome lattice Heisenberg antiferromagnets are known to be highly sensitive to perturbations caused by structural disorder. NMR is a local probe ideally suited for investigating such disorder-induced effects, but in practice large distributions in the conventional one-dimensional NMR data make it difficult to distinguish the intrinsic behavior expected for pristine kagome quantum spin liquids from disorder induced effects. Here we report the development of a two-dimensional NMR data acquisition scheme applied to Zn-barlowite (Zn$_{0.95}$Cu$_{0.05}$)Cu$_{3}$(OD)$_{6}$FBr kagome lattice, and successfully correlate the distribution of the low energy spin excitations with that of the local spin susceptibility. We present evidence for the gradual growth of domains with a local spin polarization induced by 5\% Cu$^{2+}$ defect spins occupying the interlayer non-magnetic Zn$^{2+}$ sites. These spin polarized domains account for $\sim60$\% of the sample volume at 2~K, where gapless excitations induced by interlayer defects dominate the low energy sector of spin excitations within the kagome planes.

cond-mat.str-el