Searcharxiv⌕ Search

arXiv subjects

Zhengtao Wang

Publications and source records attributed to Zhengtao Wang.

11 recordsLinked to original sources

EMoG: Emotion-Modulated Gait Generation for Expressive Humanoid Locomotion

Existing humanoid locomotion systems primarily focus on stability and task execution, while integrating expressiveness with explicit locomotion control remains challenging. We propose EMoG, an emotion-modulated gait generation framework for expressive humanoid locomotion. EMoG introduces an emotional-style code with continuously adjustable intensity. Conditioned on this code and physical commands, a lightweight MLP generates expressive, command-consistent periodic gait trajectories in real time, which are tracked by a unified reinforcement learning policy for physical execution. To support training, we collect a large-scale emotion-annotated gait dataset from professional performers and develop an automated pipeline to extract physically consistent periodic gait cycles. EMoG also integrates an LLM-based parser that converts free-form language into emotional style and motion parameters for interactive control. Experiments demonstrate that our system achieves continuous gait-style modulation with perceptible expressive cues while maintaining command tracking. EMoG provides a practical approach to parameterized emotional-style walking for human-robot interaction.

cs.RO↗

Rapid and high-sensitive NV-based microwave field imaging via digital lock-in amplification for on-chip microstrip diagnostics

High-resolution, high-sensitivity microwave (MW) magnetic field imaging is indispensable for non-destructive integrated circuit (IC) testing, radio-frequency device characterization, and spintronic research. Yet, the practical utility of these techniques is severely constrained by the pervasive challenge of isolating weak magnetic signatures from intense optical and electronic noise, which fundamentally limits both acquisition speed and detection sensitivity. Here, we overcome this barrier by introducing a wide-field imaging scheme based on an ensemble of diamond nitrogen-vacancy (NV) centers, synergistically combined with digital lock-in amplification (DLA). By exploiting digital demodulation, the DLA precisely extracts the MW-field response at a specific modulation frequency from background noise (e.g., laser intensity fluctuations), dramatically improving the signal-to-noise ratio (SNR). Consequently, our system attains a magnetic field sensitivity of 126 nT/$\sqrt(Hz)$. Critically, the unprecedented SNR permits a pixel dwell time of under one millisecond, allowing full-field images to be acquired within seconds-more than an order of magnitude faster than state-of-the-art NV-based wide-field techniques. This combination of speed, sensitivity, and micron-scale spatial resolution (1.6 $μ$m) paves the way for quasi-real-time, non-invasive diagnostics of dynamic MW devices and integrated circuits.

physics.optics↗

On-chip Radio Frequency Maser

Room-temperature solid-state masers offer exceptional frequency selectivity and ultra-low noise for weak-signal detection. However, their reliance on bulky metallic resonators has significantly hindered integration, miniaturization, and extension to lower frequencies. Here, we demonstrate the first on-chip radio-frequency maser operating at room temperature, exploiting optically pumped triplet states of pentacene. The device produces stimulated emission at 106.62 MHz and enables ultra-sensitive microwave magnetic-field detection with a sensitivity of ($\sim 10\,\rm{fT/\sqrt{Hz}}$), functioning simultaneously as a local oscillator and a sensor. By actively controlling microwave dissipation, we achieve efficient regulation of the maser output, revealing a key mechanism for tuning emission in open cavity-free systems. This work extends pentacene-based masers into the radio-frequency regime and establishes a highly integrated on-chip architecture for room-temperature masers, offering a new pathway toward portable quantum devices.

quant-ph↗

Detecting Axion Dark Matter with an Organic Molecular Maser

We present a novel quantum sensing approach to search for axion-electron interactions around the axion mass of 6 \mueV. In this region, laboratory searches are relatively scarce, and our direct experiment measuring the axion-electron coupling constant reaches the sensitivity of 8 \times 10^{-6} GeV^{-1}. The method, based on an organic molecular maser establishes a proof-of-principle for quantum-enhanced detection, with a corresponding magnetic field sensitivity of 0.85 fT/\sqrt{\rm{Hz}}. The methodology is generic and can be readily extended to other physical systems, further broadening its applicability in quantum sensing and dark matter searches.

hep-ph↗

Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

cs.CL↗

Kimi-Audio Technical Report

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.

eess.AS↗

MoonCast: High-Quality Zero-Shot Podcast Generation

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize natural podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To generate long audio, we adopt a long-context language model-based audio modeling approach utilizing large-scale long-context speech data. To enhance spontaneity, we utilize a podcast generation module to generate scripts with spontaneous details, which have been empirically shown to be as crucial as the text-to-speech modeling itself. Experiments demonstrate that MoonCast outperforms baselines, with particularly notable improvements in spontaneity and coherence.

eess.AS↗

MDAN: Multi-level Dependent Attention Network for Visual Emotion Analysis

Visual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the granularity of emotions increases, the affective gap increases as well. Existing deep approaches try to bridge the gap by directly learning discrimination among emotions globally in one shot without considering the hierarchical relationship among emotions at different affective levels and the affective level of emotions to be classified. In this paper, we present the Multi-level Dependent Attention Network (MDAN) with two branches, to leverage the emotion hierarchy and the correlation between different affective levels and semantic levels. The bottom-up branch directly learns emotions at the highest affective level and strictly follows the emotion hierarchy while predicting emotions at lower affective levels. In contrast, the top-down branch attempt to disentangle the affective gap by one-to-one mapping between semantic levels and affective levels, namely, Affective Semantic Mapping. At each semantic level, a local classifier learns discrimination among emotions at the corresponding affective level. Finally, We integrate global learning and local learning into a unified deep framework and optimize the network simultaneously. Moreover, to properly extract and leverage channel dependencies and spatial attention while disentangling the affective gap, we carefully designed two attention modules: the Multi-head Cross Channel Attention module and the Level-dependent Class Activation Map module. Finally, the proposed deep framework obtains new state-of-the-art performance on six VEA benchmarks, where it outperforms existing state-of-the-art methods by a large margin, e.g., +3.85% on the WEBEmo dataset at 25 classes classification accuracy.

cs.CV↗

Towards thinner convolutional neural networks through Gradually Global Pruning

Deep network pruning is an effective method to reduce the storage and computation cost of deep neural networks when applying them to resource-limited devices. Among many pruning granularities, neuron level pruning will remove redundant neurons and filters in the model and result in thinner networks. In this paper, we propose a gradually global pruning scheme for neuron level pruning. In each pruning step, a small percent of neurons were selected and dropped across all layers in the model. We also propose a simple method to eliminate the biases in evaluating the importance of neurons to make the scheme feasible. Compared with layer-wise pruning scheme, our scheme avoid the difficulty in determining the redundancy in each layer and is more effective for deep networks. Our scheme would automatically find a thinner sub-network in original network under a given performance.

cs.CV↗

Attribute-controlled face photo synthesis from simple line drawing

Face photo synthesis from simple line drawing is a one-to-many task as simple line drawing merely contains the contour of human face. Previous exemplar-based methods are over-dependent on the datasets and are hard to generalize to complicated natural scenes. Recently, several works utilize deep neural networks to increase the generalization, but they are still limited in the controllability of the users. In this paper, we propose a deep generative model to synthesize face photo from simple line drawing controlled by face attributes such as hair color and complexion. In order to maximize the controllability of face attributes, an attribute-disentangled variational auto-encoder (AD-VAE) is firstly introduced to learn latent representations disentangled with respect to specified attributes. Then we conduct photo synthesis from simple line drawing based on AD-VAE. Experiments show that our model can well disentangle the variations of attributes from other variations of face photos and synthesize detailed photorealistic face images with desired attributes. Regarding background and illumination as the style and human face as the content, we can also synthesize face photos with the target style of a style photo.

cs.CV↗

Every Filter Extracts A Specific Texture In Convolutional Neural Networks

Many works have concentrated on visualizing and understanding the inner mechanism of convolutional neural networks (CNNs) by generating images that activate some specific neurons, which is called deep visualization. However, it is still unclear what the filters extract from images intuitively. In this paper, we propose a modified code inversion algorithm, called feature map inversion, to understand the function of filter of interest in CNNs. We reveal that every filter extracts a specific texture. The texture from higher layer contains more colours and more intricate structures. We also demonstrate that style of images could be a combination of these texture primitives. Two methods are proposed to reallocate energy distribution of feature maps randomly and purposefully. Then, we inverse the modified code and generate images of diverse styles. With these results, we provide an explanation about why Gram matrix of feature maps \cite{Gatys_2016_CVPR} could represent image style.

cs.CV↗