SearcharxivSearch

arXiv subjects

Zijian Lin

Publications and source records attributed to Zijian Lin.

18 recordsLinked to original sources

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments. Project videos and supplementary materials are available at: https://sites.google.com/view/vista-policy.

cs.RO

BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs

Pretrained neural speech codecs typically use a fixed residual vector quantization (RVQ) depth for all frames, ignoring temporal variation in quantization difficulty. We propose BAMU, a bitstream-aware dynamic RVQ allocation framework for frozen pretrained codecs. A lightweight, rate-independent predictor estimates frame- and layer-wise marginal latent-distortion reductions, while a constrained allocator selects prefix-valid depths under an exact serialized-size budget. Experiments on EnCodec and DAC over LibriSpeech, together with VCTK evaluation, show consistent EnCodec gains and DAC improvements mainly at medium and high rates. A 30-listener study confirms a MOS improvement from 3.449 to 3.780 over matched fixed-depth coding.

eess.AS

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, traceability, and structured assurance arguments. The deployment layer maintains claim validity through runtime monitoring, authority management, intervention, incident response, and controlled updates. Because assumptions and failures propagate across these layers, neither model capability, isolated safeguards, nor benchmark performance alone can establish end-to-end trustworthiness. Drawing on embodied AI, robotics, control, dependable computing, distributed systems, and autonomous driving, we further propose a non-normative hierarchy of trustworthiness levels. This hierarchy grades the strength of bounded deployment claims across task capability, safety, system assurance, operational governance, and supporting evidence, providing a basis for bounded deployment, comparative evaluation, research prioritization, and future standardization.

cs.RO

Qwen-Music Technical Report

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.

cs.SD

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.

cs.RO

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

Conventional neural speech codecs suffer from severe intelligibility degradation at ultra-low bitrates, where the bottleneck transitions from acoustic distortion to semantic loss. To address this issue, this paper conducts a systematic investigation into the role and fundamental limits of integrating frozen semantic priors -- specifically HuBERT and Whisper -- into neural speech coding. We introduce and quantitatively validate a novel Semantic Retirement phenomenon: while semantic constraints reduce the Word Error Rate (WER) by up to ~10% relatively at 1.5 kbps, their benefits rapidly diminish beyond 6 kbps, indicating a practical capacity boundary. We further uncover a clear trade-off between different prior types: acoustic-rich priors (HuBERT) better preserve prosodic and timbral details, whereas high-level linguistic priors (Whisper) effectively suppress phonetic hallucinations in noisy environments (reducing hallucination rates by 26 percent) and substantially narrow the generalization gap for unseen speakers. Building on these findings, we propose a bitrate-aware regulation strategy that dynamically adjusts prior strength to optimize the trade-off between semantic consistency and perceptual naturalness. Extensive experimental evaluations confirm that our approach achieves competitive intelligibility and noise robustness compared to existing baselines, offering a principled pathway toward ultra-low-bitrate generative speech coding.

eess.AS

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To address this, we propose TTS-PRISM, a multi-dimensional diagnostic framework for Mandarin. First, we establish a 12-dimensional schema spanning stability to advanced expressiveness. Second, we design a targeted synthesis pipeline with adversarial perturbations and expert anchors to build a high-quality diagnostic dataset. Third, schema-driven instruction tuning embeds explicit scoring criteria and reasoning into an efficient end-to-end model. Experiments on a 1,600-sample Gold Test Set show TTS-PRISM outperforms generalist models in human alignment. Profiling six TTS paradigms establishes intuitive diagnostic flags that reveal fine-grained capability differences. TTS-PRISM is open-source, with code and checkpoints at https://github.com/xiaomi-research/tts-prism.

cs.CL

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observations. We uncover this gap via diagnostic analyses showing the bottleneck is a missing view-consistent spatial interface rather than insufficient visual features or weak reasoning. To bridge this, we introduce \textbf{3ViewSense}, a framework that grounds spatial reasoning in Orthographic Views. Drawing on engineering cognition, we propose a ``Simulate-and-Reason'' mechanism that decomposes complex scenes into canonical orthographic projections to resolve geometric ambiguities. By aligning egocentric perceptions with these allocentric references, our method facilitates explicit mental rotation and reconstruction. Empirical results on spatial reasoning benchmarks demonstrate that our method significantly outperforms existing baselines, with consistent gains on occlusion-heavy counting and view-consistent spatial reasoning. The framework also improves the stability and consistency of spatial descriptions, offering a scalable path toward stronger spatial intelligence in multimodal systems.~\footnote{https://github.com/Jasaxion/3ViewSense}

cs.CV

A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional intelligence. Existing methods for NVs acquisition are either costly and unscalable (relying on manual annotation/recording) or unnatural (relying on rule-based synthesis). To address these limitations, we propose a highly scalable automatic annotation framework to label non-verbal phenomena from natural speech, which is low-cost, easily extendable, and inherently diverse and natural. This framework leverages a unified detection model to accurately identify NVs in natural speech and integrates them with transcripts via temporal-semantic alignment method. Using this framework, we created and released \textbf{NonVerbalSpeech-38K}, a diverse, real-world dataset featuring 38,718 samples across 10 NV categories collected from in-the-wild media. Experimental results demonstrate that our dataset provides superior controllability for NVs generation and achieves comparable performance for NVs understanding.

cs.SD

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their deployment in scenarios where inference speed is critical. In this work, we propose Speech Speculative Decoding (SSD), a novel framework for autoregressive speech synthesis acceleration. Specifically, our method employs a lightweight draft model to generate candidate token sequences, which are subsequently verified in parallel by the target model using the proposed SSD framework. Experimental results demonstrate that SSD achieves a significant speedup of 1.4x compared with conventional autoregressive decoding, while maintaining high fidelity and naturalness. Subjective evaluations further validate the effectiveness of SSD in preserving the perceptual quality of the target model while accelerating inference.

cs.SD

Interlayer Hopping between Surface Mott Insulator and Bulk Band Insulator in layered 1T-TaS_{2}

In condensed matter physics, various mechanisms give rise to distinct insulating phases. The competition and interplay between these phases remain elusive, even for the seemingly most distinguishable band and Mott insulators. In multilayer systems, such interplay is mediated by interlayer hopping, which competes with the Coulomb repulsion to determine the nature of insulators. The layered compound 1T-TaS_{2} provides an ideal platform for investigating this phenomenon, as it naturally hosts coexisting Mott and band insulating states. However, distinguishing these distinct insulating states and characterizing the evolution remain challenging. In this study, we employ a dual approach utilizing surface-sensitive High-Resolution Electron Energy Loss Spectroscopy (HREELS) and bulk-sensitive Fourier-transform Infrared Spectroscopy (FTIR) to investigate the electronic excitation spectrum of 1T-TaS_{2}. Our methodology effectively identifies the features originating from the Mott and band insulators by analyzing the differences in their bulk and surface spectral weights, along with their energy distinctions. Based on the previous identification, we further investigate the evolution of insulating state features in the homostructure as they are modulated by temperature. The measurements and Dynamical Mean-Field Theory (DMFT) calculations suggest that the softening and broadening of Hubbard excitations in the Mott state with increasing temperature result from enhanced interlayer hopping between the Mott and band insulators.

cond-mat.mtrl-sci

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

This paper describes the zero-shot spontaneous style TTS system for the ISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose a LLaMA-based codec language model with a delay pattern to achieve spontaneous style voice cloning. To improve speech intelligibility, we introduce the Classifier-Free Guidance (CFG) strategy in the language model to strengthen conditional guidance on token prediction. To generate high-quality utterances, we adopt effective data preprocessing operations and fine-tune our model with selected high-quality spontaneous speech data. The official evaluations in the CoVoC constrained track show that our system achieves the best speech naturalness MOS of 3.80 and obtains considerable speech quality and speaker similarity results.

cs.SD

Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue Systems

Efficient knowledge retrieval plays a pivotal role in ensuring the success of end-to-end task-oriented dialogue systems by facilitating the selection of relevant information necessary to fulfill user requests. However, current approaches generally integrate knowledge retrieval and response generation, which poses scalability challenges when dealing with extensive knowledge bases. Taking inspiration from open-domain question answering, we propose a retriever-generator architecture that harnesses a retriever to retrieve pertinent knowledge and a generator to generate system responses.~Due to the lack of retriever training labels, we propose relying on feedback from the generator as pseudo-labels to train the retriever. To achieve this, we introduce a dual-feedback mechanism that generates both positive and negative feedback based on the output of the generator. Our method demonstrates superior performance in task-oriented dialogue tasks, as evidenced by experimental results on three benchmark datasets.

cs.CL

Temperature-Dependent Collective Excitations in a Three-Dimensional Dirac System ZrTe$_{5}$

Zirconium pentatelluride (ZrTe$_{5}$), a system with a Dirac linear band across the Fermi level and anomalous transport features, has attracted considerable research interest for it is predicted to be located at the boundary between strong and weak topological insulators separated by a topological semimetal phase. However, the experimental verification of the topological phase transition and the topological ground state in ZrTe$_{5}$ is full of controversies, mostly due to the difficulty of precisely capturing the small gap evolution with single-particle band structure measurements. Alternatively, the collective excitations of electric charges, known as plasmons, in Dirac systems exhibiting unique behavior, can well reflect the topological nature of the band structure. Here, using reflective high-resolution electron energy loss spectroscopy (HREELS), we investigate the temperature-dependent collective excitations of ZrTe$_{5}$, and discover that the plasmon energy in ZrTe$_{5}$ is proportional to the $1/3$ power of the carrier density $n$, which is a unique feature of plasmons in three-dimensional Dirac systems. Based on this conclusion, the origin of the resistivity anomaly of ZrTe$_{5}$ can be attributed to the temperature-dependent chemical potential shift in extrinsic Dirac semimetals.

cond-mat.str-el

A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance

Domain generalization studies the problem of training a model with samples from several domains (or distributions) and then testing the model with samples from a new, unseen domain. In this paper, we propose a novel approach for domain generalization that leverages recent advances in large vision-language models, specifically a CLIP teacher model, to train a smaller model that generalizes to unseen domains. The key technical contribution is a new type of regularization that requires the student's learned image representations to be close to the teacher's learned text representations obtained from encoding the corresponding text descriptions of images. We introduce two designs of the loss function, absolute and relative distance, which provide specific guidance on how the training process of the student model should be regularized. We evaluate our proposed method, dubbed RISE (Regularized Invariance with Semantic Embeddings), on various benchmark datasets and show that it outperforms several state-of-the-art domain generalization methods. To our knowledge, our work is the first to leverage knowledge distillation using a large vision-language model for domain generalization. By incorporating text-based information, RISE improves the generalization capability of machine learning models.

cs.CV

Dramatic Plasmon Response to the Charge-Density-Wave Gap Development in $1\textit{T}-{\mathrm{TiSe}}_{2}$

1T-TiSe2 is one of the most studied charge density wave (CDW) systems, not only because of its peculiar properties related to the CDW transition, but also due to its status as a promising candidate of exciton insulator signaled by the proposed plasmon softening at the CDW wave vector. Using high-resolution electron energy loss spectroscopy, we report a systematic study of the temperature-dependent plasmon behaviors of 1T-TiSe2. We unambiguously resolve the plasmon from phonon modes, revealing the existence of Landau damping to the plasmon at finite momentums, which does not support the plasmon softening picture for exciton condensation. Moreover, we discover that the plasmon lifetime at zero momentum responds dramatically to the bandgap evolution associated with the CDW transition. The interband transitions near the Fermi energy in the normal phase is demonstrated serving as a strong damping channel of plasmons, while such a channel in the CDW phase is suppressed due to the CDW gap opening, which results in the dramatic tunability of the plasmon in semimetals or small-gap semiconductors.

cond-mat.str-el

Geometric Effect of High-Resolution Electron Energy Loss Spectroscopy on the Identification of Plasmons: An Example of Graphene

High-resolution electron energy loss spectroscopy (HREELS) is one of the most powerful methods to detect the dispersion of plasmons. However, we find that in the HREELS measurement, the scattering geometric configuration will seriously affect the identification of plasmons. Here, taking graphene as an example, using the HREELS capable of two-dimensional energy-momentum mapping combined with the intensity distribution calculations, we visually display the intensity distribution of the scattering geometric factor. We demonstrate that the energy loss peaks from the scattering geometric effect may be misinterpreted as the features of an acoustic plasmon. In any HREELS measurement, it is necessary to evaluate the effect of the scattering geometry quantitatively to identify the intrinsic surface excitations.

cond-mat.mes-hall

Real-Space Investigation of the Charge Density Wave in VTe2 Monolayer with Rotational and Mirror Symmetries Broken

Recently the charge density wave (CDW) in vanadium dichalcogenides have attracted increasing research interests, but a real-space investigation on the symmetry breaking of the CDW state in VTe2 monolayer is still lacking. We have investigated the CDW of VTe2 monolayer by low energy electron diffraction (LEED) and scanning tunneling microscope (STM). While the LEED experiments revealed a (4X4) CDW transition at 192+-2 K, our low-temperature STM experiments resolved the (4X4) lattice distortions and charge-density modulation in real space, and further unveiled a 1D modulation that breaks the three-fold rotational and mirror symmetries in the CDW state. In accordance with the CDW state at low temperature, a CDW gap of 12 meV was detected by scanning tunneling spectroscopy (STS) at 4.9 K. Our work provides real-space evidence on the symmetry breaking of the (4X4) CDW state in VTe2 monolayer, and implies there is a certain mechanism, beyond the conventional Fermi surface nesting or the q-dependent electron-phonon coupling, is responsible for the formation of CDW state in VTe2 monolayer.

cond-mat.mtrl-sci