SearcharxivSearch

arXiv subjects

Chenchen Yang

Publications and source records attributed to Chenchen Yang.

16 recordsLinked to original sources

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data

Vision-language models (VLMs) are powerful general-purpose reasoners, yet converting them into robot control policies (VLAs) is surprisingly difficult. The root cause is a two-fold gap: VLMs are trained on internet-scale images with language-understanding objectives, while VLAs must perceive robot scenes and predict motor actions. Fine-tuning a VLM directly on robot action data forces the model to cross both gaps at once -- the learning curve is steep and the rich generalizations learned during pretraining tend to degrade rather than transfer. We argue that this gap can be bridged gradually with the right intermediate data. We introduce \emph{embodied trajectory-coupled (ETC) data} -- vision-language supervision derived from the same robot scenes and trajectories used for action learning. Because ETC data shares the visual context of robot operation while retaining familiar language-understanding objectives, it provides a natural stepping stone between VLM pretraining and VLA fine-tuning. Building on this, we design a three-stage training recipe. Distribution Bridging first adapts the VLM to embodied visual-language semantics. Objective Bridging then gradually shifts the model toward action prediction while preserving the acquired representations. Retentive Adaptation finally specializes the policy to the target deployment domain. We further show that mixing task-relevant out-of-distribution ETC data with a small amount of action data enables the model to generalize to novel visual-language conditions without requiring additional robot demonstrations. Simulation and real-robot experiments confirm that this gradual bridging strategy is the key to transferring VLM generalization into robust, deployable robot policies.

cs.RO

World Action Models: The Next Frontier in Embodied AI

Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.

cs.RO

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational assistants, making it a significant task for modern Text-to-Speech models. However, existing models are largely trained on carefully recorded studio data, which produces speech that is clean and well-articulated, yet lacks the lived-in qualities of real human voices. To address these limitations, we present MOSS-VoiceGenerator, an open-source instruction-driven voice generation model that creates new timbres directly from natural language prompts. Motivated by the hypothesis that exposure to real-world acoustic variation produces more perceptually natural voices, we train on large-scale expressive speech data sourced from cinematic content. Subjective preference studies demonstrate its superiority in overall performance, instruction-following, and naturalness compared to other voice design models.

cs.SD

MOVA: Towards Scalable and Synchronized Video-Audio Generation

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique challenges in architecture, data, and training. Moreover, the closed-source nature of existing systems limits progress in the field. In this work, we introduce MOVA (MOSS Video and Audio), an open-source model capable of generating high-quality, synchronized audio-visual content, including realistic lip-synced speech, environment-aware sound effects, and content-aligned music. MOVA employs a Mixture-of-Experts (MoE) architecture, with a total of 32B parameters, of which 18B are active during inference. It supports IT2VA (Image-Text to Video-Audio) generation task. By releasing the model weights and code, we aim to advance research and foster a vibrant community of creators. The released codebase features comprehensive support for efficient inference, LoRA fine-tuning, and prompt enhancement.

cs.CV

WESR: Scaling and Evaluating Word-level Event-Speech Recognition

Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet under-explored challenge. Current methods suffer from insufficient task definitions with limited category coverage and ambiguous temporal granularity. They also lack standardized evaluation frameworks, hindering the development of downstream applications. To bridge this gap, we first develop a refined taxonomy of 21 vocal events, with a new categorization into discrete (standalone) versus continuous (mixed with speech) types. Based on the refined taxonomy, we introduce WESR-Bench, an expert-annotated evaluation set (900+ utterances) with a novel position-aware protocol that disentangles ASR errors from event detection, enabling precise localization measurement for both discrete and continuous events. We also build a strong baseline by constructing a 1,700+ hour corpus, and train specialized models, surpassing both open-source audio-language models and commercial APIs while preserving ASR quality. We anticipate that WESR will serve as a foundational resource for future research in modeling rich, real-world auditory scenes.

cs.CL

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.

cs.CL

Ultraviolet and Near-Infrared Dual Band Selective-Harvesting Transparent Luminescent Solar Concentrators

Visibly transparent luminescent solar concentrators (TLSCs) can optimize both power production and visible transparency by selectively harvesting the invisible portion of the solar spectrum. Since the primary applications of TLSCs include building envelopes, greenhouses, automobiles, signage, and mobile electronics, maintaining aesthetics and functionalities is as important as achieving high power conversion efficiencies (PCEs) in practical deployment. In this work, we combine massive-downshifting phosphorescent nanoclusters and fluorescent organic molecules into a TLSC system as ultraviolet (UV) and near-infrared (NIR) selective-harvesting luminophores, respectively, demonstrating UV and NIR dual-band selective-harvesting TLSCs with PCE over 3%, average visible transmittance (AVT) exceeding 75% and color metrics suitable for the window industry. With distinct wavelength-selectivity and effective utilization of the invisible portion of the solar spectrum, this work reports the highest light utilization efficiency (PCE x AVT) of 2.6 for a TLSC system, the highest PCE of any transparent photovoltaic device with AVT greater than 70%, and outperforms the practical limit for non-wavelength-selective transparent photovoltaics.

physics.app-ph

Room-Temperature Processing of Inorganic Perovskite Films to Enable Flexible Solar Cells

Inorganic lead halide perovskite materials have attracted great attention recently due to their potential for greater thermal stability compared to hybrid organic perovskites. However, the high processing temperature to convert from the non-perovskite phase to cubic perovskite phase in many of these systems has limited their application in flexible optoelectronic devices. Here, we report a room temperature processed inorganic PSC based on CsPbI2Br as the light harvesting layer. By combing this composition with key precursor solvents, we show that the inorganic perovskite film can be prepared by the vacuum-assist method under room temperature conditions in air. Unencapsulated devices achieved the power conversion efficiency up to 8.67% when measured under 1-sun irradiation. Exploiting this room temperature process, flexible inorganic PSCs based on an inorganic metal halide perovskite material is demonstrated.

physics.app-ph

Optimization and Analysis of Probabilistic Caching in $N$-tier Heterogeneous Networks

In this paper, we study the probabilistic caching for an $N$-tier wireless heterogeneous network (HetNet) using stochastic geometry. A general and tractable expression of the successful delivery probability (SDP) is first derived. We then optimize the caching probabilities for maximizing the SDP in the high signal-to-noise ratio (SNR) regime. The problem is proved to be convex and solved efficiently. We next establish an interesting connection between $N$-tier HetNets and single-tier networks. Unlike the single-tier network where the optimal performance only depends on the cache size, the optimal performance of $N$-tier HetNets depends also on the BS densities. The performance upper bound is, however, determined by an equivalent single-tier network. We further show that with uniform caching probabilities regardless of content popularities, to achieve a target SDP, the BS density of a tier can be reduced by increasing the cache size of the tier when the cache size is larger than a threshold; otherwise the BS density and BS cache size can be increased simultaneously. It is also found analytically that the BS density of a tier is inverse to the BS cache size of the same tier and is linear to BS cache sizes of other tiers.

cs.IT

Opportunistic Channel Sharing in Stochastic Networks with Dynamic Traffic

In this paper, we consider the stochastic network with dynamic traffic. The spatial distribution of access points (APs) and users are first modeled as mutually independent Poisson point processes (PPPs). Different from most previous literatures which assume all the APs are fully loaded, we consider the fact that APs having no data to transmit do not generate interference to users. The APs opportunistically share the channel according to the existence of the packet to be transmitted and the proposed interference suppression strategy. In the interference suppression region, only one AP can be active at a time to transmit the packet on the channel and the other adjacent APs keep silent to reduce serious interference. The idle probability of any AP, influenced by the traffic load and availability of the channels, is analyzed. The density of simultaneously active APs in the network is obtained and the packet loss rate is further elaborated. We reveal the impacts of network features (e.g., AP density, user density and channel state) and service features (e.g., user request, packet size) on the network performance. Simulation results validate our proposed model.

cs.NI

Interference Cancellation at Receivers in Cache-Enabled Wireless Networks

In this paper, we propose to exploit the limited cache packets as side information to cancel incoming interference at the receiver side. We consider a stochastic network where the random locations of base stations and users are modeled using Poisson point processes. Caching schemes to reap both the local caching gain and the interference cancellation gain for the users are developed based on two factors: the density of different user subsets and the packets cached in the corresponding subsets. The packet loss rate (PLR) is analyzed, which depends on both the cached packets and the channel state information (CSI) available at the receiver. Theoretical results reveal the tradeoff between caching resource and wireless resource. The performance for different caching schemes are analyzed and the minimum achievable PLR for the distributed caching is derived.

cs.IT

Modeling and Analysis for Cache-Enabled Networks with Dynamic Traffic

Instead of assuming fully loaded cells in the analysis on cache-enabled networks with tools of stochastic geometry, we focus on the dynamic traffic in this letter. With modeling traffic dynamics of request arrivals and departures, probabilities of full-, free-, and modest-load cells in the large-scale cache-enabled network are elaborated based on the traffic queue state. Moreover, we propose to exploit the packets cached at cache-enabled users as side information to cancel the incoming interference. Then the packet loss rates for both the cache-enabled and cache-untenable users are investigated. The simulation results verify our analysis.

cs.IT

Modeling and Analysis for Cache-enabled Cognitive D2D Communications in Cellular Networks

Exploiting cognition to the cache-enabled device-to-device (D2D) communication underlaying the multi-channel cellular network is the main focus of this paper. D2D pairs perform direct communications via sensing the available cellular channels, bypassing the base station (BS). Dynamic service is considered and the network performance is evaluated with the stochastic geometry. Node locations are first modeled as mutually independent Poisson Point Processes, and the service queueing process is formulated. Then the corresponding tier association and cognitive access protocol are developed. The delay and the length for the queue at the BS and D2D transmitter are further elaborated, with modeling the traffic dynamics of request arrivals and departures as the discrete-time multiserver queue with priorities. Moreover, impacts of the physical layer and content-centric features on the system performance are jointly investigated to provide a valuable insight.

cs.IT

Optimal Caching Placement for D2D Assisted Wireless Caching Networks

In this paper, we devise the optimal caching placement to maximize the offloading probability for a two-tier wireless caching system, where the helpers and a part of users have caching ability. The offloading comes from the local caching, D2D sharing and the helper transmission. In particular, to maximize the offloading probability we reformulate the caching placement problem for users and helpers into a difference of convex (DC) problem which can be effectively solved by DC programming. Moreover, we analyze the two extreme cases where there is only help-tier caching network and only user-tier. Specifically, the placement problem for the helper-tier caching network is reduced to a convex problem, and can be effectively solved by the classical water-filling method. We notice that users and helpers prefer to cache popular contents under low node density and prefer to cache different contents evenly under high node density. Simulation results indicate a great performance gain of the proposed caching placement over existing approaches.

cs.IT

When ICN Meets C-RAN for HetNets: An SDN Approach

In this paper, we contribute to novelly proposing and elaborating the integration of the ICN, C-RAN and SDN for the HetNet to achieve win-win situation. The vision of the proposed system is demonstrated, followed by the advantages and challenges. We further present the hybrid system with a large-scale wireless heterogeneous campus network.

cs.NI

Analysis on Cache-enabled Wireless Heterogeneous Networks

Caching the popular multimedia content is a promising way to unleash the ultimate potential of wireless networks. In this paper, we contribute to proposing and analyzing the cache-based content delivery in a three-tier heterogeneous network (HetNet), where base stations (BSs), relays and device-to-device (D2D) pairs are included. We advocate to proactively cache the popular contents in the relays and parts of the users with caching ability when the network is off-peak. The cached contents can be reused for frequent access to offload the cellular network traffic. The node locations are first modeled as mutually independent Poisson Point Processes (PPPs) and the corresponding content access protocol is developed. The average ergodic rate and outage probability in the downlink are then analyzed theoretically. We further derive the throughput and the delay based on the \emph{multiclass processor-sharing queue} model and the continuous-time Markov process. According to the critical condition of the steady state in the HetNet, the maximum traffic load and the global throughput gain are investigated. Moreover, impacts of some key network characteristics, e.g., the heterogeneity of multimedia contents, node densities and the limited caching capacities, on the system performance are elaborated to provide a valuable insight.

cs.IT