SearcharxivSearch

arXiv subjects

Jiajun He

Publications and source records attributed to Jiajun He.

At least 19 recordsLinked to original sources

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

cs.CL

Tri-Hybrid Beamforming Design for DMA-Aided Secure ISAC Systems

This paper proposes a tri-hybrid beamforming scheme for secure integrated sensing and communication (ISAC) with a dynamic metasurface antenna (DMA) architecture, where the base station (BS) is capable of communicating with legitimate users and sensing the target. There is also an eavesdropper in the system intending to eavesdrop on the confidential information. The tri-hybrid beamforming design problem is formulated with the objective of maximizing the sensing signal-to-noise ratio (SNR) under the constraints of secrecy spectral efficiency (SSE), transmit power, and physical structure limitations. We first solve the problem and obtain the optimized fully-digital beamforming solution through successive convex approximation (SCA) and semidefinite relaxation (SDR) approaches. A triple alternating optimization scheme is then developed to iteratively optimize the digital, analog, and DMA beamformers, progressively approximating the fully-digital solution. Numerical results demonstrate that the proposed secure tri-hybrid beamforming design for DMA-aided ISAC improves the sensing SNR by approximately 3 dB compared to a tri-hybrid beamforming scheme with a fixed DMA electromagnetic design.

eess.SP

Positioning with Flexible Reflectors: Solution and Performance Analysis

Flexible reflectors (FRs) have emerged as a low-cost and energy-efficient solution for reshaping electromagnetic propagation environments across a wide range of applications. This paper investigates FR-swarm-assisted target localization in scenarios where line-of-sight (LoS) paths are unavailable. By leveraging the virtual LoS paths created by the FRs, a simple yet accurate estimator is proposed for localization under severe blockage conditions. To characterize the performance limits of the proposed scheme, we derive the Cramer-Rao lower bound (CRLB) and use it to optimize the positions and orientations of the FRs. Furthermore, by accounting for random FR deployment, we characterize the CRLB distribution and reveal how different network configurations affect localization accuracy. Simulation results demonstrate that the developed scheme closely approaches the CRLB performance, while the derived analytical results provide useful guidelines for FR deployment and network design.

cs.IT

Recipes for Steering and Scaling LLMs via Sampling

Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms -- one based on Sequential Monte Carlo (SMC) and one based on Replica Exchange (RE) -- that steer generation toward powering, product or tilting of the base model distribution. We illustrate this framework through scaling the generation quality of LLMs without external supervision or reward models. Experimental results demonstrate our methods scale more favorably than Best-of-N and standard MCMC baselines. Overall, this paper offers a systematic recipe for probabilistic inference with LLMs via sampling.

cs.CL

Towards Diverse Scientific Hypothesis Search with Large Language Models

Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses. Yet in many discovery settings, the goal is not to identify a single best hypothesis since validation can be noisy and expensive, and scientists benefit from a set of high-quality alternative hypotheses that hedge against downstream uncertainty for the best solutions. Nevertheless, commonly used evolutionary search recipes tend to prioritize optimization over exploration in hypothesis generation, and the resulting selection pressure during the search process leads to diversity collapse. Motivated by these limitations, we formulate hypothesis search as a sampling problem, where the objective is to efficiently produce diverse, high-quality hypotheses under a fixed validation budget. Building on this perspective, we propose \ours, an evolutionary framework inspired by the classical parallel tempering algorithm that searches hypotheses at multiple temperature levels and enables principled information exchange across temperatures to improve exploration without disrupting convergence. Across domains including molecular discovery, equation discovery, and algorithm discovery, our approach consistently improves both hypothesis quality and diversity under the same validation budget, and produces candidates that remain robust under more expensive downstream computational validations.

cs.LG

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, limited phonetic variation, unnatural prosody, and temporal shifts, degrading naturalness and intelligibility. Although sequence-to-sequence (seq2seq) voice conversion (VC) based EL-speech-to-normal-speech conversion (EL2SP) is promising, substantial mismatches between EL and normal speech inevitably cause cumulative mapping errors that limit performance. To address this, we describe a novel representation learning framework integrating speech and text representations to improve mapping and reconstruction quality within a seq2seq VC model. Methods: our methodology comprises two main stages: 1) representation integration and learning, and 2) reconstruction training. A network capable of incorporating auxiliary text information is first constructed with pretrained modules to learn speech--text-based integrated representations. Then, an autoencoder-style reconstruction strategy finalizes EL2SP model to inherit these representations without increasing model complexity. We introduce three fusion strategies including middle-, input-, and hybrid-level fusion strategies that progressively enhance learning. Moreover, besides standard seq2seq VC objectives, an additional reconstruction loss on the integrated representation is introduced to refine representation transfer. Results: experiments under different EL2SP datasets consistently demonstrate that our methods, combined with data augmentations, outperform baselines relying solely on speech representations. Furthermore, progressive improvements with system design depth validate the effectiveness of our methods. Significance: the proposed methods provide an extensible and practical methodology for EL speech enhancement and assistive communication technologies.

eess.AS

Free energy Estimation on Any State Space

Free energy estimation is a fundamental yet challenging problem, from physics to statistics. Classical approaches rely on thermodynamic transformations, ranging from direct estimation, quasistatic integration, to finite-time averaging. Recent work [He and Du et al., 2025] learns neural transports to significantly accelerate the efficiency in the finite-time regime. In this paper, we generalize this framework to arbitrary state spaces. Building on this view, we develop a generalized neural transport learning approach for efficient estimation. Experiments validate the effectiveness and efficiency of the proposed method beyond continuous settings, extending to discrete and multimodal spaces as well as autoregressive settings. Beyond free energy estimation, we establish algebraic identities and reveal a group-theoretic structure linking infinitesimal time reversal and generalized Doob's $h$-transforms, showing that their compositions form a generalized dihedral group.

stat.ML

Multiresonant Membrane Metasurfaces for Multifunctional Fingerprint Recognition and Real-time Biochemical Tracking

Label-free identification and real-time tracking of biochemical substances became critical for molecular diagnostics and chemical analysis, yet conventional resonant terahertz metasurface sensing relies on a single resonance, limiting spectral selectivity and dynamic capability. Here, we suggest multiresonant membrane metasurfaces and implement them for simultaneous static molecular fingerprint retrieval and dynamic reaction monitoring within a single pixel. We consider a membrane metasurface supporting multiple quasi-bound states in the continuum designed at target frequencies and enabling the tailoring of the field enhancement and frequency-selective interaction with target analytes. As a proof-of-concept, we achieve label-free detection of the dual fingerprint absorption features of pefloxacin at 0.78 THz and 0.99 THz, and real-time tracking of vitamin C oxidation and denaturation under ambient conditions. The kinetic profiles extracted from the THz amplitude evolution show excellent agreement with nonlinear reaction models, demonstrating quantitative biochemical tracking capabilities. Our results establish a versatile and scalable THz photonic platform that unifies static fingerprint identification and dynamic reaction monitoring, paving the way toward integrated on-chip biochemical analytics and multifunctional metasurface sensors.

physics.optics

Rare Event Analysis via Stochastic Optimal Control

Rare events such as conformational changes in biomolecules, phase transitions, and chemical reactions are central to the behavior of many physical systems, yet they are extremely difficult to study computationally because unbiased simulations seldom produce them. Transition Path Theory (TPT) provides a rigorous statistical framework for analyzing such events: it characterizes the ensemble of reactive trajectories between two designated metastable states (reactant and product), and its central object--the committor function, which gives the probability that the system will next reach the product rather than the reactant--encodes all essential kinetic and thermodynamic information. We introduce a framework that casts committor estimation as a stochastic optimal control (SOC) problem. In this formulation the committor defines a feedback control--proportional to the gradient of its logarithm--that actively steers trajectories toward the reactive region, thereby enabling efficient sampling of reactive paths. To solve the resulting hitting-time control problem we develop two complementary objectives: a direct backpropagation loss and a principled off-policy Value Matching loss, for which we establish first-order optimality guarantees. We further address metastability, which can trap controlled trajectories in intermediate basins, by introducing an alternative sampling process that preserves the reactive current while lowering effective energy barriers. On benchmark systems, the framework yields markedly more accurate committor estimates, reaction rates, and equilibrium constants than existing methods.

stat.ML

Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models

Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations of visual signals, \textit{e.g.}, an 81-frame video, can further be hashed into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.

cs.LG

VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation

Visual content creation tasks demand a nuanced understanding of design conventions and creative workflows-capabilities challenging for general models, while workflow-based agents lack specialized knowledge for autonomous creative planning. To overcome these challenges, we propose VisionCreator, a native visual-generation agentic model that unifies Understanding, Thinking, Planning, and Creation (UTPC) capabilities within an end-to-end learnable framework. Our work introduces four key contributions: (i) VisGenData-4k and its construction methodology using metacognition-based VisionAgent to generate high-quality creation trajectories with explicit UTPC structures; (ii) The VisionCreator agentic model, optimized through Progressive Specialization Training (PST) and Virtual Reinforcement Learning (VRL) within a high-fidelity simulated environment, enabling stable and efficient acquisition of UTPC capabilities for complex creation tasks; (iii) VisGenBench, a comprehensive benchmark featuring 1.2k test samples across diverse scenarios for standardized evaluation of multi-step visual creation capabilities; (iv) Remarkably, our VisionCreator-8B/32B models demonstrate superior performance over larger closed-source models across multiple evaluation dimensions. Overall, this work provides a foundation for future research in visual-generation agentic systems.

cs.CV

Topology-Aware Integrated Communication, Sensing, and Power Transfer for SAGIN

The space-air-ground integrated network (SAGIN) has garnered significant attention in recent years due to its capability to extend communication networks from terrestrial environments to near-ground and space contexts. The application of SAGIN enables to achieve a high-quality, multi-functional, and complex communication requirements, which are essential for sixth-generation communication systems. This paper presents a topology aware (TA) framework to leverage the topological structure in SAGIN to address the multi-functional communication challenge, particularly the integrated sensing, communication, and power transfer (ISCPT) problem. To take advantage of the topological structure, we initially establish the topology according to the criteria of visibility and channel strength. The ISCPT problem can be reformulated into a topological structure as a mixed integer linear program, providing valuable insights from the objectives and constraints. Results demonstrate the superior performance of our solution compared to the benchmarks.

cs.IT

Topology-Aware Coordination for Multi-Functional Low-Altitude Wireless Networks

Low-altitude wireless networks (LAWNs) are expected to consist of multi-tier, heterogeneous terrestrial and non-terrestrial devices, where effective coordination is essential to fully unlock the complementary capabilities of diverse systems from different vendors. To address this issue, we propose a novel multi-functional coordination framework that enables seamless cooperation within the LAWN while supporting efficient execution of diverse network functions. In the proposed architecture, each device or infrastructure element is assigned to a specific functional role, namely, edge mobile terminal (E-MT), distributed MT (D-MT), or computing center. E-MTs are equipped with lightweight, independent signal processing and computing capabilities, while D-MTs and the computing center handle regional and global coordination, respectively. To enhance the overall network efficiency, we model the LAWN as a sparse graph, where nodes represent network nodes and edges are defined according to a set of controllable connection rules. This topology-aware (TA) representation allows for efficiently solving various coordination tasks across the network. Numerical results show that the proposed TA coordination framework outperforms baseline approaches that lack topological insights, achieving higher efficiency in multi-task coordination. Finally, we discuss key technical challenges and outline potential solutions for future deployment.

cs.IT

Delay Alignment Modulation for Secure ISAC Systems

This paper introduces delay-alignment modulation (DAM) for secure integrated sensing and communication (ISAC). Due to the broadcast nature of multi-user downlinks, communications are vulnerable to eavesdropping. DAM applies controlled per-path symbol delays at the transmitter to coherently align the multipath components at the intended user, enhancing the received signal power, while simultaneously creating delay misalignment at the eavesdropper (Eve). To mitigate sensing degradation caused by multipath propagation, we propose a two-stage protocol that first estimates the angle and then the delay of the line-of-sight (LoS) path after suppressing multipath interference. We derive the secrecy spectral efficiency (SSE) and the Cramer-Rao (CRB) of the target delay. Finally, we develop a path-based zero-forcing (ZF) precoding framework and formulate a max-min SSE design under CRB and power constraints. Simulation results show DAM significantly outperforms the strongest-path (SP) benchmark in terms of SSE, while meeting sensing requirements, since intentional delay alignment at legitimate users degrades reception at Eve.

cs.IT

Assessing generative modeling approaches for free energy estimates in condensed matter

The accurate estimation of free energy differences between two states is a long-standing challenge in molecular simulations. Traditional approaches generally rely on sampling multiple intermediate states to ensure sufficient overlap in phase space and are, consequently, computationally expensive. Boltzmann Generators and related generative-model-based methods have recently addressed this challenge by learning a direct probability density transform between two states. However, it remains unclear which approach provides the best trade-off between efficiency, accuracy, and scalability. In this work, we review and benchmark selected generative approaches for condensed-matter systems, including discrete and continuous normalizing flows for targeted free energy perturbation and FEAT (Free Energy Estimators with Adaptive Transport) combined with the escorted Jarzynski equality, using coarse-grained monatomic ice and Lennard-Jones solids as benchmark systems. All models yield highly accurate free energy estimates and, depending on the system, may require fewer energy evaluations than traditional methods. Continuous flows and FEAT are most efficient in energy evaluations, whereas discrete flows have substantially lower inference cost. By releasing all data together with our results, we enable future benchmarking of free energy estimation methods in condensed-phase systems.

cond-mat.stat-mech

Evaluating Large Language Models in Scientific Discovery

Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.

cs.AI

Training-Free Vector Quantization via Gaussian VAEs

Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images into discrete tokens. However, they are difficult to train due to discretization. In this paper, we propose a simple yet effective technique dubbed Gaussian Quant (GQ), which first trains a Gaussian VAE under certain constraints and then converts it into a VQ-VAE without additional training. For conversion, GQ generates random Gaussian noise as a codebook and finds the closest noise vector to the posterior mean. Theoretically, we prove that when the logarithm of the codebook size exceeds the bits-back coding rate of the Gaussian VAE, a small quantization error is guaranteed. Practically, we propose a heuristic to train Gaussian VAEs for effective conversion, named the target divergence constraint (TDC). Empirically, we show that GQ outperforms previous VQ-VAEs, such as VQGAN, FSQ, LFQ, and BSQ, on both UNet and ViT architectures. Furthermore, TDC also improves previous Gaussian VAE discretization methods, such as TokenBridge. The source code is provided in https://github.com/tongdaxu/VQ-VAE-from-Gaussian-VAE.

cs.LG

Pinching-Antenna System-Assisted Localization: A Stochastic Geometry Perspective

This paper proposes a novel localization framework underpinned by a pinching-antenna (PA) system, in which the target location is estimated using received signal strength (RSS) measurements obtained from downlink signals transmitted by the PAs. To develop a comprehensive analytical framework, we employ stochastic geometry to model the spatial distribution of the PAs, enabling tractable and insightful network-level performance analysis. Closed-form expressions for target localizability and the Cramer-Rao lower bound (CRLB) distribution are analytically derived, enabling the evaluation of the fundamental limits of PA-assisted localization systems without extensive simulations. Furthermore, the proposed framework provides practical guidance for selecting the optimal waveguide number to maximize localization performance. Numerical results also highlight the superiority of the PA-assisted approach over conventional fixed-antenna systems in terms of the CRLB.

eess.SP