SearcharxivSearch

arXiv subjects

Junxiang Zhang

Publications and source records attributed to Junxiang Zhang.

At least 19 recordsLinked to original sources

BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation Models

fMRI foundation models increasingly aggregate heterogeneous data across brain states, cohorts, and acquisition settings, yet pretraining domains are commonly treated as a flat mixture and downstream tasks are adapted independently. We study whether measured learning relations can organize both stages without modifying the backbone. During pretraining, a lightweight Brain-DiT proxy estimates difficulty and directed facilitation across ten fMRI domains, yielding a priority-guided cumulative domain curriculum combined with high-to-low-noise timestep scheduling and joint consolidation. During adaptation, controlled first- and higher-order transfer across fifteen tasks constructs a directed taskonomy, from which budgeted integer programming (BIP) selects directly supervised source tasks and target-specific routes. The joint priority-domain and high-to-low-timestep curriculum reduces v-NMSE, PSD-NMSE, and FC-MSE by 6.5%, 16.3%, and 10.5%, respectively, relative to uniform sampling over both dimensions, and shows strong downstream performance across six in- and out-of-domain tasks. The taskonomy reveals asymmetric, target-dependent transfer, while exploratory sealed-test evaluation shows larger descriptive gains for BIP policies when higher-order route spaces are available than for matched random controls. Together, these findings support organizing fMRI pretraining and adaptation by measured learning relations rather than treating domains and tasks as independent flat sets.

cs.CV

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

cs.LG

BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics

Whole-brain 4D fMRI generation is valuable for modeling functional brain dynamics, yet existing fMRI foundation models mainly target representation learning and downstream prediction rather than conditional predictive generation. We introduce BrainWorld, a structural-prior-conditioned generative model for whole-brain 4D fMRI dynamics. BrainWorld uses sMRI as subject-level anatomical context to guide future fMRI generation, integrating structural information into the denoising process rather than treating it as a parallel modality. Evaluated on 22 datasets spanning diverse cohorts and brain states, BrainWorld generates stable 4D fMRI trajectories up to 400 frames, improves downstream performance through generated-example augmentation, and learns transferable multimodal representations that outperform baselines. Together, these results establish BrainWorld as a condition-aware generative framework for long-horizon brain dynamics modeling and multimodal representation learning.

cs.CV

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely on inversion-based editing. While inversion-free editing is appealing as it decreases computational overhead and reconstruction errors, it remains largely unexplored for audio editing. The key challenge is to construct a source-to-target editing path through diffusion denoising dynamics. In this paper, we introduce DirectAudioEdit, the first attempt to develop a training-free and inversion-free method for audio editing. Experiments on music and event-level benchmarks across two backbones show that DirectAudioEdit reduces macro-averaged FAD and KL by 15.9% and 15.8% compared with DDPM inversion, while achieving up to 64.5% editing speedup.

cs.SD

On the Emotion Understanding of Synthesized Speech

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible reward or evaluation metric for assessing emotional expressiveness in speech synthesis. In this work, we critically examine this assumption by systematically evaluating Speech Emotion Recognition (SER) on synthesized speech across datasets, discriminative and generative SER models, and diverse synthesis models. We find that current SER models can not generalize to synthesized speech, largely because speech token prediction during synthesis induces a representation mismatch between synthesized and human speech. Moreover, generative Speech Language Models (SLMs) tend to infer emotion from textual semantics while ignoring paralinguistic cues. Overall, our findings suggest that existing SER models often exploit non-robust shortcuts rather than capturing fundamental features, and paralinguistic understanding in SLMs remains challenging.

cs.CL

Omni-fMRI: A Universal Atlas-Free fMRI Foundation Model

Self-supervised fMRI foundation models have shown promising transfer performance, yet most rely on predefined region-level parcellations that discard fine-grained voxel information and introduce atlas-dependent biases. We propose Omni-fMRI, an atlas-free foundation model that operates directly on voxel-level signals. To enable scalable pretraining on 49,497 fMRI sessions across nine datasets, Omni-fMRI introduces a dynamic patching mechanism that substantially reduces computational cost while preserving informative spatial structure. To support reproducibility and fair comparison, we establish a comprehensive benchmark suite spanning 11 datasets and a diverse set of resting-state and task-based fMRI tasks. Experimental results demonstrate that Omni-fMRI consistently outperforms existing foundation models, providing a scalable and reproducible framework for atlas-free brain representation learning. Code and logs are available.

cs.CE

SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose \texttt{SageLM}, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce \textit{SpeechFeedback}, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79\% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42\% and 26.20\%, respectively.

cs.CL

Surface Passivation for Halide Optoelectronics: Comparing Optimization and Reactivity of Amino-Silanes with Formamidinium

Amino-silane-based surface passivation schemes are gaining attention in halide perovskite optoelectronics, with varying levels of success. We compare surface treatments using (3-aminopropyl)trimethoxysilane (APTMS) and [3-(2-aminoethylamino)propyl]trimethoxysilane (AEAPTMS), applied via room-temperature vacuum deposition, to the perovskite FA0.78Cs0.22Pb(I0.85Br0.15)3 (FA = formamidinium). Both molecules improve thin-film photoluminescence properties and photovoltaic device performance, although their effectiveness depends strongly on deposition time. We show AEAPTMS has a wider, more robust processing window and yields higher performance under optimized conditions. In contrast, over-exposure, particularly with APTMS, reduces performance, with notable reductions in photoluminescence lifetime and absorbance. To probe the underlying chemistry, we employ nuclear magnetic resonance (NMR) spectroscopy and depth-resolved time-of-flight secondary ion mass spectrometry (ToF-SIMS), demonstrating that both amino-silanes react with formamidinium (FA+) cations in solution and in the solid state. This work underscores the importance of optimizing deposition conditions to balance effective passivation with potential performance loss and elucidates previously unrecognized reactive chemistry between amino-silane passivating agents and halide perovskites.

cond-mat.mtrl-sci

SDTN and TRN: Adaptive Spectral-Spatial Feature Extraction for Hyperspectral Image Classification

Hyperspectral image classification plays a pivotal role in precision agriculture, providing accurate insights into crop health monitoring, disease detection, and soil analysis. However, traditional methods struggle with high-dimensional data, spectral-spatial redundancy, and the scarcity of labeled samples, often leading to suboptimal performance. To address these challenges, we propose the Self-Adaptive Tensor- Regularized Network (SDTN), which combines tensor decomposition with regularization mechanisms to dynamically adjust tensor ranks, ensuring optimal feature representation tailored to the complexity of the data. Building upon SDTN, we propose the Tensor-Regularized Network (TRN), which integrates the features extracted by SDTN into a lightweight network capable of capturing spectral-spatial features at multiple scales. This approach not only maintains high classification accuracy but also significantly reduces computational complexity, making the framework highly suitable for real-time deployment in resource-constrained environments. Experiments on PaviaU datasets demonstrate significant improvements in accuracy and reduced model parameters compared to state-of-the-art methods.

cs.CV

Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques

In the field of deep learning, traditional attention mechanisms face significant challenges related to high computational complexity and large memory consumption when processing long sequence data. To address these limitations, we propose Opt-GPTQ, an optimized Gradient-based Post Training Quantization (GPTQ) combining the Grouped Query Attention (GQA) mechanism with paging memory management, optimizing the traditional Multi-Head Attention (MHA) mechanism by grouping query heads and sharing key-value vectors. Optimized GQA (Opt-GQA) effectively reduces computational complexity, minimizes memory fragmentation, and enhances memory utilization for large-scale models. Opt-GPTQ is optimized for Data Center Units (DCUs) and integrated into the vLLM model to maximize hardware efficiency. It customizes GPU kernels to further enhance attention computation by reducing memory access latency and boosting parallel computing capabilities. Opt-GQA integrates Attention with Linear Biases (ALiBi) to reduce overhead and enhance long-sequence processing. Experimental results show that Opt-GPTQ significantly reduces computation time and memory usage while improving model performance.

cs.DC

Observation of topological prethermal strong zero modes

Symmetry-protected topological phases cannot be described by any local order parameter and are beyond the conventional symmetry-breaking paradigm for understanding quantum matter. They are characterized by topological boundary states robust against perturbations that respect the protecting symmetry. In a clean system without disorder, these edge modes typically only occur for the ground states of systems with a bulk energy gap and would not survive at finite temperatures due to mobile thermal excitations. Here, we report the observation of a distinct type of topological edge modes, which are protected by emergent symmetries and persist even up to infinite temperature, with an array of 100 programmable superconducting qubits. In particular, through digital quantum simulation of the dynamics of a one-dimensional disorder-free "cluster" Hamiltonian, we observe robust long-lived topological edge modes over up to 30 cycles at a wide range of temperatures. By monitoring the propagation of thermal excitations, we show that despite the free mobility of these excitations, their interactions with the edge modes are substantially suppressed in the dimerized regime due to an emergent U(1)$\times$U(1) symmetry, resulting in an unusually prolonged lifetime of the topological edge modes even at infinite temperature. In addition, we exploit these topological edge modes as logical qubits and prepare a logical Bell state, which exhibits persistent coherence in the dimerized and off-resonant regime, despite the system being disorder-free and far from its ground state. Our results establish a viable digital simulation approach to experimentally exploring a variety of finite-temperature topological phases and demonstrate a potential route to construct long-lived robust boundary qubits that survive to infinite temperature in disorder-free systems.

quant-ph

Disorder-tunable entanglement at infinite temperature

Emerging quantum technologies hold the promise of unraveling difficult problems ranging from condensed matter to high energy physics, while at the same time motivating the search for unprecedented phenomena in their setting. Here we utilize a custom-built superconducting qubit ladder to realize non-thermalizing states with rich entanglement structures in the middle of the energy spectrum. Despite effectively forming an "infinite" temperature ensemble, these states robustly encode quantum information far from equilibrium, as we demonstrate by measuring the fidelity and entanglement entropy in the quench dynamics of the ladder. Our approach harnesses the recently proposed type of non-ergodic behavior known as "rainbow scar", which allows us to obtain analytically exact eigenfunctions whose ergodicity-breaking properties can be conveniently controlled by randomizing the couplings of the model, without affecting their energy. The on-demand tunability of quantum correlations via disorder allows for in situ control over ergodicity breaking and it provides a knob for designing exotic many-body states that defy thermalization.

quant-ph

Learning to Generate Pseudo Personal Mobility

The importance of personal mobility data is widely recognized in various fields. However, the utilization of real personal mobility data raises privacy concerns. Therefore, it is crucial to generate pseudo personal mobility data that accurately reflects real-world mobility patterns while safeguarding user privacy. Nevertheless, existing methods for generating pseudo mobility data, such as mechanism-based and deep-learning-based approaches, have limitations in capturing sufficient individual heterogeneity. To address these gaps, taking pseudo-person(avatar) as ground-zero, a novel individual-based human mobility generator called GeoAvatar has been proposed - which considers individual heterogeneity in spatial and temporal decision-making, incorporates demographic characteristics, and provides interpretability. Our method utilizes a deep generative model to simulate heterogeneous individual life patterns, a reliable labeler for inferring individual demographic characteristics, and a Bayesian approach for generating spatial choices. Through our method, we have achieved the generation of heterogeneous individual human mobility data without accessing individual-level personal information, with good quality - we evaluated the proposed method based on physical features, activity patterns, and spatial-temporal characteristics, demonstrating its good performance, compared to mechanism-based modeling and black-box deep learning approaches. Furthermore, this method maintains extensibility for broader applications, making it a promising paradigm for generating human mobility data.

cs.CY

Enhanced opposite Imbert-Fedorov shifts of vortex beams for precise sensing of temperature and thickness

Imbert-Fedorov (IF) shift, which refers to a tiny transverse splitting induced by spin-orbit interaction at a reflection/refraction interface, is sensitive to the refractive index of a medium and momentum state of incident light. Most of studies have focused on the shift for an incident light beam with a spin angular momentum (SAM) i.e., circular polarization. Compared to SAM, orbital angular momentum (OAM) has infinite dimensions in theory as a new degree of freedom of light and plays an important role in light-matter coupling. We demonstrate experimentally that the relative IF shifts of vortex beams with large opposite OAMs are highly enhanced in resonant structures when light refracts through a double-prism structure (DPS), in which the thickness and temperature of the air gap are precisely sensed via the observed relative IF shifts. The thickness and temperature sensitivities increase as the absolute value of opposite OAMs increases. Our results offer a technological and practical platform for applications in sensing of thickness and temperature, ingredients of environment gas, spatial displacement, chemical substances and deformation structure.

physics.optics

Quantum PT-Phase Diagram in a Non-Hermitian Photonic Structure

Photonic structures have an inherent advantage to realize PT-phase transition through modulating the refractive index or gain-loss. However, quantum PT properties of these photonic systems have not been comprehensively studied yet. Here, in a bi-photonic structure with loss and gain simultaneously existing, we analytically obtained the quantum PT-phase diagram under the steady state condition. To characterize the PT-symmetry or -broken phase, we define an Hermitian exchange operator expressing the exchange between quadrature variables of two modes. If inputting several-photon Fock states into a PT-broken bi-waveguide splitting system, most photons will concentrate in the dominant waveguide with some state distributions. Quantum PT-phase diagram paves the way to the quantum state engineering, quantum interferences, and logic operations in non-Hermitian photonic systems.

physics.optics

Generation of Arbitrarily Multiple Entangled Fields via Mechanical Oscillator Displacement

We present a convenient and efficient scheme to generate arbitrarily multipartite continuous-variable entanglement via mechanical oscillator displacement induced by two strong input pump fields in the conventional single-cavity optomechanical system. It is shown that multipartite entanglement among the outputs of the two pump fields and any number of relatively weak probe fields can be realized and optimized when the two pump fields with suitable amplitude ratio and relative initial phase are, respectively, tuned to the red and blue mechanical sidebands of a single cavity mode and each probe field is red-detuned by the mechanical frequency with respect to a different neighboring cavity mode. This method can, in principle, be extended to flexibly and conveniently generate arbitrarily multiple nondegenerate bright entangled fields by using only coherent laser fields, and may find promising applications in realistic quantum communication and networks.

quant-ph

Controllable spin-Hall and related effects of light in an atomic medium via coupling fields

We show the existence of spin-Hall effect of light (SHEL) in an atomic medium which is made anisotropic via electromagnetically induced transparency. The medium is made birefringent by applying an additional linearly polarized coupling light beam. The refractive index and the orientation of the optics axis are controlled by the coupling beam. We show that after transmitting the atomic medium, a linearly polarized probe light beam splits into its two spin components by opposite transverse shifts. With proper choice of parameters and atomic density of about $2.5\times10^{17}\,\mathrm{m}^{-3}$, the shifts are about the order of wavelength and can be larger than the wavelength by increasing the atomic density. We propose a novel measurement scheme based on a balanced homodyne detection (BHD). By properly choosing the polarization, phase, and transverse mode of the local oscillator of the BHD, one can independently measure (i) the SHEL shifts of the two spin components; (ii) the spatial and angular shifts; (iii) the transverse and longitudinal shifts. The measurement can reach the quantum limit of precision by detecting signals at the modulation frequency of the electro-optic modulator used to modulate the input probe beam. The precision is estimated to be at the nanometer level limited by the quantum noise.

physics.optics

On vortex strength and beam propagation factor of fractional vortex beams

Fractional vortex beams (FVBs) with non-integer topological charges attract much attention due to unique features of propagations, but there still exist different viewpoints on the change of their total vortex strength. Here we have experimentally demonstrated the distribution and number of vortices contained in FVBs at Fraunhofer diffraction region. We have verified that the jumps of total vortex strength for FVBs happens only when non-integer topological charge is before and after (but very close to) any even integer number, which originates from two different mechanisms for generation and movement of vortices on focal plane. Meanwhile, we have also measured the beam propagation factor (BPF) of such FVBs, and have found that their BPF values almost increase linearly in one component and oscillate increasingly in another component. Our experimental results are in good agreement with numerical results.

physics.optics