Searcharxiv⌕ Search

arXiv subjects

Yichen Zhang

Publications and source records attributed to Yichen Zhang.

At least 55 records · Page 3Linked to original sources

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

A recent cutting-edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes and visual representations, making it non-trivial to jointly optimize within a shared feature space. In this work, we present Cheers, a unified multimodal model that decouples patch-level details from semantic representations, thereby stabilizing semantics for multimodal understanding and improving fidelity for image generation via gated detail residuals. Cheers includes three key components: (i) a unified vision tokenizer that encodes and compresses image latent states into semantic tokens for efficient LLM conditioning, (ii) an LLM-based Transformer that unifies autoregressive decoding for text generation and diffusion decoding for image generation, and (iii) a cascaded flow matching head that decodes visual semantics first and then injects semantically gated detail residuals from the vision tokenizer to refine high-frequency content. Experiments on popular benchmarks demonstrate that Cheers matches or surpasses advanced UMMs in both visual understanding and generation. Cheers also achieves 4x token compression, enabling more efficient high-resolution image encoding and generation. Notably, Cheers outperforms the Tar-1.5B on the popular benchmarks GenEval and MMBench, while requiring only 20% of the training cost, indicating effective and efficient (i.e., 4x token compression) unified multimodal modeling. We will release all code and data for future research.

cs.CV↗

Disk Wind Feedback from High-mass Protostars. V. Application of Multi-Modal Machine Learning to Characterize Outflow Properties

Characterizing protostellar outflows is fundamental to understanding star formation feedback, yet traditional methods are often hindered by projection effects and complex morphologies. We present a multi-modal deep learning framework that jointly leverages spatial and spectral information from CO observations to infer protostellar mass, inclination, and position angle ($PA$). Our model, trained on synthetic ALMA observations generated from 3D magnetohydrodynamic simulations, utilizes a cross-attention fusion mechanism to integrate morphological and kinematic features with probabilistic uncertainty estimation. Our results demonstrate that Vision Transformer architectures significantly outperform convolutional networks, showing remarkable robustness to reduced spatial resolution. Interpretability analysis reveals a physically consistent hierarchy: spatial features dominate across all parameters, whereas spectral profiles provide secondary constraints for mass and inclination. Applied to observational ALMA data, the framework delivers stable mass and $PA$ estimates with exceptionally tightly constrained inclination angles. This study establishes multi-modal deep learning as a powerful, interpretable tool for overcoming projection biases in high-mass star formation studies.

astro-ph.GA↗

xLLM Technical Report

We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locates online and offline tasks through unified elastic scheduling to maximize cluster utilization. This module also relies on a workload-adaptive dynamic Prefill-Decode (PD) disaggregation policy and a novel Encode-Prefill-Decode (EPD) disaggregation policy designed for multimodal inputs. Furthermore, it incorporates a distributed architecture to provide global KV Cache management and robust fault-tolerant capabilities for high availability. At the engine layer, xLLM-Engine co-optimizes system and algorithm designs to fully saturate computing resources. This is achieved through comprehensive multi-layer execution pipeline optimizations, an adaptive graph mode and an xTensor memory management. xLLM-Engine also further integrates algorithmic enhancements such as optimized speculative decoding and dynamic EPLB, collectively serving to substantially boost throughput and inference efficiency. Extensive evaluations demonstrate that xLLM delivers significantly superior performance and resource efficiency. Under identical TPOT constraints, xLLM achieves throughput up to 1.7x that of MindIE and 2.2x that of vLLM-Ascend with Qwen-series models, while maintaining an average throughput of 1.7x that of MindIE with Deepseek-series models. xLLM framework is publicly available at https://github.com/jd-opensource/xllm and https://github.com/jd-opensource/xllm-service.

cs.DC↗

Online Tensor Inference

Contemporary applications, such as recommendation systems and mobile health monitoring, require real-time processing and analysis of sequentially arriving high-dimensional tensor data. Traditional offline learning, involving the storage and utilization of all data in each computational iteration, becomes impractical for these tasks. Furthermore, existing low-rank tensor methods lack the capability for online statistical inference, which is essential for real-time predictions and informed decision-making. This paper addresses these challenges by introducing a novel online inference framework for low-rank tensors. Our approach employs Stochastic Gradient Descent (SGD) to enable efficient real-time data processing without extensive memory requirements. We establish a non-asymptotic convergence result for the online low-rank SGD estimator, nearly matches the minimax optimal estimation error rate of offline models. Furthermore, we propose a simple yet powerful online debiasing approach for sequential statistical inference. The entire online procedure, covering both estimation and inference, eliminates the need for data splitting or storing historical data, making it suitable for on-the-fly hypothesis testing. In our analysis, we control the sum of constructed super-martingales to ensure estimates along the entire solution path remain within the benign region. Additionally, a novel spectral representation tool is employed to address statistical dependencies among iterative estimates, establishing the desired asymptotic normality.

stat.ML↗

Magnetic field-induced momentum-dependent symmetry breaking in a kagome superconductor

When multiple degrees of freedom share similar energy scales in quantum materials, intertwined electronic orders, which exhibit broken symmetries, are often strongly coupled. Recent studies on kagome superconductors such as CsV$_3$Sb$_5$ report rotational and time-reversal symmetry breaking linked to a charge density wave. Here, we observe a momentum-selective response of the electronic structure of CsV$_3$Sb$_5$ to an external magnetic field. By performing angle-resolved photoemission spectroscopy in a tuneable magnetic field, we demonstrate that the response of the electronic structure is compatible with piezomagnetism along with strong orbital selectivity. Our results show that the origin of the time-reversal symmetry breaking is associated with the vanadium Van Hove singularities at the onset of the charge density wave order. We also demonstrate the presence of fluctuations beyond the charge ordering temperature. Our results reveal that magnetic fields can be used as tuning knobs for disentangling intertwined orders in the momentum space for quantum materials.

cond-mat.str-el↗

Atomically-sharp magnetic soliton in the square-net lattice EuRhAl$_{4}$Si$_{2}$

Topological spin textures are hallmark manifestations of competing interactions in magnetic matter. Their effective description by nonlinear field theories reflects an energetic frustration that destabilizes uniform order while selecting finite-size, topologically nontrivial configurations as stationary states. Among the most extreme realizations are atomically-sharp domain wall excitations, namely one-dimensional (1D) magnetic solitons, which represent the ultimate scaling limit of magnetic textures. Such solitons may emerge in magnetic systems where effective exchange interactions compete directly with uniaxial magnetic anisotropy. Here we show that the square-net rare earth compound EuRhAl$_{4}$Si$_{2}$ realizes a very susceptible regime where the magnetic anisotropy competes with highly frustrated exchange interactions stabilizing a rare ferrimagnetic $\uparrow\uparrow\downarrow$ state that, under applied magnetic field, supports the formation of atomically-sharp soliton defects. We confirm the bulk response of the 1D magnetic solitons via magnetization and electrical transport measurements. We establish both the zero- and in-field $\uparrow\uparrow\downarrow$ order via neutron diffraction, while magnetic force microscopy visualizes its real-space evolution into a stripe-like array. To elucidate the microscopic origin of the soliton, we relate the Ruderman-Kittel-Kasuya-Yosida (RKKY)-driven exchange interactions and the magnetic anisotropy through density functional theory, and we construct an effective 1D $J_{1}$-$J_{2}$-$K$ model whose atomistic spin dynamics simulations reproduce the observed soliton states as a function of external field. Our results demonstrate that EuRhAl$_{4}$Si$_{2}$ hosts atomically-sharp, field-driven 1D magnetic solitons, providing a new platform for studying 1D topological excitations at the atomic length scale.

cond-mat.str-el↗

Practical continuous-variable quantum key distribution using dynamic digital signal processing: security proof and experimental demonstration

Digital signal processing technology has paved the way for the realization of high-speed continuous-variable quantum key distribution systems. However, existing security proofs are limited to static digital signal processing algorithms, while practical systems rely on dynamic multiple-input multiple-output algorithms to compensate for time-varying channel impairments. Our analysis reveals that the conventional dynamic algorithm, due to its non-unitary nature, systematically underestimates the excess noise, which in turn leads to security issues and the generation of insecure keys. To close this gap, we propose a secure algorithm model, mapping the dynamic algorithm to an equivalent physical optical model whose security can be rigorously assessed. Simulations illustrate the algorithm's non-unitary property and provide a quantitative analysis of the excess noise underestimation caused by the conventional algorithm. We further experimentally validate the necessity of the proposed modeling for dynamic digital signal processing, achieving a secret key rate of 14.4 Mbps based on estimated excess noise of 0.07 shot noise unit; whereas the conventional algorithm would have dangerously overestimated the key rate to 28.2 Mbps with noise of 0.008 shot noise unit. This work provides the essential security framework for dynamic digital signal processing, overcoming a critical impediment for the development of high-performance continuous-variable quantum key distribution systems.

quant-ph↗

MAGellanic Outflow and chemistry Survey (MAGOS): Hot cores in the LMC

The Large Magellanic Cloud (LMC) provides a key laboratory for exploring the diversity of star formation and interstellar chemistry under subsolar metallicity conditions. We present the results of a hot core survey toward 30 massive protostellar objects in the LMC using the Atacama Large Millimeter/submillimeter Array (ALMA) at 350 GHz. Continuum imaging reveals 36 compact sources in total, among which line analyses identify 9 hot cores and 1 hot-core candidate, including two newly identified sources. We detect CO, HCO+, H13CO+, HC15N, HC3N, SiO, SO, SO+, NS, SO2, 34SO2, 33SO2, CH3OH, 13CH3OH, HCOOH, HCOOCH3, CH3OCH3, C2H5OH, H2CCO (tentative), and hydrogen recombination lines from hot cores. CH3OCH3, a complex organic molecule larger than CH3OH, is detected for the first time in a hot core outside the LMC bar region. All hot cores show stronger emission in the high-excitation SO line compared to non-hot-core sources, suggesting that its strong detection will be useful for identifying hot-core candidates in the LMC. Chemical analysis reveals a spread of more than two orders of magnitude in CH3OH abundances, with some sources deficient in COMs. In contrast, SO2 is detected in all hot cores, and its abundance shows a good correlation with rotational temperature. The hot cores without CH3OH detections are all located outside the LMC bar region and are characterized by either high luminosity or active star formation in their surroundings. A combination of locally low metallicity, active star formation in the vicinity, and high protostellar luminosity may jointly trigger the COM-poor hot core chemistry observed in the LMC.

astro-ph.GA↗

Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent

With the fast development of big data, learning the optimal decision rule by recursively updating it and making online decisions has been easier than before. We study the online statistical inference of model parameters in a contextual bandit framework of sequential decision-making. We propose a general framework for an online and adaptive data collection environment that can update decision rules via weighted stochastic gradient descent. We allow different weighting schemes of the stochastic gradient and establish the asymptotic normality of the parameter estimator. Our proposed estimator significantly improves the asymptotic efficiency over the previous averaged SGD approach via inverse probability weights. We also conduct an optimality analysis on the weights in a linear regression setting. We provide a Bahadur representation of the proposed estimator and show that the remainder term in the Bahadur representation entails a slower convergence rate compared to classical SGD due to the adaptive data collection.

stat.ML↗

SGD with Dependent Data: Optimal Estimation, Regret, and Inference

This work investigates the performance of the final iterate produced by stochastic gradient descent (SGD) under temporally dependent data. We consider two complementary sources of dependence: $(i)$ martingale-type dependence in both the covariate and noise processes, which accommodates non-stationary and non-mixing time series data, and $(ii)$ dependence induced by sequential decision making. Our formulation runs in parallel with classical notions of (local) stationarity and strong mixing, while neither framework fully subsumes the other. Remarkably, SGD is shown to automatically accommodate both independent and dependent information under a broad class of stepsize schedules and exploration rate schemes. Non-asymptotically, we show that SGD simultaneously achieves statistically optimal estimation error and regret, extending and improving existing results. In particular, our tail bounds remain sharp even for potentially infinite horizon $T=+\infty$. Asymptotically, the SGD iterates converge to a Gaussian distribution with only an $O_{\PP}(1/\sqrt{t})$ remainder, demonstrating that the supposed estimation-regret trade-off claimed in prior work can in fact be avoided. We further propose a new ``conic'' approximation of the decision region that allows the covariates to have unbounded support. For online sparse regression, we develop a new SGD-based algorithm that uses only $d$ units of storage and requires $O(d)$ flops per iteration, achieving the long term statistical optimality. Intuitively, each incoming observation contributes to estimation accuracy, while aggregated summary statistics guide support recovery.

math.ST↗

Designing Information Delays in Supply Chains

This paper studies how a downstream retailer in a decentralized two-tier supply chain can implicitly transmit demand information to an upstream supplier through the structure of its order stream in the absence of an explicit information-sharing mechanism. We distinguish our work from prior work by introducing the notion of information delay and by linking optimal implicit information sharing to the group delay of the retailer's ordering transfer function. We show that pure delay is strictly suboptimal, while fractional-delay mechanisms can reshape the order autocorrelation to improve supplier forecastability and reduce system-wide inventory costs. Using Hardy-space factorization, we develop a tractable family of invertible ARMA policies that approximates the theoretically optimal (but non-rational) limiting filter derived by Caldentey et al. (2025) and preserves its informational delay properties. This construction yields sharp guidance on how policy complexity, as measured by the degrees of the ARMA policies, impacts supply chain costs. We further extend the analysis to memory-constrained suppliers and characterize how the complexity of the retailer's policy should scale with the supplier's finite forecasting window, highlighting when, perhaps counterintuitively, increasing policy complexity can become counterproductive.

math.OC↗

TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning

Current autonomous driving systems often favor end-to-end frameworks, which take sensor inputs like images and learn to map them into trajectory space via neural networks. Previous work has demonstrated that models can achieve better planning performance when provided with a prior distribution of possible trajectories. However, these approaches often overlook two critical aspects: 1) The appropriate trajectory prior can vary significantly across different driving scenarios. 2) Their trajectory evaluation mechanism lacks policy-driven refinement, remaining constrained by the limitations of one-stage supervised training. To address these issues, we explore improvements in two key areas. For problem 1, we employ MoE to apply different trajectory priors tailored to different scenarios. For problem 2, we utilize Reinforcement Learning to fine-tune the trajectory scoring mechanism. Additionally, we integrate models with different perception backbones to enhance perceptual features. Our integrated model achieved a score of 51.08 on the navsim ICCV benchmark, securing third place.

cs.CV↗

LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs

Visual encoding followed by token condensing has become the standard architectural paradigm in multi-modal large language models (MLLMs). Many recent MLLMs increasingly favor global native- resolution visual encoding over slice-based methods. To investigate this trend, we systematically compare their behavior on vision-language understanding and attention patterns, revealing that global encoding enhances overall capability but at the expense of greater computational overhead. To address this issue, we present LLaVA-UHD v3, an MLLM centered upon our proposed Progressive Visual Compression (PVC) method, which can be seamlessly integrated into standard Vision Transformer (ViT) to enable efficient native-resolution encoding. The PVC approach consists of two key modules: (i) refined patch embedding, which supports flexible patch-size scaling for fine-grained visual model- ing, (ii) windowed token compression, hierarchically deployed across ViT layers to progressively aggregate local token representations. Jointly modulated by these two modules, a widely pretrained ViT can be reconfigured into an efficient architecture while largely preserving generality. Evaluated across extensive benchmarks, the transformed ViT, termed ViT-UHD, demonstrates competitive performance with MoonViT while reducing TTFT (time-to-first-token) by 2.4x, when developed within an identical MLLM architecture. Building upon ViT-UHD, LLaVA-UHD v3 also achieves competitive performance to Qwen2-VL, while further reducing TTFT by 1.9x. We will release all code and checkpoints to support future research on efficient MLLMs.

cs.CV↗

Spin Excitations and Flat Electronic Bands in a Cr-based Kagome Superconductor

In the quest for topology- and correlation-driven quantum states, kagome lattice materials have garnered significant interest for their band structures, featuring flat bands (FBs) from the quantum destructive interference of the electronic wavefunction. Tuning an FB to the chemical potential could induce electronic instabilities and emergent orders. Despite extensive studies, direct evidence of FBs tuned to the chemical potential and their role in emergent orders in bulk materials remains lacking. Using angle-resolved photoemission spectroscopy, resonant inelastic X-ray scattering, and density functional theory, we show that the low-energy structure of the Cr-based kagome metal superconductor {\Cr} is dominated by FBs at the Fermi level. We also observe low-energy magnetic excitations evolving across the low-temperature transition, largely consistent with the FB shift. Our results suggest that the low-temperature order contains a magnetic origin and that the kagome FBs may play a role in the emergence of this order.

cond-mat.str-el↗

The ALMA-QUARKS survey: Evidence of a candidate high-mass prestellar core aside a bright-rimmed cloud IRAS 18290-0924

Although frequently reported in observations, the definitive confirmation of high-mass prestellar cores has remained elusive, presenting a persistent challenge in star formation studies. Using two-band observational data from the 3mm ATOMS and 1.3mm QUARKS surveys, we report a high-mass prestellar core candidate, C2, located on the side of the bright-rimmed cloud IRAS 18290-0924. The C2 core identified from the 3mm continuum data of the ATOMS survey ($\sim$2 arcsecond, $\rm\sim 10000~au$ at 5.3 kpc) has a mass ranging from 27-68 $M_{\odot}$ for temperatures 10-22K within a radius of $\sim$2800 au. The highest-resolution ($\sim$0.3 arcsecond, $\rm\sim 1500 au$) observations of this source presented to date from the QUARKS survey reveal no evidence of further fragmentation. Further analysis of a total $\sim$10 GHz band width of molecular line survey does not find star-formation activity (e.g., outflows, ionized gas) associated with the core, with a few molecular lines of cold gas detected only. Additionally, virial analysis indicates the C2 core is gravitationally bound ($α_{\rm vir} \sim0.1-0.3$) and thus could be undergoing collapse toward star formation. These results strongly establish a candidate for a high-mass prestellar core, contributing to the very limited number of such sources known to date.

astro-ph.GA↗

Tunable Asymmetric Delay Attack in Quantum Clock Synchronization

Quantum clock synchronization underpins modern secure communications and critical infrastructure, yet its fundamental dependence on channel reciprocity introduces an exploitable vulnerability to asymmetric delay attacks. Current attack strategies rely on static delays, limiting their ability to target application-specific stability requirements. Here, we propose a tunable asymmetric delay attack (T-ADA) that dynamically controls delay parameters to induce manipulate synchronization accuracy. Through experimental implementation, we demonstrate how tailored attack trajectories can selectively compromise system stability across different scenarios. This work uncovers key vulnerabilities in synchronization protocols under customizable attacks and provide a foundation for developing secure and resilient quantum clock synchronization systems.

quant-ph↗

The SOFIA Massive (SOMA) Radio Survey. II. Radio Emission from High-Luminosity Protostars

We present centimeter continuum observations of seven high luminosity massive protostars and their surrounding sources in regions with multiple targets, as part of the SOFIA Massive (SOMA) Star Formation Survey. With data from the Very Large Array and the Australia Telescope Compact Array, we analyze the spectral index, morphology and multiplicity of the detected radio sources. The high-sensitivity, high-resolution observations allow us to resolve many sources; 65$\%$ of the reported sources are resolved at least within the synthesized beam. We report thirteen new detections and two previously known detections that we observed for the first time in radio frequencies. We use the observations to build radio spectral energy distributions (SEDs) to calculate spectral indices. With radio morphologies and the spectral indices, we give assessments on the nature of the sources, highlighting six sources that display a radio jet-like morphology and a spectral index consistent with ionized jets. Combining with the SOMA Radio I sample, we present the radio - bolometric luminosity relation, especially probing the regime from $L_{\rm bol}\sim 10^4$ to $10^6\:L_\odot$. Here we find a steep rise in radio luminosity, which is expected by models that transition from shock ionization to photoionization.

astro-ph.SR↗

Quantum oscillations and anisotropic magnetoresistance in the quasi-two-dimensional Dirac nodal line superconductor $\mathrm{YbSb_2}$

Recent interest in quantum materials has focused on systems exhibiting both superconductivity and non-trivial band topology as material candidates to realize topological or unconventional superconducting states. So far, superconductivity in most topological materials has been identified as type II. In this work, we present magnetotransport studies on the quasi-two-dimensional type I superconductor $\mathrm{YbSb_2}$. Combined ab initio DFT calculations and quantum oscillation measurements confirm that $\mathrm{YbSb_2}$ is a Dirac nodal line semimetal in the normal state. The complex Fermi surface morphology is evidenced by the non-monotonic angular dependence of both the quantum oscillation amplitude and the magnetoresistance. Our results establish $\mathrm{YbSb_2}$ as a candidate material platform for exploring the interplay between band topology and superconductivity.

cond-mat.supr-con↗