SearcharxivSearch

arXiv subjects

Yukun Zhang

Publications and source records attributed to Yukun Zhang.

At least 19 recordsLinked to original sources

Quantum Advantage with Adaptive Shallow Circuits

Quantum advantage is widely expected to require sufficiently deep circuits, where correlations and global computational structure can grow beyond the reach of efficient classical simulation. This expectation is especially stark for constant-depth circuits with local readout: the expectation value of any fixed local observable lies within a bounded backward lightcone and is therefore classically tractable. Here we show that measurement feedback changes this picture. We establish a strict hierarchy of computational power: at fixed coherent depth, increasing the number of feedback outcomes strictly enlarges the class of functions accessible through a local expectation value. The two ends of this hierarchy exhibit distinct computational regimes. With logarithmic feedback, local expectation values for product-state inputs are efficiently classically simulable. Polynomial feedback, by contrast, enables an explicit family of adaptive shallow circuits to encode prime-field discrete logarithm problem~(DLP) into a fixed single-qubit expectation. Assuming the standard worst-case classical hardness of DLP, estimating this expectation value is classically hard. These results reveal a feedback-driven complexity transition, with further implications for resource lower bounds on DLP and the complexity of local-observable estimation under area-law entanglement. Our results open a new route to quantum advantage with shallow quantum circuits.

quant-ph

Efficient Lindbladian Learning from Constant-Time Pauli Responses

Learning the generator of an open many-body system is more challenging than Hamiltonian learning: local responses, which can directly reveal coherent interaction terms in closed-system dynamics, may also contain dissipative contributions in open-system dynamics. In this paper, we address this challenge by developing an efficient Lindbladian learning framework for a known local candidate generator dictionary with bounded dissipative support and either bounded dual-interaction-graph degree or bounded unweighted local strength. The framework resolves the coherent-dissipative ambiguity by treating local Pauli responses as a linear system over both types of generator terms. Inverting this response system separates their contributions and makes the individual Lindbladian coefficients accessible from local response data in a fixed short-time window. Within this framework, we develop two efficient learning algorithms: Chebyshev--Lobatto response interpolation, which uses logarithmically many short evolution times and has a post-mean cost linear in $M$, with the stated dependence on $\epsilon$, and Single-time projected response contraction, which uses a single fixed evolution time and globally inverts a truncated response function. Both procedures estimate $M$ candidate coefficients to entrywise accuracy $\epsilon$ using $\widetilde{\mathcal{O}}(M/\epsilon^2)$ sample and classical post-processing complexity. Our theoretical results establish local response inversion as a scalable paradigm for learning, calibrating, and diagnosing complex quantum systems from experimentally accessible short-time data.

quant-ph

CausalMix: Data Mixture as Causal Inference for Language Model Training

In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.

cs.LG

Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents

AI agents increasingly operate inside digital accounts by exercising privileges that users already hold, raising a new control question: whether an existing account entitlement must be exercised manually or may be exercised through a user-authorized automated proxy. We define \emph{delegation rights} as the revocable, identity-preserving, scope-limited, and mode-specific authority of an account holder to authorize such proxy execution. We develop a three-party incomplete-contracts model with a User, an AI Agent provider, and a Platform. The contested object is not platform ownership, account transferability, data portability, or unrestricted API access, but residual control over the mode of account execution. Under Platform Control, the platform can protect infrastructure, identity systems, privacy boundaries, and third parties, but its discretionary veto weakens the User--Agent coalition's disagreement payoff and depresses relationship-specific investment. Under User Control, hold-up is reduced, but security, privacy, congestion, and third-party risks may remain insufficiently internalized. We then analyze \emph{Certified Delegation}, under which access protection is conditional on verifiable authorization, revocability, auditability, rate-limit compliance, data minimization, and risk mitigation. Certification is therefore not merely a technical safety screen; it is a conditional allocation of residual control. Illustrative mechanism simulations show how this regime can reduce deadweight loss by restoring delegation incentives while bounding residual risk.

econ.EM

PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting

Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables remains challenging due to scale differences. To address this, we introduce patch-mean decoupling (PMD), which separates the trend and residual shape information by subtracting the mean of each patch, preserving the original structure and ensuring that the attention mechanism captures true shape similarities. Futhermore, to more effectively model long-range dependencies and capture cross-variable relationships, we propose Trend Restoration Attention (TRA) and Proximal Variable Attention (PVA). The former module reintegrates the decoupled trend from PMD while calculating attention output. And the latter focuses cross-variable attention on the most relevant, recent time segments to avoid overfitting on outdated correlations. Combining these components, we propose PMDformer, a model designed to effectively capture shape similarity in long-term forecasting scenarios. Extensive experiments indicate that PMDformer outperforms existing state-of-the-art methods in stability and accuracy across multiple LTSF benchmarks. The code is available at https://github.com/aohu1105/PMDformer.

cs.AI

Simultaneous Estimation of Partial-Transpose Moments with Active Memory Independent of the Moment Order

We study the simultaneous estimation of partial-transpose moments $p_j(\rho_{AB})=\mathrm{Tr}[(\rho_{AB}^{T_B})^j]$, $j=2,\ldots,K$, of an unknown bipartite $n$-qubit state from independent copies under an explicit active-memory constraint. We give a sequential qubit-reuse realization of the partial-transpose permutation that uses at most $2n+1$ active qubits, independent of $K$, and estimates all moments $p_2,\ldots,p_K$ to uniform additive error $\epsilon$ with total copy complexity $O(K\log K/\epsilon^2)$. We also prove two converse bounds. First, any uniformly accurate simultaneous estimator requires $\Omega(K/\epsilon^2)$ copies in the worst case. Second, the same scaling holds on an explicit isospectral two-qubit negative-partial-transpose (NPT) family whose ordinary moments are constant while the partial-transpose moments vary. These results characterize the copy complexity of the partial-transpose moment hierarchy up to a logarithmic factor and extend simultaneous nonlinear-functional estimation from ordinary state powers to partial-transpose spectral data under active quantum memory independent of the target moment order.

quant-ph

Hardware-Efficient and Performance-Enhanced Joint Pulse Shaping and Dispersion Compensation for Coherent Data Center Interconnects

With the explosion of data traffic triggered by 5G/6G and Generative artificial intelligence, coherent optical communication is moving towards higher baud rates and more complex modulation formats. This leads to a significant increase in the computational complexity and power consumption of digital signal processing (DSP) at the transmitter and receiver ends, especially in the chromatic dispersion(CD) Compensation and low roll-off shaping filter modules. We propose a joint shaping filtering and CD compensation (JFS-CD) algorithm. This algorithm moves the CD compensation to the transmitter side and utilizes the characteristics of discrete fourier transform and the spectral features of shaping filtering for integrated processing. Aiming at the high peak-to-average power ratio (PAPR) problem caused by chromatic dispersion pre-compensation, we propose a low-complexity square boundary clipping algorithm(SBC). Simulation results show that, under the premise of maintaining unchanged performance, JFS-CD can reduce the real multiplication complexity by about 46%. Meanwhile, benefiting from the suppression of the effects of system nonlinearity and receiver IQ imbalance, the joint JFS-CD and SBC scheme improves the Q-factor by about 0.3 dB in experiments compared to the traditional post-chromatic dispersion compensation scheme. This research provides a highly potential transmitter DSP solution for next-generation low-power and high-performance data center interconnects (DCI).

cs.IT

Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise manipulation in diffusion-based methods to token-level operations, VAR-based approaches achieve better background preservation and significantly faster inference. However, existing VAR-based editing methods still face two key challenges: accurately localizing editable tokens and maintaining structural consistency in the edited results. In this work, we propose a novel text-guided image editing framework rooted in an analysis of intermediate feature distributions within VAR models. First, we introduce a coarse-to-fine token localization strategy that can refine editable regions, balancing editing fidelity and background preservation. Second, we analyze the intermediate representations of VAR models and identify structure-related features, by which we design a simple yet effective feature injection mechanism to enhance structural consistency between the edited and source images. Third, we develop a reinforcement learning-based adaptive feature injection scheme that automatically learns scale- and layer-specific injection ratios to jointly optimize editing fidelity and structure preservation. Extensive experiments demonstrate that our method achieves superior structural consistency and editing quality compared with state-of-the-art approaches, across both local and global editing scenarios.

cs.CV

Efficient Noisy Quantum State and Process Tomography

Efficiently characterizing large quantum states and processes is a central yet notoriously challenging task in quantum information science, as conventional tomography methods typically require resources that grow exponentially with system size. Here, we introduce a structure-agnostic learning framework for noisy $n$-qubit quantum circuits under~i.i.d.~single-qubit noise. We first prove that quantum states with unital noise channels admit an efficient learnable representation in the logarithmic-depth regime. We then extend this framework to quantum process tomography under constant noise, deriving a unified protocol that applies to both unital and non-unital noisy channels and retains efficient guarantees for logarithmic-depth circuits. This process-learning formulation is input-agnostic and imposes no distributional assumptions on the input quantum states. We further study a more general regime with arbitrary noise strength. In this setting, low-weight Pauli propagation induces a terminal truncation whose threshold depends logarithmically on the inverse accuracy, leading to quasi-polynomial complexity and near-unit success probability in the average case. In contrast to the preceding two results, this arbitrary-noise guarantee does not impose any restriction on the circuit depth, and therefore covers arbitrary-depth circuits, including both the noiseless limit ($\gamma = 0$) and the strong-decoherence regime ($\gamma = \Theta(1)$). Numerical simulations of two-dimensional Hamiltonian dynamics further demonstrate the accuracy and robustness of the approach, including for structured circuits beyond the random-circuit setting assumed in the theoretical analysis. These results provide a scalable and practically relevant route toward characterizing large-scale noisy quantum devices, addressing a key bottleneck in the development of quantum technologies.

quant-ph

How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning

When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how does this structure evolve across layers? We address this question with a layer-wise framework based on Partial Information Decomposition (PID) that decomposes the predictive information at each Transformer layer into redundant, vision-unique, language-unique, and synergistic components. To make PID tractable for high-dimensional neural representations, we introduce \emph{PID Flow}, a pipeline combining dimensionality reduction, normalizing-flow Gaussianization, and closed-form Gaussian PID estimation. Applying this framework to LLaVA-1.5-7B and LLaVA-1.6-7B across six GQA reasoning tasks, we uncover a consistent \emph{modal transduction} pattern: visual-unique information peaks early and decays with depth, language-unique information surges in late layers to account for roughly 82\% of the final prediction, and cross-modal synergy remains below 2\%. This trajectory is highly stable across model variants (layer-wise correlations $>$0.96) yet strongly task-dependent, with semantic redundancy governing the detailed information fingerprint. To establish causality, we perform targeted Image$\rightarrow$Question attention knockouts and show that disrupting the primary transduction pathway induces predictable increases in trapped visual-unique information, compensatory synergy, and total information cost -- effects that are strongest in vision-dependent tasks and weakest in high-redundancy tasks. Together, these results provide an information-theoretic, causal account of how vision becomes language in multimodal Transformers, and offer quantitative guidance for identifying architectural bottlenecks where modality-specific information is lost.

cs.AI

Where to Add PDE Diffusion in Transformers

Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of locality-inducing modules in hybrid architectures has largely been empirical. We study a simple deterministic PDE diffusion layer implemented as one explicit Euler step of one-dimensional heat smoothing using a discrete Neumann Laplacian under a spectral stability constraint, and ask a structural question: where should diffusion be inserted relative to attention? Our central claim is that diffusion and attention generally do not commute, so inserting the same local operator before versus after attention leads to qualitatively different behaviors. We develop a three-layer operator-theoretic framework that (1) establishes unconditional guarantees for the diffusion subsystem, including spectral non-expansiveness and monotone Dirichlet-energy dissipation when the diffusion step size is smaller than one half, (2) derives compositional perturbation bounds linking insertion effects to representation roughness and downstream amplification, and (3) uses diffusion-attention non-commutativity as a diagnostic for structural double-mixing conflicts. Guided by theory, we evaluate seven insertion positions on the Long Range Arena benchmark. Early diffusion acts as effective pre-regularization, improving average accuracy by 4.1 percentage points when applied after embedding, while post-attention diffusion degrades performance by 2.5 percentage points, consistent with the predicted conflict. A multi-scale diffusion variant yields consistent gains under the same global stability constraint. Our analysis provides a general template for reasoning about local-global compositions in sequence models by separating provable guarantees, compositional bounds, and mechanistic diagnostics.

cs.LG

SMES: Towards Scalable Multi-Task Recommendation via Expert Sparsity

Industrial recommender systems typically rely on multi-task learning to estimate diverse user feedback signals and aggregate them for ranking. Recent advances in model scaling have shown promising gains in recommendation. However, naively increasing model capacity imposes prohibitive online inference costs and often yields diminishing returns for sparse tasks with skewed label distributions. This mismatch between uniform parameter scaling and heterogeneous task capacity demands poses a fundamental challenge for scalable multi-task recommendation. In this work, we investigate parameter sparsification as a principled scaling paradigm and identify two critical obstacles when applying sparse Mixture-of-Experts (MoE) to multi-task recommendation: exploded expert activation that undermines instance-level sparsity and expert load skew caused by independent task-wise routing. To address these challenges, we propose SMES, a scalable sparse MoE framework with progressive expert routing. SMES decomposes expert activation into a task-shared expert subset jointly selected across tasks and task-adaptive private experts, explicitly bounding per-instance expert execution while preserving task-specific capacity. In addition, SMES introduces a global multi-gate load-balancing regularizer that stabilizes training by regulating aggregated expert utilization across all tasks. SMES has been deployed in Kuaishou large-scale short-video services, supporting over 400 million daily active users. Extensive online experiments demonstrate stable improvements, with GAUC gain of 0.29% and a 0.31% uplift in user watch time.

cs.IR

The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox

This paper develops a micro-founded economic theory of the AI industry by modeling large language models as a distinct asset class-Digital Intelligence Capital-characterized by data-compute complementarities, increasing returns to scale, and relative (rather than absolute) valuation. We show that these features fundamentally reshape industry dynamics along three dimensions. First, because downstream demand depends on relative capability, innovation by one firm endogenously depreciates the economic value of rivals' existing capital, generating a persistent innovation pressure we term the Red Queen Effect. Second, falling inference prices induce downstream firms to adopt more compute-intensive agent architectures, rendering aggregate demand for compute super-elastic and producing a structural Jevons paradox. Third, learning from user feedback creates a data flywheel that can destabilize symmetric competition: when data accumulation outpaces data decay, the market bifurcates endogenously toward a winner-takes-all equilibrium. We further characterize conditions under which expanding upstream capabilities erode downstream application value (the Wrapper Trap). A calibrated agent-based model confirms these mechanisms and their quantitative implications. Together, the results provide a unified framework linking intelligence production upstream with agentic demand downstream, offering new insights into competition, scalability, and regulation in the AI economy.

econ.GN

Generative AI as a Non-Convex Supply Shock: Market Bifurcation and Welfare Analysis

The diffusion of Generative AI (GenAI) constitutes a supply shock of a fundamentally different nature: while marginal production costs approach zero, content generation creates congestion externalities through information pollution. We develop a three-layer general equilibrium framework to study how this non-convex technology reshapes market structure, transition dynamics, and social welfare. In a static vertical differentiation model, we show that the GenAI cost shock induces a kinked production frontier that bifurcates the market into exit, AI, and human segments, generating a ``middle-class hollow'' in the quality distribution. To analyze adjustment paths, we embed this structure in a mean-field evolutionary system and a calibrated agent-based model with bounded rationality. The transition to the AI-integrated equilibrium is non-monotonic: rather than smooth diffusion, the economy experiences a temporary ecological collapse driven by search frictions and delayed skill adaptation, followed by selective recovery. Survival depends on asymmetric skill reconfiguration, whereby humans retreat from technical execution toward semantic creativity. Finally, we show that the welfare impact of AI adoption is highly sensitive to pollution intensity: low congestion yields monotonic welfare gains, whereas high pollution produces an inverted-U relationship in which further AI expansion reduces total welfare. These results imply that laissez-faire adoption can be inefficient and that optimal governance must shift from input regulation toward output-side congestion management.

econ.GN

The Impact of Generative AI on Content Platforms: A Two-Sided Market Analysis with Multi-Dimensional Quality Heterogeneity

This paper presents a unified computational framework to examine how generative AI (GenAI) reshapes welfare, inequality, and diversity in content platform economies. By integrating welfare economics with agent-based simulations, we model the co-evolutionary dynamics among AI generators, human creators, and consumers within a two-sided market characterized by multi-dimensional quality heterogeneity. Unlike static models, our framework endogenizes AI learning as a function of human data synthesis and models human adaptation as a strategic reallocation of skills toward creative niches. The results reveal that while GenAI significantly enhances consumer surplus through technical quality gains and price depression, it triggers a skill-biased displacement of human incumbents and intensifies market concentration. Through the evaluation of six governance regimes, we identify a fundamental ``Policy Trilemma'' where platforms must navigate non-trivial trade-offs between allocative efficiency, distributional equity, and ecosystem sustainability. Our findings highlight that algorithmic diversity and pro-creative commission structures function as essential economic mechanisms for sustaining long-tail participation and inclusive social welfare in the generative AI era.

cs.CY

Low-Complexity Monitoring and Compensation of Transceiver IQ Imbalance by Multi-dimensional Architecture for Dual-Polarization 16 Quadrature Amplitude Modulation

In this paper, a low-complexity multi-dimensional architecture for IQ imbalance compensation is proposed, which reduces the effects of in-phase (I) and quadrature (Q) imbalance. The architecture use a transceiver IQ skew estimation structure to compensate for IQ skew, and then use a low-complexity MIMO equalizer to compensate for IQ amplitude/phase imbalance. In the transceiver IQ skew estimation structure, the receiver(RX) IQ skew is estimated by Gardner's phase detector, and the transmitter TX skew is estimated by finding the value that yields the lowest equalizer error. The low-complexity MIMO equalizer consists of a complex-valued MIMO (CV-MIMO) and a two-layer multimodulus algorithm real-valued MIMO (TMMA-RV-MIMO), which employ a butterfly and a non-butterfly structure, respectively. The CV-MIMO is used to perform polarization demultiplexing and the TMMA-RV-MIMO equalizes each of the two polarizations. In addition, the TMMA-RV-MIMO can recovery the carrier phase. A 100 km transmission simulation and experiment with 36 Gbaud dual-polarization 16 quadrature amplitude modulation (DP-16QAM) signals showed that, with the TX/RX IQ skew estimation, the estimation error is less than 0.9/0.25 ps. The low-complexity MIMO equalizer can tolerate 0.1 TX IQ amplitude imbalance and 5 degrees at a 0.3 dB Q-factor penalty. The number of real multiplications is reduced by 55% compared with conventional cases in total.

cs.NI

The Economics of Information Pollution in the Age of AI: General Equilibrium, Welfare, and Policy Design

The advent of Large Language Models (LLMs) represents a fundamental shock to the economics of information production. By asymmetrically collapsing the marginal cost of generating low-quality, synthetic content while leaving high-quality production costly, AI systematically incentivizes information pollution. This paper develops a general equilibrium framework to analyze this challenge. We model the strategic interactions among a monopolistic platform, profit-maximizing producers, and utility-maximizing consumers in a three-stage game. The core of our model is a production technology with differential elasticities of substitution (σ_L > 1 > σ_H), which formalizes the insight that AI is a substitute for labor in low-quality production but a complement in high-quality creation. We prove the existence of a unique "Polluted Information Equilibrium" and demonstrate its inefficiency, which is driven by a threefold market failure: a production externality, a platform governance failure, and an information commons externality. Methodologically, we derive a theoretically-grounded Information Pollution Index (IPI) with endogenous welfare weights to measure ecosystem health. From a policy perspective, we show that a first-best outcome requires a portfolio of instruments targeting each failure. Finally, considering the challenges of deep uncertainty, we advocate for an adaptive governance framework where policy instruments are dynamically adjusted based on real-time IPI readings, offering a robust blueprint for regulating information markets in the age of AI.

cs.CY

Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs

This research project addresses the errors of financial numerical reasoning Question Answering (QA) tasks due to the lack of domain knowledge in finance. Despite recent advances in Large Language Models (LLMs), financial numerical questions remain challenging because they require specific domain knowledge in finance and complex multi-step numeric reasoning. We implement a multi-retriever Retrieval Augmented Generators (RAG) system to retrieve both external domain knowledge and internal question contexts, and utilize the latest LLM to tackle these tasks. Through comprehensive ablation experiments and error analysis, we find that domain-specific training with the SecBERT encoder significantly contributes to our best neural symbolic model surpassing the FinQA paper's top model, which serves as our baseline. This suggests the potential superior performance of domain-specific training. Furthermore, our best prompt-based LLM generator achieves the state-of-the-art (SOTA) performance with significant improvement (>7%), yet it is still below the human expert performance. This study highlights the trade-off between hallucinations loss and external knowledge gains in smaller models and few-shot examples. For larger models, the gains from external facts typically outweigh the hallucination loss. Finally, our findings confirm the enhanced numerical reasoning capabilities of the latest LLM, optimized for few-shot learning.

cs.CL