SearcharxivSearch

arXiv subjects

Zhichao Chen

Publications and source records attributed to Zhichao Chen.

At least 19 recordsLinked to original sources

Integrated yoctosecond-precision timing detector

Precise timing detection is essential for exploring ultrafast phenomena in fields ranging from free-electron lasers to ultra-high-power laser facilities. However, achieving simultaneous high resolution and large dynamic range remains a fundamental challenge, and state of the art systems are often constrained by their physical size, power requirements, and limited scalability. Here we introduce an integrated dual electro optic sub cycle timing detector (DEST) that overcomes these limitations. In a proof of principle measurement, the device resolves timing jitter as small as 11 yoctoseconds (ys, $10^{-24}$ s) at 1 MHz--equivalent to the transit time of light across two protons--while maintaining an unambiguous measurement range of 6.15 ps and a dynamic range exceeding 155 dB. The core detection unit is miniaturized to chip scale dimensions of 18 mm x 2 mm x 1 mm on a thin film lithium niobate platform, ensuring inherent stability and immunity to environmental disturbances. Moreover, the architecture naturally lends itself to massive parallelization through array integration, with the potential to push timing precision to the sub 10 rontosecond (rs, $10^{-27}$ s) level within a 1 $m^2$ footprint. This combination of extreme sensitivity, wide dynamic range, compact size, and scalability opens new avenues for detecting previously inaccessible weak signals, including those from gravitational waves, quantum vacuum fluctuations, and beyond.

physics.optics

Unfolding the Interdisciplinary Complexities of Climate Science: Fuxi-Climate Foundational Model

Climate research and decision-making require integrating evidence across physical processes, socio-economic dynamics and policy responses. Large language models (LLMs) have been explored for accessing and synthesizing climate knowledge, but their ability to support structured interdisciplinary reasoning is still limited. Here we present the Fuxi-Climate Foundation Model (CFM), a climate-specialized LLM designed to support consistent reasoning across domains. CFM maintains more stable analytical behavior as interdisciplinary complexity increases, whereas performance in other models becomes more variable. On expert-designed climate transition tasks, CFM produces more structured analyses that explicitly address trade-offs and uncertainty, achieving 45% trade-off coverage and 47.27% uncertainty-aware reasoning. These results indicate that CFM can support more realistic analysis of climate risks and transition pathways, and provide a basis for agent-based systems to explore complex policy and decision scenarios. The model is openly available at https://huggingface.co/SII-yuning/cfm.

cs.CY

Mutation-preserving generalized cluster algebras and Laurent mutation invariants

We introduce mutation-preserving generalized cluster algebras, for which the generalized cluster mutation in each direction is independent of the seed in the mutation equivalence class. We classify all irreducible generalized cluster algebras with this property. Then, a Markov-type Diophantine equation $x^2+y^2+z^2+2yz=kxyz$ is studied, which has a structure of the mutation-preserving generalized cluster algebra. We prove that positive integer solutions exist if and only if $1\leq k\leq 5$ and determine all their orbits under the associated generalized cluster mutation groups. In particular, multiple orbits occur for each $k=1,3$, whereas the solutions form a single orbit for each $ k=2,4,5$. We then classify all generalized Markov Laurent mutation invariants. As an application, a conjecture proposed by Chen-Li is proved, showing that every Laurent mutation invariant of irreducible sign-equivalent cluster algebras is essentially a polynomial in the corresponding basic invariant.

math.RA

Battery-Swapping Station Operation Under Forecast Uncertainty: A Scenario-Based Stochastic MPC Framework

Battery-swapping stations (BSSs) can shorten electric-vehicle energy replenishment while using centrally managed battery inventories as flexible grid-connected storage. Realizing both benefits requires the station to schedule charging, grid discharge, and swapping service before future customer demand and electricity prices are known. This paper develops a forecast-aware rolling-horizon operating framework for this problem. A lightweight DLinear model predicts 24-hour price and demand trajectories, Stein variational gradient descent quantifies their uncertainty through representative scenarios, and a two-stage stochastic model predictive controller converts those scenarios into station decisions. The controller accounts for service shortfall, terminal readiness, a protected service buffer, and electrochemical degradation without assuming perfect future information. The application contribution is an implementable controller that coordinates the station's mobility-service and energy-storage roles. The methodological contribution is a modular forecast-to-control interface that separates the operational value of mean-forecast accuracy from that of uncertainty representation. In a 120-day closed-loop evaluation, DLinear-SVGD SMPC achieves the lowest cost among the implementable controllers. Relative to deterministic DLinear MPC, it reduces final cost by 1.2\% and service-shortfall hours by 80.7\%, with 99.10\% of the evaluated hours free of shortfall.

math.OC

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from human-written rubrics, preference data, or sampled responses. Direct query-to-rubric generation avoids these resources, but provides no explicit check that a plausible rubric is useful. Such a rubric may fail to distinguish answer quality, reward an optional style, or penalize a valid alternative strategy. We introduce Rubrics on Trial, a query-only framework that evolves a rubric set from an empty set without external annotations or model training. It derives supervision solely from synthetic rubric-conditioned response pairs and validates each proposed rubric before adding it, screening out non-discriminative, over-specific, and style-only candidate rubrics. Experiments across five preference benchmark suites demonstrate the effectiveness of Rubrics on Trial, which achieves the best average accuracy and leads on six of seven evaluation sets.

cs.CL

Rethinking the Flow-Based Gradual Domain Adaptation: A Semi-Dual Optimal Transport Perspective

Gradual domain adaptation (GDA) aims to mitigate domain shift by progressively adapting models from the source domain to the target domain via intermediate domains. However, real intermediate domains are often unavailable or ineffective, necessitating the synthesis of intermediate samples. Flow-based models have recently been used for this purpose by interpolating between source and target distributions. Notably, their training typically relies on sample-based log-likelihood estimation, which can discard useful information and thus degrade GDA performance. The key to addressing this limitation is constructing the intermediate domains via samples directly. To this end, we propose an Entropy-regularized Semi-dual Unbalanced Optimal Transport (E-SUOT) framework to construct intermediate domains. Specifically, we reformulate flow-based GDA as a Lagrangian dual problem and derive an equivalent semi-dual objective that circumvents the need for likelihood estimation. However, the dual problem leads to an unstable min-max training procedure. To alleviate this issue, we further introduce the entropy regularization to convert it into a more stable sequential optimization procedure. Based on this, we propose a novel GDA training framework and provide theoretical analysis in terms of stability and generalization. Finally, extensive experiments are conducted to demonstrate the efficacy of the E-SUOT framework.

cs.LG

Geometric structures of $G$-fans associated with rank $3$ cluster-cyclic exchange matrices

In this paper, we investigate the geometric structures of $G$-fans associated with rank $3$ real cluster-cyclic exchange matrices. In this class, a simple recursion for tropical signs was found, which enables us to study the detailed properties of $c$-, $g$-vectors. We introduce two kinds of upper bounds of the $G$-fans. The first one is the global upper bound, which comes from a hyperbolic surface containing all $g$-vectors after an initial mutation. The second one is the local upper bound, which reflects the internal separateness structure. As applications, we prove that there is no periodicity among $g$-vectors, and we completely determine the sign of $g$-vectors. We also prove the monotonicity of $g$-vectors under the minimum assumption. Moreover, we show that the three global upper bounds can be simplified to a single uniform upper bound.

math.CO

Real $C$-, $G$-structures and sign-coherence of cluster algebras

We generalize the theory of integer $C$-, $G$-matrices in cluster algebras to the real case. By a skew-symmetrizing method, we can reduce the problem of skew-symmetrizable patterns to skew-symmetric patterns. In this sense, the sign-coherence of a more general real class called of quasi-integer type can be inherited directly from that of integer $C$-, $G$-matrices proved by Gross-Hacking-Keel-Kontsevich. However, the sign-coherence of real $C$-, $G$-matrices does not always hold in general. For this purpose, we classify all the rank $2$ case and the finite type case via the Coxeter diagrams. We also give two conjectures about the real exchange matrices and $C$-, $G$-matrices. Under these conjectures, the dual mutation, $G$-fan structure and synchronicity property hold. As an application, the isomorphism of several kinds of exchange graphs is studied.

math.RT

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

We propose Intrinsic Quality (IQ), a validation-free metric designed to estimate the inherent potential of face recognition (FR) datasets to produce high-performance models without the need for full-scale training. IQ integrates two components: (i) a Neighbor-Consistency Score that quantifies local identity label agreement via nearest neighbors, and (ii) Global Representation Subspace Complexity (Effective Rank, ER), which captures the underlying embedding geometry and dataset diversity. IQ allows for rapid evaluation using lightweight proxy models or data subsets, facilitating dataset diagnosis and curation prior to resource-intensive full-scale training. We describe an experimental protocol tailored to clean, noisy, and mixed-quality FR datasets, and outline evaluation methodologies to validate IQ's predictive power for downstream performance.

cs.CV

Log-concavity and unimodality of cluster monomials of type $A_3$

The log-concavity of cluster variables of type $A_n$ and cluster monomials of type $A_2$ was established by Chen-Huang-Sun. It is still a conjecture for the cluster monomials of higher rank. In this paper, we prove the log-concavity and unimodality of the cluster monomials of type $A_3$, a substantially more intricate case. Moreover, we refine and extend this conjecture by considering the unimodality and the strongly isomorphism of cluster algebras.

math.RT

Fractal phenomenon in $c$- and $g$-vectors of the Markov quiver

We study the $C$- and $G$-patterns associated with rank $3$ skew-symmetrizable matrices of $B$-invariant type, including the Markov quiver. Motivated by the self-contained simple mutations in Markov-type cluster algebras, we prove that large classes of subpatterns of modified $c$- and $g$-vectors are linearly isomorphic, yielding a fractal structure of the corresponding $G$-fan. We further derive explicit recursive formulas for all modified $c$- and $g$-vectors in terms of integer pairs satisfying a recursion analogous to the Calkin-Wilf tree, which leads to a parameterization by coprime integers. As an application, we describe all connected components of the complement of the support of the $G$-fan, and show that they are generated recursively by three kinds of linear maps.

math.RT

The Logarithmic Asymptotic Phenomenon for Generalized Markov-Hurwitz Equations

The purpose of this paper is twofold. First, we introduce a family of generalized Markov-Hurwitz equations, extending classical Markov-Hurwitz equations with additional degree n-1 interaction terms, Gyoda and Matsushita's generalized Markov equations from 3 variables to n variables. Second, we prove a logarithmic asymptotic phenomenon for the positive integer solutions of these equations.

math.NT

DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment

Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.

cs.LG

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training

Facial representation pre-training is crucial for tasks like facial recognition, expression analysis, and virtual reality. However, existing methods face three key challenges: (1) failing to capture distinct facial features and fine-grained semantics, (2) ignoring the spatial structure inherent to facial anatomy, and (3) inefficiently utilizing limited labeled data. To overcome these, we introduce PaCo-FR, an unsupervised framework that combines masked image modeling with patch-pixel alignment. Our approach integrates three innovative components: (1) a structured masking strategy that preserves spatial coherence by aligning with semantically meaningful facial regions, (2) a novel patch-based codebook that enhances feature discrimination with multiple candidate tokens, and (3) spatial consistency constraints that preserve geometric relationships between facial components. PaCo-FR achieves state-of-the-art performance across several facial analysis tasks with just 2 million unlabeled images for pre-training. Our method demonstrates significant improvements, particularly in scenarios with varying poses, occlusions, and lighting conditions. We believe this work advances facial representation learning and offers a scalable, efficient solution that reduces reliance on expensive annotated datasets, driving more effective facial analysis systems.

cs.CV

Towards Intrinsically Calibrated Uncertainty Quantification in Industrial Data-Driven Models via Diffusion Sampler

In modern process industries, data-driven models are important tools for real-time monitoring when key performance indicators are difficult to measure directly. While accurate predictions are essential, reliable uncertainty quantification (UQ) is equally critical for safety, reliability, and decision-making, but remains a major challenge in current data-driven approaches. In this work, we introduce a diffusion-based posterior sampling framework that inherently produces well-calibrated predictive uncertainty via faithful posterior sampling, eliminating the need for post-hoc calibration. In extensive evaluations on synthetic distributions, the Raman-based phenylacetic acid soft sensor benchmark, and a real ammonia synthesis case study, our method achieves practical improvements over existing UQ techniques in both uncertainty calibration and predictive accuracy. These results highlight diffusion samplers as a principled and scalable paradigm for advancing uncertainty-aware modeling in industrial applications.

cs.LG

Entire Space Counterfactual Learning for Reliable Content Recommendations

Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias. To handle both challenges, the entire space multitask models are employed to decompose the user behavior track into a sequence of exposure $\rightarrow$ click $\rightarrow$ conversion, constructing surrogate learning tasks for CVR estimation. However, these methods suffer from two significant defects: (1) intrinsic estimation bias (IEB), where the CVR estimates are higher than the actual values; (2) false independence prior (FIP), where the causal relationship between clicks and subsequent conversions is potentially overlooked. To overcome these limitations, we develop a model-agnostic framework, namely Entire Space Counterfactual Multitask Model (ESCM$^2$), which incorporates a counterfactual risk minimizer within the ESMM framework to regularize CVR estimation. Experiments conducted on large-scale industrial recommendation datasets and an online industrial recommendation service demonstrate that ESCM$^2$ effectively mitigates IEB and FIP defects and substantially enhances recommendation performance.

cs.LG

Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation

Heterogeneous treatment effect (HTE) estimation from observational data poses significant challenges due to treatment selection bias. Existing methods address this bias by minimizing distribution discrepancies between treatment groups in latent space, focusing on global alignment. However, the fruitful aspect of local proximity, where similar units exhibit similar outcomes, is often overlooked. In this study, we propose Proximity-enhanced CounterFactual Regression (CFR-Pro) to exploit proximity for enhancing representation balancing within the HTE estimation context. Specifically, we introduce a pair-wise proximity regularizer based on optimal transport to incorporate the local proximity in discrepancy calculation. However, the curse of dimensionality renders the proximity measure and discrepancy estimation ineffective -- exacerbated by limited data availability for HTE estimation. To handle this problem, we further develop an informative subspace projector, which trades off minimal distance precision for improved sample complexity. Extensive experiments demonstrate that CFR-Pro accurately matches units across different treatment groups, effectively mitigates treatment selection bias, and significantly outperforms competitors. Code is available at https://github.com/HowardZJU/CFR-Pro.

cs.LG

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

Recent advances in large language models (LLMs) have attracted significant interest in extending their capabilities to multimodal scenarios, particularly for speech-to-speech conversational systems. However, existing multimodal models handling interleaved audio and text rely on autoregressive (AR) methods, overlooking that text depends on target-target relations whereas audio depends mainly on source-target relations. In this work, we propose Text-to-Talk (TtT), a unified audio-text framework that integrates AR text generation with non-autoregressive (NAR) audio diffusion in a single Transformer. By leveraging the any-order AR property of absorbing discrete diffusion, our approach provides a unified training objective for text and audio. To support this hybrid generation paradigm, we design a modality-aware attention mechanism that enforces causal decoding for text while allowing bidirectional modeling within audio spans, and further introduce three training strategies that reduce train-test discrepancies. During inference, TtT employs block-wise diffusion to synthesize audio in parallel while flexibly handling variable-length outputs. Comprehensive experiments on Audio-QA, ASR, AAC and speech-to-speech benchmarks show that TtT consistently surpasses strong AR and NAR baselines, with additional ablation and training-strategy analyses confirming the contribution of each component. We will open-source our models, data and code to facilitate future research in this direction.

cs.CL