SearcharxivSearch

arXiv subjects

Yi Zheng

Publications and source records attributed to Yi Zheng.

At least 19 recordsLinked to original sources

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical challenges such as backtest overfitting, deep neural networks and Generative AI have introduced mechanical nondeterminism rooted in hardware and architecture. This survey provides a systems perspective on reproducibility failures across three modalities now dominant in financial AI: tabular models (post-hoc explanation variance), graph networks (stochastic sampling and temporal asynchrony), and LLM-based agentic workflows (batch-dependent divergence and trajectory drift). We supplement the literature analysis with first-party experiments on public financial datasets -- quantifying explanation rank instability in credit scoring, prediction flip rates in GNN-based fraud detection, and tensor-parallel-induced output divergence in LLM entity extraction. We propose a layered evaluation framework linking modality-specific metrics (RBO, D_cos, TDI, PSD) to audit readiness, and report where these measures overlap rather than complement one another.

cs.AI

Beyond Electrons: A Theoretical Framework for Near-Field Radiative Thermal Computing and Neural-Network-Inspired Processing

Near field radiative heat transfer provides a route for information processing in which thermal radiation, rather than charge transport, serves as the physical carrier of signals. Here, we propose and theoretically analyze a programmable near field radiative thermal computing framework in which radiative coupling, phase change nonlinearity, and thermal state memory are mapped onto neural network inspired operations. The framework is constructed from near field radiative thermal diodes, transistors, and multi terminal logic units separated by nanoscale gaps. Radiative heat flux represents the propagated thermal information, while geometry and material dependent radiative coupling provides physically constrained weighting, and the temperature dependent optical response of phase-change materials enables nonlinear modulation and logic state control. Based on these primitives, we formulate a radiative thermal convolutional network for spatial information processing and a radiative thermal recurrent network for history dependent computation. The recurrent response is associated with radiative feedback, thermal relaxation, and phase change hysteresis, with VO2 providing history dependent short term memory and GST offering a possible route toward non volatile phase storage. We further distinguish the physical radiative networks from a separate software based inverse identification study, in which recurrent machine-learning models are trained on simulated near field heat flux temperature characteristics to recover structural parameters. By establishing a bottom up connection between fluctuational electrodynamics, radiative thermal logic, programmable thermal states, and neural network inspired computation, this study provides a physically grounded basis for exploring non contact thermal information processing at the near field limit.

cond-mat.mes-hall

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.

cs.CV

Harnessing thermo-optic dynamics for frequency-agile soliton microcombs

Dissipative Kerr soliton microcombs enable compact and scalable frequency comb sources for precision metrology, spectroscopy, communications and coherent LiDAR, where broad and reliable frequency tuning is essential. Thermo-optic response can support thermal locking during soliton operation, enabling resonance tracking and thereby extending the tuning range, albeit modestly. However, it also induces pronounced thermal instability during soliton initiation, hindering reliable access to this extended operating regime and limiting practical deployment in applications requiring frequency agility. Here we show that strong mode coupling reshapes the effective detuning trajectory governing soliton formation, establishing a distinct operating regime in which thermo-optic response is significantly reinforced and constructively harnessed. In this regime, soliton formation proceeds without the thermal instability inherent to conventional operation, enabling robust soliton generation in material platforms previously limited by strong thermal effects. Importantly, the enhanced thermo-optic response strengthens thermal locking during soliton operation, enabling more effective resonance tracking and substantially extending the tuning range. Leveraging this regime in AlGaAs-on-insulator multimode microresonators, we demonstrate soliton generation with a tuning range approaching 100 GHz at a pump power of 32 mW. The same mechanism further enables frequency-agile operation through direct pump-frequency tuning without auxiliary stabilization, allowing massively parallel chirped comb generation with more than 90 channels exhibiting frequency excursions exceeding 10 GHz. These results establish a general operating principle for transforming thermo-optic effects from a limiting factor into an active resource, enabling robust and frequency-agile integrated soliton microcombs.

physics.optics

Fractals in rate-induced tipping

When parameters of a dynamical system change sufficiently fast, critical transitions can take place even in the absence of bifurcations. This phenomenon is known as rate-induced tipping and has been reported in a variety of systems, from simple ordinary differential equations and maps to mathematical models in climate sciences and ecology. In most examples, the transition happens at a critical rate of parameter change, a rate-induced tipping point, and is associated with a simple unstable orbit (edge state). In this work, we show how this simple picture changes when non-attracting fractal sets exist in the autonomous system, a ubiquitous situation in non-linear dynamics. We show that these fractals in phase space induce fractals in parameter space, which control the rates and parameter changes that result in tipping. We explain how such rate-induced fractals appear and how the fractal dimensions of the different sets are related to each other. We illustrate our general theory in three paradigmatic systems: a piecewise linear one-dimensional map, the two-dimensional Hénon map, and a forced pendulum.

nlin.CD

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latency and power limits. Due to an observed low-frequency shift in scene context and its relevant vocabulary, we present AdaDINO, an adaptive framework that makes on-device VFM inference efficient by matching execution to the current scene and task. We build on a known phenomenon, that the accuracy drop of shrinking model sizes depends on the task, and turn it into task-level adaptive execution. AdaDINO integrates neural architecture search (NAS) into a language-aligned VFM backbone distilled from DINOv2, training a single family of subnets for efficient execution during runtime. A multimodal large language model (LLM) on the cloud, invoked at low frequency, refines the candidate class set from scene context, while a learned selector activates the least-cost subnet predicted to retain a target fraction of accuracy. With the backbone and semantic pipeline held fixed, learned selection alone reduces average compute by $37\%$ over the best fixed subnet at equal segmentation accuracy. Across zero-shot classification and open-vocabulary segmentation, AdaDINO establishes a strong accuracy-efficiency frontier, improving over evaluated models of comparable sizes by up to $7.9\%$ in acc@1 on IN1K and $5.2\%$ mIoU on ADE20K, and reducing average FLOPs by up to $74.9\%$ at similar accuracy.

cs.CV

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony

Continuous emotional arousal quantification remains bottlenecked by time-consuming and labor-intensive manual annotation. This work investigates group-level EEG dynamic neural synchrony (DNS) as a principled signal for continuous arousal quantification that bypasses per-subject manual labeling. Using Correlated Component Analysis (CorrCA) with sliding-window computation across four EEG datasets spanning 142 subjects and over 207 hours, we systematically evaluate DNS as a group-level marker for emotional arousal dynamics. Three key findings emerge. First, DNS exhibits significant emotion information from valence-dependent differences (all p<0.003), with positive emotions eliciting higher synchrony. Second, DNS correlates more strongly with the first-order derivative of arousal than with raw arousal values, revealing that neural synchrony captures the rate of emotional change rather than static intensity. Third, we provide the first systematic characterization of how DNS-arousal coupling depends on key methodological choices, finding that moderate windows (10-30 s), positive lags (0-10 steps), and First-order Difference feature of EEG from the dominant CorrCA component yield consistently strong coupling. Subject-split replication and block permutation tests confirm these associations are not statistical artifacts. Our findings establish DNS as an empirically validated group-level marker toward annotation-efficient continuous emotional arousal quantification.

cs.HC

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the cost of efficiency and noisy cross-modal attention. To address these limitations, we propose CRISP, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context. CRISP works in a two-stage pipeline: Stage 1 first identifies text-aligned visual tokens, and Stage 2 enhances contextual completeness through semantic diversity. Extensive experiments on LLaVA-1.5 and LLaVA-NeXT demonstrate that CRISP achieves superior performance retention under aggressive pruning ratios, maintaining up to 99.5% accuracy while reducing inference cost and latency by more than 2 times. CRISP serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.

cs.CV

ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports

Reproduction tests help developers confirm reported issues and provide executable feedback for issue resolution, yet issue reports in open-source projects rarely include such tests. Recent studies have explored generating issue reproduction tests from issue reports with large language models, but existing approaches largely rely on prompt-based pipelines that retrieve textual context and generate tests. This limits their ability to understand how reported issues behave in repository-scale codebases and to flexibly organize the construction of reproduction tests. In this paper, we propose ReProAgent, a multi-stage agent framework for reproduction test generation from issue reports. ReProAgent decomposes the task into four agent stages: bug localization, root cause analysis, test planning, and test generation. To support these stages, ReProAgent integrates task-specific tools for task decomposition and reflection, context retrieval from both textual sources and repository graphs, and runtime interaction with the execution environment. Experiments on SWT-bench-lite and SWT-bench-verified show that ReProAgent successfully reproduces 58.43% and 70.30% of issues, outperforming all baselines, with an average cost of $0.14 per instance. For example, when equipped with GPT-5-mini, ReProAgent exceeds OpenHands with the same backbone by 20.43 and 7.90 percentage points, respectively. ReProAgent also generalizes across multiple backbone LLMs and improves downstream issue resolution performance when integrated with existing repair approaches.

cs.SE

Forecasting the E_G measurements from the photometric and spectroscopic surveys of Chinese Space Station Survey Telescope (CSST)

We present forecasts for the $E_G$ statistic using redshift distributions of realistic mock galaxy samples from the upcoming Chinese Space Station Survey Telescope (CSST). The dominant uncertainty in $E_G$ stems from the redshift space distortion parameter $β$, whose precision limits the overall constraining power. Our analysis shows that CSST will nevertheless achieve $E_G$ constraints at the few-percent level ($3\%-9\%$) over $0 < z < 1.2$, an improvement by a factor of several to an order of magnitude over current observations. Within the $μ-Σ$ modified gravity framework, the parameter $Σ_0$, associated with the effective gravitational constant of the Weyl potential, can be constrained to $\sim 5\%$ precision. In a plausible scenario where upcoming spectroscopic surveys determine $β$ to $1\%$ accuracy, $E_G$ constraints tighten to the percent level, and $Σ_0$ becomes measurable at $\sim 1\%$. For representative modified gravity scenarios, we find that the potential deviations from the Hu--Sawicki $f(R)$ model and the normal-branch Dvali--Gabadadze--Porrati (nDGP) model remain detectable within the expected sensitivity of CSST. These results demonstrate that CSST will serve as a powerful facility for testing gravity and underscore the essential synergy between photometric weak lensing and spectroscopic surveys in probing cosmic acceleration.

astro-ph.CO

A User-Friendly Python Interface for the Numerical Relativity Code AMSS-NCKU

Numerical relativity has brought about profound and wide-ranging influences on modern astrophysics and gravitational-wave astronomy. In this study, we present a user-friendly Python interface for the numerical relativity code AMSS-NCKU. This interface facilitates the automation of initializing and executing the AMSS-NCKU simulations, as well as the automatic visualization of the output data. The Python interface can significantly reduce the operational complexity of the AMSS-NCKU simulation workflow, lowering the technical barriers for new users. To show the utility of this Python interface, we present two representative examples of numerical relativity simulations (the binary black hole and triple black hole merger processes), obtaining stable numerical results and the expected physical behaviors for black hole systems. Keywords: Numerical Relativity, Gravitational Waves, Black Holes, Python

gr-qc

Synthesis of single-layered fluorographdiyne nanosheets via selective on-surface 2D covalent polymerization

Two-dimensional conjugated polymers (2DCPs) are significant macromolecular materials with intriguing and tunable physicochemical properties that depend on their geometries. Graphdiyne and its derivatives are exemplary 2DCPs featuring sp-sp2 hybridized skeletons. However, achieving single-layered, large-domain/regular graphdiyne and its derivatives on surfaces remains a formidable challenge due to the lack of selective 2D covalent polymerization methods. Here, we report a selective on-surface 2D covalent polymerization method via the combination of cobalt catalysis and coronene templating, achieving the synthesis of single-layered fluorographdiyne nanosheets up to 60*60 nm2 on Au(111) surface. Using scanning probe techniques, we visualize the sequential polymerization process and characterize cobalt-activated coupling intermediates at the atomic level. Experimental and theoretical analyses suggest that strong d-π coupling between cobalt and alkynyl transforms a robust Csp-Au bond into a weaker Csp2-Au bond, thereby facilitating the demetallization C-C coupling. Besides, the templating effect of coronene suppresses kinetically trapped defects and improves the selectivity of hexagonal-ring formation in the complex 2D covalent polymerization process.

cond-mat.mtrl-sci

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models

Supervised fine-tuning (SFT) followed by reinforcement learning (RL) has become a standard post-training paradigm for large language models. This paradigm provides a cold-start for RL exploration, avoiding the inefficiency of pure RL where on-policy sampling yields insufficient positive samples. However, in practice, existing approaches often use a small amount of data for SFT initialization compared to the RL phase, which can cause the model to fit the limited samples and shift away from its pre-trained distribution. This distribution shift impedes the model's ability to effectively explore during subsequent RL training. To address this challenge, we propose that in low-data regimes, SFT should prioritize activating task-relevant capabilities rather than memorizing specific content. Along this line, we propose EKSFT (Entropy-KL Selective Fine-Tuning), which selectively masks tokens that exhibit either high entropy or high KL divergence from a reference model. By excluding these high-uncertainty, distribution-shifting tokens from imitation, EKSFT injects task-specific knowledge while preserving the integrity of the model's pre-trained distribution. Empirical evaluations on mathematical reasoning benchmarks demonstrate that EKSFT consistently outperforms standard SFT. Further RL fine-tuning from the EKSFT model yields consistently better post-RL performance, indicating improved exploration for the RL stage. Our codes and datasets are available at https://github.com/MINE-USTC/EKSFT.

cs.AI

AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

Video production workflows offer a rich and demanding arena for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this end, we introduce AgenticVBench, a benchmark of 100 agentic tasks across 4 task families spanning the real world post-production workflow, constructed from real production workflows contributed by 20 industry experts averaging 6 years of professional experience. Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics. We evaluate frontier vision-language models (VLMs) with both vendor-native and open-source harnesses. The best evaluated agent stack barely crosses 30%, far below human expert performance on the same tasks. We further find that the choice of harness substantially affects model behavior, including scores, tool-use patterns, and failure modes. AgenticVBench provides a foundation for diagnosing and improving both models and harnesses for agentic video production. Benchmark website: https://agenticvbench.com.

cs.CR

Cross-Subject EEG Emotion Recognition Based on Temporal Asynchronous Alignment Contrastive Learning

With the advancement of science and technology, the importance of emotion research has become increasingly evident. Electroencephalography (EEG)-based emotion recognition has emerged as an active research area in recent years, owing to its objectivity and high temporal resolution. However, most existing methods focus on optimizing encoder structures to enhance feature extraction capabilities, while paying relatively little attention to similarity calculation strategies, particularly overlooking the potential temporal misalignment of responses among different subjects. To address these shortcomings, this paper draws inspiration from the late interaction mechanism of ColBERT in natural language processing (NLP) and proposes a Temporal Asynchronous Alignment-based Contrastive Learning (TA2CL) framework. This method transforms the traditional global "hard alignment" similarity calculation approach into a fine-grained local matching mechanism, enabling the model to adaptively search for and align "locally highly correlated" segments between two EEG signals, thereby effectively mitigating the effects of inter-subject differences and temporal delays. Experimental results demonstrate that the proposed method achieves strong performance across multiple public datasets. Specifically, on the FACED dataset, it achieves an accuracy of 64.5% for the nine-class classification task and 79.5% for the binary classification task, while on the SEED and SEED-V datasets, it achieves accuracies of 86.4% and 70.1%, respectively, validating the method's effectiveness and generalization capability.

cs.HC

Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling

Recent advances in foundation models have enabled conversational agents that aim for sustained companionship rather than mere task completion. Yet most still remain unable to support natural, long-term companion-like interactions, resulting in experiences that feel episodic and inauthentic. We argue that current agents overlooked cross-temporal modeling of agents' social behaviors and internal emotions: generated behaviors rarely influence an agent's emotional state, and emotional states seldom shape subsequent behaviors. We present Cross-Temporal Emotion Modeling (CTEM), a framework that links long-term behavioral history to moment-to-moment emotional expression. CTEM establishes a closed loop where past experiences update an evolving emotional state; this state conditions immediate interactions; and user feedback continually revises both memory and emotional state, enabling reflection and anticipation. We instantiate CTEM as Auri, a companion agent on an instant-messaging platform, and report a 21-day in-the-wild study showing that CTEM shows improvements in perceived naturalness, coherence, and emotional harmony.

cs.HC

Scalable Generation of Massive Schrödinger Cat States via Quantum Tunneling

Massive objects in spatial superposition may provide insights into the interplay between quantum mechanics and gravity. Cold atomic interferometers offer a promising platform due to extended matter-wave coherence times and precise controllability. However, high-mass spatial superpositions beyond single atoms have yet to be generated in such setups. Here, we report the scalable realization of high-mass spatial entanglement via quantum tunneling of ultracold atoms in optical lattices. We observe coherent tunneling of bound clusters, forming a composite object with a mass of 608~amu. Full control of the model parameters allows us to mitigate the usual suppression of tunneling with increasing mass. Furthermore, we construct an interferometer to certify the entanglement and use spatially distributed Schrödinger cat states to perform quantum-enhanced measurements. These results establish an approach to generating and detecting massive superposition states relevant to studies of quantum gravity.

cond-mat.quant-gas

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, effectively mitigating their inherent knowledge limitations. However, RAG remains vulnerable to poisoning attacks that manipulate retrieved texts to mislead model outputs. Existing defense mechanisms often lack theoretical robustness guarantees and perform unreliably when the LLM has limited knowledge of the retrieved content. In this work, we propose PRA-RAG, a provably robust retrieval aggregation algorithm designed to defend against poisoning attacks on retrieved texts. PRA-RAG samples multiple combinations of retrieved texts and utilizes geometric structures in the embedding space to identify a robust subset, from which a stable aggregated representation is derived. We provide theoretical bounds on the maximum impact of poisoned retrieved content and establish a quantitative measure of RAG's robustness. Experiments across multiple benchmarks and RAG architectures demonstrate that PRA-RAG reduces the attack success rate to as low as 1% while maintaining an accuracy of 71%, significantly outperforming representative state-of-the-art methods.

cs.IR