SearcharxivSearch

arXiv subjects

Enze Zhang

Publications and source records attributed to Enze Zhang.

At least 19 recordsLinked to original sources

A Compact Selective State-Space Model for Cross-Sectional Stock Return Ranking from Raw Intraday Bars

We present STRATA (Staggered-Timescale Residual Architecture), a 244,633-parameter sequence model that maps five trading days of raw five-minute bar and order-book data directly to a next-day cross-sectional return ranking, with no hand-crafted features. The raw-input setting has a structural obstacle: price series are non-stationary and differ across stocks by orders of magnitude, so a model easily latches onto price level rather than dynamics. STRATA addresses it with a stem of five branches--four learnable causal depthwise convolutions whose effective kernels are initialised to sum to zero, plus one cross-field linear contrast--followed by four selective state-space blocks whose decay biases are staggered across the stack and a four-path readout. Because a score that merely tilts toward common style factors scores well on raw rank correlations, every model's scores are residualised against eight price-volume style factors before any metric is computed. Trained on four years of data covering roughly one thousand mid-capitalisation Chinese A-shares and evaluated once on a held-out year, STRATA reaches a style-residualised rank information coefficient of 0.0728 (information ratio 1.128, signal long-short Sharpe 12.85), ahead of six parameter-matched sequence baselines on all four reported metrics; on rank IC the day-level paired gap against every baseline is significant at p < 0.001, and among the arms competitive on predictive power STRATA's scores are the least explained by the controls. The close-to-close target opens before the score exists: measured instead from the first executable price, the decile spread is indistinguishable from zero, while the ordering of the seven architectures is unchanged and STRATA's margin widens.

cs.CE

ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.

cs.CL

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural signals mainly with abstract text, a lossy translation that may discard fine-grained perceptual information encoded in brain activity. We propose Generative Visual Grounding (GVG), a framework that visualizes the invisible by using an EEG-to-image generative model as a visual translator. Instead of forcing EEG into text alone, GVG hallucinates instance-specific proxy images for non-visual EEG, providing structured visual contexts that allow MLLMs to exploit their visual priors for clinical-state interpretation. We validate this idea on two MLLM backbones, GVG-X-Omni and GVG-Janus. Image-only alignment is already competitive: the lightweight GVG-X-Omni matches 1.7B-parameter text-aligned baselines while tuning only 170M parameters on a frozen 7B backbone. We further extend GVG-Janus with trimodal Image+Text alignment, where text supplies categorical semantic anchors and visual proxies enrich neural representations with perceptual details. Experiments show consistent gains in EEG understanding and visual generation, suggesting visual proxy grounding as an effective complement to textual alignment.

cs.AI

HarmThoughts: A Benchmark for Fine-Grained Harmful Behavior Detection in Reasoning Traces

Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm emerges during reasoning. When jailbroken, harm does not appear instantaneously but unfolds through distinct behavioral steps such as suppressing refusal, rationalizing compliance, decomposing harmful tasks, and concealing risk. However, no existing benchmark captures this process at sentence-level granularity within reasoning traces -- a key step toward reliable safety monitoring, interventions, and systematic failure diagnosis. To address this gap, we introduce HarmThoughts, a benchmark for step-wise safety evaluation of reasoning traces. HarmThoughts is built around our proposed harm taxonomy, comprising 16 functional reasoning behavior categories that capture how reasoning steps contribute to or mitigate harmful outcomes. The dataset consists of 56,931 sentences from 1,018 reasoning traces generated by four model families, each annotated with fine-grained sentence-level behavioral labels. Using HarmThoughts, we analyze harm propagation by composing taxonomy behaviors into safety-failure patterns that characterize how reasoning transitions into harmful execution. We compare white-box and black-box monitors for identifying fine-grained taxonomy behaviors, and further evaluate supervised fine-tuning. While off-the-shelf monitors degrade sharply as behavioral granularity increases, fine-tuning substantially improves performance, highlighting both the difficulty and learnability of fine-grained process-level safety monitoring.

cs.CL

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, period) from visual input remains underexplored. We introduce a multi-category, cross-cultural benchmark for this task and evaluate VLMs using an LLM-as-Judge framework that measures semantic alignment with reference annotations. To assess cultural reasoning, we report exact-match, partial-match, and attribute-level accuracy across cultural regions. Results show that models capture fragmented signals and exhibit substantial performance variation across cultures and metadata types, leading to inconsistent and weakly grounded predictions. These findings highlight the limitations of current VLMs in structured cultural metadata inference beyond visual perception.

cs.CV

State-Action Inpainting Diffuser for Continuous Control with Delay

Signal delay poses a fundamental challenge in continuous control and reinforcement learning (RL) by introducing a temporal gap between interaction and perception. Current solutions have largely evolved along two distinct paradigms: model-free approaches which utilize state augmentation to preserve Markovian properties, and model-based methods which focus on inferring latent beliefs via dynamics modeling. In this paper, we bridge these perspectives by introducing State-Action Inpainting Diffuser (SAID), a framework that integrates the inductive bias of dynamics learning with the direct decision-making capability of policy optimization. By formulating the problem as a joint sequence inpainting task, SAID implicitly captures environmental dynamics while directly generating consistent plans, effectively operating at the intersection of model-based and model-free paradigms. Crucially, this generative formulation allows SAID to be seamlessly applied to both online and offline RL. Extensive experiments on delayed continuous control benchmarks demonstrate that SAID achieves state-of-the-art and robust performance. Our study suggests a new methodology to advance the field of RL with delay.

cs.AI

MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy

Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into interpretations, decisions, or claims whose specificity, certainty, severity, and actionability remain warranted. Existing benchmarks mainly evaluate specific clinical roles or final answers, leaving the reliability of explicit reasoning trajectories under-specified. We introduce MiraMind, a unified benchmark spanning six task families and 13 datasets across appraisal, diagnosis, intervention, abstraction, and verification. MiraMind evaluates both task outcomes and the trajectories connecting evidence to judgment through usability, logical structure, and informational contribution. Evaluating 20 LLMs reveals a restraint gap, in which the specificity or certainty of model judgments exceeds what limited evidence supports. We further train Mindora, an 8B model that targets evidence-to-judgment transitions through hard-case supervision, structured trajectory rewriting, and consistency-aware optimization. Mindora achieves the best average rank on MiraMind, improves over its backbone across all six task families, and produces more balanced reasoning trajectories. These results show that MiraMind can expose shared restraint failures and evaluate whether targeted post-training improves evidence-constrained mental-health reasoning.

cs.CL

DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation

Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on surface-level metrics that fail to capture the distinctive traits of this genre. To address these gaps, we introduce DITING, the first comprehensive evaluation framework for web novel translation, assessing narrative and cultural fidelity across six dimensions: idiom translation, lexical ambiguity, terminology localization, tense consistency, zero-pronoun resolution, and cultural safety, supported by over 18K expert-annotated Chinese-English sentence pairs. We further propose AgentEval, a reasoning-driven multi-agent evaluation framework that simulates expert deliberation to assess translation quality beyond lexical overlap, achieving the highest correlation with human judgments among seven tested automatic metrics. To enable metric comparison, we develop MetricAlign, a meta-evaluation dataset of 300 sentence pairs annotated with error labels and scalar quality scores. Comprehensive evaluation of fourteen open, closed, and commercial models reveals that Chinese-trained LLMs surpass larger foreign counterparts, and that DeepSeek-V3 delivers the most faithful and stylistically coherent translations. Our work establishes a new paradigm for exploring LLM-based web novel translation and provides public resources to advance future research.

cs.CL

Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation

Offline-to-online Reinforcement Learning (O2O RL) aims to perform online fine-tuning on an offline pre-trained policy to minimize costly online interactions. Existing work used offline datasets to generate data that conform to the online data distribution for data augmentation. However, generated data still exhibits a gap with the online data, limiting overall performance. To address this, we propose a new data augmentation approach, Classifier-Free Diffusion Generation (CFDG). Without introducing additional classifier training overhead, CFDG leverages classifier-free guidance diffusion to significantly enhance the generation quality of offline and online data with different distributions. Additionally, it employs a reweighting method to enable more generated data to align with the online data, enhancing performance while maintaining the agent's stability. Experimental results show that CFDG outperforms replaying the two data types or using a standard diffusion model to generate new data. Our method is versatile and can be integrated with existing offline-to-online RL algorithms. By implementing CFDG to popular methods IQL, PEX and APL, we achieve a notable 15% average improvement in empirical performance on the D4RL benchmark such as MuJoCo and AntMaze.

cs.LG

Room-temperature nonlinear transport and microwave rectification in antiferromagnetic MnBi$_2$Te$_4$ films

The discovery of the nonlinear Hall effect provides an avenue for studying the interplay among symmetry, topology, and phase transitions, with potential applications in signal doubling and high-frequency rectification. However, practical applications require devices fabricated on large area thin film as well as room-temperature operation. Here, we demonstrate robust room-temperature nonlinear transverse response and microwave rectification in MnBi$_2$Te$_4$ films grown by molecular beam epitaxy. We observe multiple sign-reversals in the nonlinear response by tuning the chemical potential. Through theoretical analysis, we identify skew scattering and side jump, arising from extrinsic spin-orbit scattering, as the main mechanisms underlying the observed nonlinear signals. Furthermore, we demonstrate radio frequency (RF) rectification in the range of 1-8 gigahertz at 300 K. These findings not only enhance our understanding of the relationship between nonlinear response and magnetism, but also expand the potential applications as energy harvesters and detectors in high-frequency scenarios.

cond-mat.mtrl-sci

Room-temperature van der Waals 2D ferromagnet switching by spin-orbit torques

Emerging wide varieties of the two-dimensional (2D) van der Waals (vdW) magnets with atomically thin and smooth interfaces holds great promise for next-generation spintronic devices. However, due to the lower Curie temperature of the vdW 2D ferromagnets than room temperature, electrically manipulating its magnetization at room temperature has not been realized. In this work, we demonstrate the perpendicular magnetization of 2D vdW ferromagnet Fe3GaTe2 can be effectively switched at room temperature in Fe3GaTe2/Pt bilayer by spin-orbit torques (SOTs) with a relatively low current density of 1.3 10^7A/cm2. Moreover, the high SOT efficiency of ξ_{DL}~0.22 is quantitatively determined by harmonic measurements, which is higher than those in Pt-based heavy metal/conventional ferromagnet devices. Our findings of room-temperature vdW 2D ferromagnet switching by SOTs provide a significant basis for the development of vdW-ferromagnet-based spintronic applications.

physics.app-ph

All-electrical switching of a topological non-collinear antiferromagnet at room temperature

Non-collinear antiferromagnetic Weyl semimetals, combining the advantages of a zero stray field and ultrafast spin dynamics as well as a large anomalous Hall effect and the chiral anomaly of Weyl fermions, have attracted extensive interests. However, the all-electrical control of such systems at room temperature, a crucial step toward practical applications, has not been reported. Here using a small writing current of around 5*10^{6} A/cm^{2}, we realize the all-electrical current-induced deterministic switching of the non-collinear antiferromagnet Mn3Sn with a strong readout signal at room temperature in the Si/SiO2/Mn3Sn/AlOx structure, without external magnetic field and injected spin current. Our simulations reveal that the switching is originated from the current-induced intrinsic non-collinear spin-orbit torques in Mn3Sn itself. Our findings pave the way for the development of topological antiferromagnetic spintronics.

cond-mat.mes-hall

Gate-tunable Intrinsic Anomalous Hall Effect in Epitaxial MnBi2Te4 Films

Anomalous Hall effect (AHE) is an important transport signature revealing topological properties of magnetic materials and their spin textures. Recently, antiferromagnetic MnBi2Te4 has been demonstrated to be an intrinsic magnetic topological insulator that exhibits quantum AHE in exfoliated nanoflakes. However, its complicated AHE behaviors may offer an opportunity for the unexplored correlation between magnetism and band structure. Here, we show the Berry curvature dominated intrinsic AHE in wafer-scale MnBi2Te4 thin films. By utilizing a high-dielectric SrTiO3 as the back-gate, we unveil an ambipolar conduction and electron-hole carrier (n-p) transition in ~7 septuple layer MnBi2Te4. A quadratic relation between the saturated AHE resistance and longitudinal resistance suggests its intrinsic AHE mechanism. For ~3 septuple layer MnBi2Te4, however, the AHE reverses its sign from pristine negative to positive under the electric-gating. The first-principles calculations demonstrate that such behavior is due to the competing Berry curvature between polarized spin-minority-dominated surface states and spin-majority-dominated inner-bands. Our results shed light on the physical mechanism of the gate-tunable intrinsic AHE in MnBi2Te4 thin films and provide a feasible approach to engineering its AHE.

cond-mat.mtrl-sci

Van der Waals Ferromagnetic Josephson Junctions

Superconductor-ferromagnet (S-F) interfaces in two-dimensional (2D) heterostructures present a unique opportunity to study the interplay between superconductivity and ferromagnetism. The realization of such nanoscale heterostructures in van der Waals (vdW) crystals remains largely unexplored due to the challenge of making an atomically-sharp interface from their layered structures. Here, we build a vdW ferromagnetic Josephson junction (JJ) by inserting a few-layer ferromagnetic insulator Cr2Ge2Te6 into two layers of superconductor NbSe2. Owing to the remanent magnetic moment of the barrier, the critical current and the corresponding junction resistance exhibit a hysteretic and oscillatory behavior against in-plane magnetic fields, manifesting itself as a strong Josephson coupling state. Through the control of this hysteresis, we can effectively trace the magnetic properties of atomic Cr2Ge2Te6 in response to the external magnetic field. Also, we observe a central minimum of critical current in some thick JJ devices, evidencing the coexistence of 0 and π phase coupling in the junction region. Our study paves the way to exploring the sensitive probes of weak magnetism and multifunctional building blocks for phase-related superconducting circuits with the use of vdW heterostructures.

cond-mat.mes-hall

The Discovery of Tunable Universality Class in Superconducting $β$-W Thin Films

The interplay between quenched disorder and critical behavior in quantum phase transitions is conceptually fascinating and of fundamental importance for understanding phase transitions. However, it is still unclear whether or not the quenched disorder influences the universality class of quantum phase transitions. More crucially, the absence of superconducting-metal transitions under in-plane magnetic fields in 2D superconductors imposes constraints on the universality of quantum criticality. Here, we discover the tunable universality class of superconductor-metal transition by changing the disorder strength in $β$-W films with varying thickness. The finite-size scaling uncovers the switch of universality class: quantum Griffiths singularity to multiple quantum criticality at a critical thickness of $t_{c \perp 1}\sim 8 nm$ and then from multiple quantum criticality to single criticality at $t_{c\perp 2}\sim 16 nm$. Moreover, the superconducting-metal transition is observed for the first time under in-plane magnetic fields and the universality class is changed at $t_{c \parallel }\sim 8 nm$. The discovery of tunable universality class under both out-of-plane and in-plane magnetic fields provides broad information for the disorder effect on superconducting-metal transitions and quantum criticality.

cond-mat.supr-con

Edge superconductivity in Multilayer WTe2 Josephson junction

WTe2, as a type-II Weyl semimetal, has 2D Fermi arcs on the (001) surface in the bulk and 1D helical edge states in its monolayer. These features have recently attracted wide attention in condensed matter physics. However, in the intermediate regime between the bulk and monolayer, the edge states have not been resolved owing to its closed band gap which makes the bulk states dominant. Here, we report the signatures of the edge superconductivity by superconducting quantum interference measurements in multilayer WTe2 Josephson junctions and we directly map the localized supercurrent. In thick WTe2 (~60 nm), the supercurrent is uniformly distributed by bulk states with symmetric Josephson effect ($\left|I_c^+(B)\right|=\left|I_c^-(B)\right|$). In thin WTe2 (10 nm), however, the supercurrent becomes confined to the edge and its width reaches up to 1.4 um and exhibits non-symmetric behavior $\left|I_c^+(B)\right|\neq \left|I_c^-(B)\right|$. The ability to tune the edge domination by changing thickness and the edge superconductivity establishes WTe2 as a promising topological system with exotic quantum phases and a rich physics.

cond-mat.supr-con

SegET: Deep Neural Network with Rich Contextual Features for Cellular Structures Segmentation in Electron Tomography Image

Electron tomography (ET) allows high-resolution reconstructions of macromolecular complexes at nearnative state. Cellular structures segmentation in the reconstruction data from electron tomographic images is often required for analyzing and visualizing biological structures, making it a powerful tool for quantitative descriptions of whole cell structures and understanding biological functions. However, these cellular structures are rather difficult to automatically separate or quantify from view owing to complex molecular environment and the limitations of reconstruction data of ET. In this paper, we propose a single end-to-end deep fully-convolutional semantic segmentation network dubbed SegET with rich contextual features which fully exploitsthe multi-scale and multi-level contextual information and reduces the loss of details of cellular structures in ET images. We trained and evaluated our network on the electron tomogram of the CTL Immunological Synapse from Cell Image library. Our results demonstrate that SegET can automatically segment accurately and outperform all other baseline methods on each individual structure in our ET dataset.

cs.CV

Evolution of Weyl orbit and quantum Hall effect in Dirac semimetal Cd3As2

Owing to the coupling between open Fermi arcs on opposite surfaces, topological Dirac semimetals exhibit a new type of cyclotron orbit in the surface states known as Weyl orbit. Here, by lowering the carrier density in Cd3As2 nanoplates, we observe a crossover from multiple- to single-frequency Shubnikov-de Haas (SdH) oscillations when subjected to out-of-plane magnetic field, indicating the dominant role of surface transport. With the increase of magnetic field, the SdH oscillations further develop into quantum Hall state with non-vanishing longitudinal resistance. By tracking the oscillation frequency and Hall plateau, we observe a Zeeman-related splitting and extract the Landau level index as well as sub-band number. Different from conventional two-dimensional systems, this unique quantum Hall effect may be related to the quantized version of Weyl orbits. Our results call for further investigations into the exotic quantum Hall states in the low-dimensional structure of topological semimetals.

cond-mat.mtrl-sci