SearcharxivSearch

arXiv subjects

Zichao Chen

Publications and source records attributed to Zichao Chen.

8 recordsLinked to original sources

Decoupled domain-texture switching from magnetic easy axis in kagome ferromagnet EuTi3Bi4

Magnetic anisotropy defines the easy axis of a magnetic material and governs the spatial arrangement of its domains. To date, anisotropy engineering has focused on reorienting the easy axis or tuning the anisotropy energy, both of which demand substantial energy input. Here, we demonstrate that magnetic domain textures can be switched without reorienting the easy axis, as observed in a kagome ferromagnet EuTi3Bi4 crystal. Using low-temperature magnetic force microscopy, we observe that the preferred orientation of magnetic domains switches from the a-axis to the b-axis upon temperature variation, and that this switching can also be triggered by an out-of-plane magnetic-field reset. Magnetization measurements and density functional theory calculations confirm a robust c-axis easy magnetization, ruling out a conventional spin-reorientation transition. Instead, the texture switching is governed by the temperature dependence of the in-plane variation of the Magnetic anisotropy energy landscape, which arises from two competing interactions with different decay rates: single-ion anisotropy favors a-oriented spin components, while nearest-neighbor anisotropic exchange favors b-oriented ones. Furthermore, the critical switching temperature is substantially elevated in a mechanically exfoliated EuTi3Bi4 flake. Our findings establish that macroscopic magnetic textures can be effectively manipulated by tuning the competition between in-plane anisotropic interactions, without the energy cost of reorienting the easy axis.

cond-mat.mtrl-sci

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move and interact from a single image is still difficult for these systems. We present BilliardPhys-Bench, a benchmark for physical reasoning in synthetic billiards environments. Its procedural engine generates randomized scenarios with friction and elastic collisions. The benchmark tests three abilities: (1) predicting ball-to-ball collisions, (2) reasoning about wall bounces, and (3) estimating final ball positions after motion stops. We evaluate recent MLLMs from the GPT, Claude, Gemini, and Qwen families. Performance drops as simulation time increases and scene geometry grows more complex. We also observe a consistent failure mode we call "stasis bias": when the correct physical outcome is harder to infer, models tend to predict no interaction. These findings show where current MLLMs break down on visual dynamics and point toward the need for better physical inductive biases in multimodal architectures.

cs.AI

Spin-mediated hysteretic switching of unidirectional charge density waves by rotating magnetic fields

Charge density waves (CDWs) are a widespread collective electronic order in quantum materials, furnishing key insights into symmetry breaking and competing phases. However, their dynamic control with external fields remains a pivotal challenge. Here, we report deterministic and hysteretic switching of unidirectional CDW orientation via in-plane magnetic field rotation in magnetic kagome metal GdTi3Bi4. Atomically resolved spectroscopy shows two types of 3a0*1a0 CDW domains, Q1 and Q2 oriented 60 degree apart along two distinct crystallographic directions and separated by atomically sharp domain walls. Rotating the magnetic field drives reversible transitions between these CDW configurations, exhibiting a robust C2-symmetric phase diagram with pronounced hysteresis. This hysteretic switching is mediated by a field-dependent reorientation of underlying antiferromagnetic spins, revealing a tunable energy landscape with stable and metastable states and modulates the electronic charge order via spin-lattice coupling. Our findings not only demonstrate the switching of CDW configurations by in-plane magnetic field but also reveal the mechanism of coupling between CDW and magnetic fields, offering new insights into CDW manipulation and versatile platform for developing a spin-mediated multistate spin-charge coupling memory and programmable quantum devices.

cond-mat.str-el

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. They do not assess whether models can reason over the global structural properties of formal diagrams, such as topology, conservation constraints, and the consistent mapping between visual patterns and algebraic expressions. We introduce FeynmanBench, a benchmark of over 2,000 tasks centered on Feynman diagrams spanning the electromagnetic, weak, and strong interactions of the Standard Model. Each instance couples a diagram image with minimal textual conventions and requires models to recover the full physical content -- vertex inventory, propagator types, topological connectivity, momentum routing, and the complete scattering amplitude. An automated generation and verification pipeline produces the diagrams, annotations, and reference answers under standardized rules. Evaluating 19 state-of-the-art multimodal LLMs, we find a consistent failure pattern: models achieve 70--95\% on local recognition (vertex and propagator identification) but collapse to 13--17\% on topological reconstruction (CP3), and near zero on full algebraic derivation (CP5). FeynmanBench offers a controlled testbed for multimodal reasoning over formal scientific diagrams and highlights fundamental limitations of current architectures in topology-sensitive scientific reasoning.

cs.AI

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchmarks due to data contamination, insufficient complexity, and prohibitive human labor costs. Here we present SPM-Bench, an original, PhD-level multimodal benchmark specifically designed for scanning probe microscopy (SPM). We propose a fully automated data synthesis pipeline that ensures both high authority and low-cost. By employing Anchor-Gated Sieve (AGS) technology, we efficiently extract high-value image-text pairs from arXiv and journal papers published between 2023 and 2025. Through a hybrid cloud-local architecture where VLMs return only spatial coordinates "llbox" for local high-fidelity cropping, our pipeline achieves extreme token savings while maintaining high dataset purity. To accurately and objectively evaluate the performance of the LLMs, we introduce the Strict Imperfection Penalty F1 (SIP-F1) score. This metric not only establishes a rigorous capability hierarchy but also, for the first time, quantifies model "personalities" (Conservative, Aggressive, Gambler, or Wise). By correlating these results with model-reported confidence and perceived difficulty, we expose the true reasoning boundaries of current AI in complex physical scenarios. These insights establish SPM-Bench as a generalizable paradigm for automated scientific data synthesis.

cs.AI

HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified

cs.CL

Blockchain Network Analysis: A Comparative Study of Decentralized Banks

Decentralized finance (DeFi) is known for its unique mechanism design, which applies smart contracts to facilitate peer-to-peer transactions. The decentralized bank is a typical DeFi application. Ideally, a decentralized bank should be decentralized in the transaction. However, many recent studies have found that decentralized banks have not achieved a significant degree of decentralization. This research conducts a comparative study among mainstream decentralized banks. We apply core-periphery network features analysis using the transaction data from four decentralized banks, Liquity, Aave, MakerDao, and Compound. We extract six features and compare the banks' levels of decentralization cross-sectionally. According to the analysis results, we find that: 1) MakerDao and Compound are more decentralized in the transactions than Aave and Liquity. 2) Although decentralized banking transactions are supposed to be decentralized, the data show that four banks have primary external transaction core addresses such as Huobi, Coinbase, and Binance, etc. We also discuss four design features that might affect network decentralization. Our research contributes to the literature at the interface of decentralized finance, financial technology (Fintech), and social network analysis and inspires future protocol designs to live up to the promise of decentralized finance for a truly peer-to-peer transaction network.

econ.GN

Visualizing Non-Fungible Token Ethics: A Case Study On CryptoPunks

As a blockchain-based application, Non-Fungible Token (NFT) has received worldwide attention over the past few years. Digital artwork is the main form of NFT that can be stored on different blockchains. Although the NFT market is rapidly developing, we observed potential ethical and racial fairness issues in the design of NFT artworks due to a lack of ethical guidelines or censorship. Therefore, we investigated CryptoPunks, the most famous collection in the NFT market, to explore and visualize its potential ethical issues. We explored the ethical issues from three aspects: design, trading transactions, and related topics on Twitter. We scraped data from Twitter and Dune Analytics using python libraries, Twitter crawler, and sentiment analysis tools. Our five visualizations implied that 1.6 times more male punks were created in the initial design process than the female ones. And the male ones have a higher average selling price than females; lighter-skinned punks tend to sell for higher prices. The results of our study and visualizations provide a preliminary exploration of CryptoPunks and further inspire future ethical-related investigation and research in the NFT domain.

cs.HC