SearcharxivSearch

arXiv subjects

Wenbin Wu

Publications and source records attributed to Wenbin Wu.

At least 19 recordsLinked to original sources

Giant-exchange-driven Vectorial Control of a Minimal Topological Magnet in Eu3In2As4

The interplay between magnetism and band topology provides a route to controlling quantum states of matter, yet its realization in materials is often constrained by weak exchange coupling and complex electronic structures. Here, a giant exchange coupling is identified in the newly predicted topological magnet Eu3In2As4, giving rise to magnetization-dependent band shifts of up to 300 meV. Together with its intrinsically soft magnetic response, this strong cou-pling enables systematic tuning of topological phases by both the magnitude and orientation of applied magnetic fields. The magneto-topological phase diagram is mapped out in which an antiferromagnetic topological insulator ground state evolves, under modest fields, into a pro-posed intermediate 2/3-ferrimagnetic phase, and further into fully polarized ferromagnetic states predicted to host either Weyl or nodal-ring semimetals. Notably, the Weyl phase corresponds to a minimal model hosting a single pair of Weyl nodes. Quantum oscillations, anomalous Hall transport and magneto-infrared spectroscopy consistently reveal exchange-driven band recon-struction across these transitions. Rotation of the magnetization theoretically provides an effi-cient means to tune the momentum-space positions and separations of the Weyl nodes. These results establish Eu3In2As4 as a model system for exploring how strong exchange coupling can be used to control topological band structures with minimal complexity.

cond-mat.mtrl-sci

SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not an explicit constraint-search additionally requires the rewrite to remain faithful to the user's stated query intent. Transplanted directly, these models learn a shortcut we term the generic-word dominance effect: they favor generic rewrites that score well on paths but drift from query intent. To address this, we propose SPEAR (Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval), which integrates three components that each target one failure mode: (1) a dual-embedding backbone with auxiliary loss and gradient isolation that shields recall-side semantics from being eroded by CTR-driven ranking signals; (2) a multiplicative gating aggregator that lets a rewrite score high only when both its confidence and item relevance are strong, eliminating the generic-word shortcut; (3) a Dynamic Rewrite Selector that jointly generates request-specific rewrite weights and user-query-conditioned scale and bias terms, allowing both rewrite preference and relevance calibration to adapt to each request. Offline evaluation on 100K held-out industrial search sessions shows that the proposed framework improves rewrite semantic similarity@10 by +18.2 and click recall@10 by +99.5 over the production baseline. In online A/B testing, SPEAR achieves +0.259 in query-view CTR and +0.733 in average reading depth, confirming that improved rewrite selection translates into stronger retrieval and deeper user engagement. The proposed SPEAR system has been fully deployed in Dewu's community search platform since 2025. Our code is available at https://github.com/mallocagi1-cell/spear.

cs.IR

In-Situ Polarimetry in Collimated Magneto-Infrared Spectroscopy System

Magneto-infrared spectroscopy under strong magnetic fields provides a powerful probe of Landau quantization and field-induced collective excitations, yet its full potential has long been constrained by the lack of in-situ polarization control, because the highly divergent infrared beam propagating through narrow light tubes undergoes multiple wall reflections, leading to severe polarization degradation. Here we report a collimated magneto-infrared spectroscopy system that integrates continuous in-situ polarimetry. The system employs incident and exit collimation chambers forming a Kepler type optical architecture, which converts the large-aperture FTIR output into a low-divergence beam and strongly suppresses multi-reflection trajectories inside long gold-plated light tubes, thereby enhancing both optical throughput and polarization fidelity. A remotely controlled polarization module, consisting of an automated linear polarizer and a switchable Fresnel rhomb positioned entirely outside the high-field region, enables continuous in-situ tuning between linear, circular, and arbitrary elliptical polarization states without thermal cycling, manual realignment, or breaking vacuum. Interchangeable compact focusing modules further support Faraday and Voigt geometries in both transmission and reflection experiments within a 50 mm magnet bore, providing efficient beam focusing and signal collection while maintaining polarization fidelity. The setup achieves a minimum root-mean-square noise of 0.0033%, an average noise of 0.0082%, and a linear polarization extinction ratio up to 40:1. We demonstrate the capability through continuous in-situ linear polarimetry and broadband circular polarimetry in the magneto-infrared spectroscopy of various single crystals. This platform establishes a robust experimental framework for in-situ polarization-resolved magneto-infrared spectroscopy.

physics.ins-det

Giant and Broadband Circular Dichroism from Particle-Hole Symmetry Breaking in Weyl Semimetals

Circular dichroism originates from symmetry breaking of material structure, leading to differential absorption of left- and right-circularly polarized light. However, circular dichroism in most materials is inherently weak and spectrally narrow, especially in the mid-to-far infrared. Here, we uncover giant infrared circular dichroism in the magnetic-field-forced Weyl semimetal Mn(Bi,Sb)2Te4, driven by extreme particle-hole symmetry breaking. Helicity-resolved magneto-infrared spectroscopy reveals circular dichroism exceeding 3000 mdeg (~130 mdeg/nm) with above-degree response extending over the 6-13 {\mu}m spectral range. The optical resonances are enhanced by a strong band nesting effect intrinsic to the Landau levels of type-II Weyl dispersion. A symmetry-based kp model reproduces these magneto-infrared responses and demonstrates that magnetization-induced asymmetric spin-orbit coupling generates particle-hole symmetry breaking, suppressing spin-up, parity-even wavefunction components in the valence Landau band and thereby producing pronounced optical helicity selectivity. Our findings establish particle-hole symmetry breaking as an effective route toward helicity-resolved optical control in quantum materials.

cond-mat.mtrl-sci

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no regulator-aligned way to measure the resulting false alarms. We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model, forecasts future exposure networks; (2) deterministic monitors and stress scenarios then turn those forecasts into typed alerts, attribution signals, and scenario evidence; and (3) data-health and confidence gates constrain escalation before DeXposure-Claw emits auditable supervisory tickets with rationales. We further develop DeXposure-Bench, a six-axis evaluation harness, whose decision axis scores tickets against a regulator-aligned absolute-loss ground truth and an explicit false-intervention rate. Experiments on five years of weekly real data fully support our system. Code is at https://github.com/EVIEHub/DeXposure-Claw.

cs.AI

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply it to Bitcoin. First, a behavioral audit of nine frontier LLMs shows that Bitcoin's ranking among money-like instruments is frame-dependent: models place it around rank 5 of 8 as "reliable money" but near the top under crisis and autonomous-agent frames, and an attribute-swap experiment shows that rankings track functional properties, not names. Second, we open a model's internals: a search across thousands of sparse-autoencoder features in Gemma 3 identifies a dominant Bitcoin-selective feature. Amplifying it shifts the model toward the asset and suppressing it shifts the model away, even when "Bitcoin" never appears in the prompt. Third, we test financial consequences: amplification raises Bitcoin's portfolio share by 5.2 percentage points while suppression lowers it by 4.6 pp, with amplification reallocating within crypto and suppression cutting total crypto exposure. We characterize this as bounded behavioral leverage (leverage meaning causal influence over outputs, not financial leverage): an identifiable internal feature can be perturbed to move financial choices, but only within measurable limits. The framework links internal representations to external recommendations, validated with random controls and mechanism boundaries. As LLMs become autonomous financial agents, this is a first step toward a behavioral layer for emerging know-your-agent (KYA) standards: knowing what an agent prefers, and how far that preference can be moved.

q-fin.GN

Excitations across the equilibrium and photoinduced `hidden' states of magnetoresistive manganites

"Hidden" phases, generated using ultrafast laser pulses (few hundred femtoseconds), with properties distinct from thermodynamic equilibrium, are appealing for technologies because they can be long-lived, with lifetimes of hours or weeks, and reversible with temperature sweeping or extra pulses. In this regard, La$_{2/3}$Ca$_{1/3}$MnO$_3$ (LCMO) stands out due to its tunability through epitaxial strain, which can drive the bulk ferromagnetic metal (FMM) into an antiferromagnetic insulator (AFI), and its susceptibility to photo-induced transitions. Indeed, AFI LCMO displays a long-lived photo-induced transition into a putative 'hidden' phase whose exact nature and excitations are still largely unknown. Here, we combine ultrafast photo-excitation in the near infrared with in situ transport, x-ray absorption (XAS), and Resonant Inelastic X-ray Scattering (RIXS) to investigate the excitations (polarons, phonons, and orbital) of the photo-excited phase of LCMO and contrast them with the thermodynamic phases achieved through strain and temperature. In the thermodynamic regime, we establish the correlation between polarons and transport, placing them in the 'strong coupling' regime of the Holstein model. Upon photo-excitation of LCMO-AFI, we uncover a long-lived phase characterized by the softening of the polaron excitations, the partial suppression of the Jahn-Teller distortion, and nearly unchanged phonons, showing the emergence of a photo-excited state absent in the equilibrium phase diagram. Finally, by varying temperature, epitaxial strain, and photo-excitation fluence, we construct a polaron phase diagram and identify the key spectroscopic signatures of each phase. Our laser-RIXS approach establishes a versatile platform for exploring photo-induced 'hidden' phases in quantum materials in non-stroboscopic conditions.

cond-mat.str-el

Aligning to Illusions: Choice Blindness in Human and AI Feedback

Reinforcement Learning from Human Feedback (RLHF) assumes annotator preferences reflect stable internal states. We challenge this through three experiments spanning the preference pipeline. In a human choice blindness study, 91% of surreptitiously swapped preferences go undetected, extending choice blindness to third-person evaluative comparison of unfamiliar text. Testing fifteen LLM judges as potential replacements, we find detection relies on shallow text matching rather than genuine self-monitoring: removing prior reasoning from context causes blindness to surge from near-zero to over 50%, while explicit social pressure induces near-universal compliance. In a dose-response experiment across two architectures from 86M to 2B parameters, one-sixth to one-third of labels must be corrupted before the reward signal halves, yet standard pairwise accuracy remains virtually unchanged. A Best-of-N evaluation confirms this translates to downstream policy degradation: at 50% corruption, reward-guided selection produces no improvement over random sampling, while the proxy model reports monotonically increasing scores. Together, these results reveal a preference construction problem: the signal entering RLHF is shaped by elicitation context in ways that neither human metacognition, LLM self-monitoring, nor standard evaluation metrics can detect.

cs.CL

Tokens All the Way Down: A Money View of Decentralized Finance

In traditional banking, repeated deposit-and-lend cycles let a single dollar of reserves support multiple dollars of claims. Decentralized finance produces an analogous structure with tokens. Constructing a Token Graph of 10,200 tokens across 200 blockchains, this paper maps the resulting hierarchy and shows that, by late 2025, each dollar of base assets supports $4.7 of total claims. An embedded yield correction disentangles two channels that raw data conflates: a compositional channel, where lending protocols concentrate in deeper tiers and mechanically raise average yields; and a liquidity channel, where each derivation step reduces secondary-market depth and depresses yields in liquidity-sensitive pools. The liquidity channel concentrates in DEX pools and vanishes in lending pools. A yield decomposition shows that the tier gradient operates entirely through fundamental protocol yields, not incentive-token emissions; quantile regressions reveal that the structural associations concentrate in the upper tail of the yield distribution, with near-zero effects at the median. These findings reframe DeFi's "double counting" as a structural risk question and identify liquidity fragmentation as the primary mechanism associated with yield variation across the token hierarchy.

econ.GN

Stability Anchors and Risk Amplifiers: Tail Spillovers Across Stablecoin Designs

This paper investigates systemic risk transmission across stablecoin markets using Quantile Vector Autoregression (QVAR). Analyzing eight major stablecoins with day data coverage from 2021 to 2025, supplemented by minute-level event studies on three additional coins experiencing major depegs until 2025, we document three findings. First, stabilization mechanism dictates tail-risk behavior: fiat-backed stablecoins function as "stability anchors" with near-zero net spillovers across quantiles, while algorithmic and crypto-collateralized designs become risk amplifiers specifically under extreme market conditions. Second, the theoretical risk isolation between fiat and crypto markets breaks down during stress: direct volatility channels emerge between the US Dollar Index and Bitcoin that bypass stablecoin intermediation. Third, Forbes-Rigobon contagion tests across four depeg events show heterogeneous transmission: after adjusting for volatility, algorithmic stablecoins exhibit significant residual contagion while fiat-backed coins show flight-to-quality effects. These findings imply that uniform stablecoin regulation is inappropriate; regulatory capital buffers for extreme losses should be 2--3x higher for non-fiat-backed stablecoins than median-based measures indicate.

econ.GN

Bitcoin Under Stress: Measuring Infrastructure Resilience 2014-2025

Bitcoin's design promises resilience through decentralization, yet the physical infrastructure supporting the network creates hidden dependencies. We present the first longitudinal study of Bitcoin's resilience to submarine cable failures, using 11 years of P2P network data (2014--2025) and 68 verified cable fault events. Applying a Buldyrev-style cascade model at country level, we find that Bitcoin's clearnet (non-TOR) critical failure threshold $p_c \approx 0.72$--$0.92$ for random failures, meaning the vast majority of inter-country cables must fail before significant node disconnection. Targeted attacks are an order of magnitude more effective ($p_c = 0.05$--$0.20$). To address the majority of nodes now using TOR with unobservable locations, we develop a 4-layer multiplex model incorporating TOR relay infrastructure. Because relay bandwidth concentrates in well-connected European countries, TOR adoption increases resilience under current relay geography ($\Delta p_c \approx +0.02$--$+0.10$) rather than introducing hidden fragility. Empirical validation confirms weak physical-layer coupling: 87% of historical cable faults caused less than 5% node impact. We contribute: (1) a multiplex percolation framework for overlay-underlay coupling, including a 4-layer TOR relay model; (2) the first empirical measurement of Bitcoin's physical-layer resilience over a decade; and (3) evidence that TOR adoption amplifies resilience, with distributional bounds quantifying uncertainty under partial observability.

cs.NI

DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks

Credit exposure in Decentralized Finance (DeFi) is often implicit and token-mediated, creating a dense web of inter-protocol dependencies. Thus, a shock to one token may result in significant and uncontrolled contagion effects. As the DeFi ecosystem becomes increasingly linked with traditional financial infrastructure through instruments, such as stablecoins, the risk posed by this dynamic demands more powerful quantification tools. We introduce DeXposure-FM, the first time-series, graph foundation model for measuring and forecasting inter-protocol credit exposure on DeFi networks, to the best of our knowledge. Employing a graph-tabular encoder, with pre-trained weight initialization, and multiple task-specific heads, DeXposure-FM is trained on the DeXposure dataset that has 43.7 million data entries, across 4,300+ protocols on 602 blockchains, covering 24,300+ unique tokens. The training is operationalized for credit-exposure forecasting, predicting the joint dynamics of (1) protocol-level flows, and (2) the topology and weights of credit-exposure links. The DeXposure-FM is empirically validated on two machine learning benchmarks; it consistently outperforms the state-of-the-art approaches, including a graph foundation model and temporal graph neural networks. DeXposure-FM further produces financial economics tools that support macroprudential monitoring and scenario-based DeFi stress testing, by enabling protocol-level systemic-importance scores, sector-level spillover and concentration measures via a forecast-then-measure pipeline. Empirical verification fully supports our financial economics tools. The model and code have been publicly available. Model: https://huggingface.co/EVIEHub/DeXposure-FM. Code: https://github.com/EVIEHub/DeXposure-FM.

cs.LG

Symmetry-engineered and electrically tunable in-plane anomalous Hall effect in oxide heterostructures

The family of Hall effects has long served as a premier probe of how symmetry, magnetic order, and topology intertwine in solids. Recently, the in-plane anomalous Hall effect (IP-AHE), a transverse Hall response driven by in-plane magnetization, has emerged as a distinct member of this family, offering innovative spintronic functionalities and illuminating intricate interplay between mirror-symmetry breaking and in-plane magnetic order. However, practical routes to deterministically and reversibly control IP-AHE remain limited. Here, we establish a symmetry-engineered IP-AHE platform, CaRuO3/La2/3Ca1/3MnO3/CaRuO3 heterostructure on NdGaO3(110), that turns strict mirror-symmetry breaking constraints into effective tuning knobs. IP-AHE in these epitaxial trilayers unambiguously couples to the CaRuO3-buffer-induced mirror-symmetry breaking and faithfully reproduces the ferromagnetic hysteresis. Ionic liquid gating further enables reversible reconfigurations of the symmetry breaking, thereby achieving electrical modulation and ON/OFF switching of IP-AHE. This highly tunable IP-AHE platform opens pathways for exploring nontrivial magnetic order and developing programmable Hall functionalities in planar geometries.

cond-mat.str-el

Isotropic Dirac fermion and anomalous oscillator strength of zeroth Landau level transition

Dirac fermions, characterized by their linear dispersion and relativistic nature, have emerged as a prominent class of quasiparticles in condensed matter physics. While the Dirac equation, initially developed in the context of high-energy physics, provides a remarkable framework for describing the electronic properties of these materials, the inherent symmetry constraints of condensed matter often lead to deviations from the idealized paradigm. In particular, three-dimensional Dirac fermions in solids often exhibit anisotropic behavior, challenging the notion of perfect symmetry inherent in the Dirac equation. Here, we report the observation of isotropic massive Dirac fermions in LaAlSi through Landau level spectroscopy. The presence of three-dimensional massive Dirac fermions across the Fermi energy is demonstrated by quantized and semiclassical analyses of the magnetic field evolution of Landau level transitions. The isotropic topological nature, Fermi velocity, and Dirac mass are evidenced by the identical magneto-infrared response among the Faraday and three Voigt geometries. Furthermore, we observe an unusually large oscillator strength in the zeroth Landau level transition of the Dirac fermion, compared to transitions with higher indices. This phenomenon, supported by model calculations, can be attributed to the combined effects of the partial excitation of Dirac fermion and the resonant dielectric coupling with the Weyl plasma. Our work provides a strategy for realizing ideal quasiparticle excitations and their coupling effects in condensed matter systems, offering a platform for exploring relativistic physics.

cond-mat.mtrl-sci

A High-Flux and High-Efficiency Setup for Magneto-Infrared Spectroscopy

We report the design and implementation of a high-flux, high-efficiency magneto-infrared spectroscopy system optimized for broadband measurements in high magnetic fields. The setup integrates a Fourier transform infrared spectrometer, a 12 T cryogen-free superconducting magnet, precision-polished and gold-plated light tubes, custom-designed reflective focusing modules for Faraday and Voigt geometries, and an external multi-detector chamber with motorized selection. Optical throughput is maximized by reducing light tube loss from 65.5%/m to 22.0%/m via abrasive flow and mechanical polishing followed by gold electroplating, and by adopting a single-on-axis parabolic-mirror Faraday module that increases the effective numerical aperture from 0.14 to 0.36, enhancing collection efficiency by nearly an order of magnitude. An eight-position motorized sample stage and fully automated control over magnetic field, temperature, optical path, and detector choice enable high-throughput measurements without repeated warm-ups. The optimized configuration achieves a root-mean-square noise level of 0.0061% in a 2-minute integration for a 40% reflectivity sample, corresponding to a signal-to-noise ratio exceeding 16000. System capabilities are demonstrated by resolving weak replica bands in EuCd2As2 and faint Landau level transitions in LaAlSi.

cond-mat.mtrl-sci

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, splitting medical multimodal reasoning into fine-grained visual understanding and multi-step reasoning to enable targeted evaluation; 2) Challenging task design, with visual understanding across three key dimensions (small-object detection, fine-detail discrimination, spatial understanding) and reasoning covering four clinically relevant scenarios (temporal prediction, causal reasoning, long-tail generalization, multi-source integration); 3) Broad, high-quality data coverage, comprising 20,653 Visual Question Answering (VQA) pairs spanning 11 organ systems and 12 imaging modalities, validated via a rigorous two-stage (human expert + model-assisted) review to ensure clinical authenticity. We evaluate 18 state-of-the-art MLLMs with Med-CMR, revealing GPT-5 as the top-performing commercial model: 57.81 accuracy on multiple-choice questions (MCQs) and a 48.70 open-ended score, outperforming Gemini 2.5 Pro (49.87 MCQ accuracy, 45.98 open-ended score) and leading open-source model Qwen3-VL-235B-A22B (49.34 MCQ accuracy, 42.62 open-ended score). However, specialized medical MLLMs do not reliably outperform strong general models, and long-tail generalization emerges as the dominant failure mode. Med-CMR thus provides a stress test for visual-reasoning integration and rare-case robustness in medical MLLMs, and a rigorous yardstick for future clinical systems.

cs.AI

DeXposure: A Dataset and Benchmarks for Inter-protocol Credit Exposure in Decentralized Financial Networks

We curate the DeXposure dataset, the first large-scale dataset for inter-protocol credit exposure in decentralized financial networks, covering global markets of 43.7 million entries across 4.3 thousand protocols, 602 blockchains, and 24.3 thousand tokens, from 2020 to 2025. A new measure, value-linked credit exposure between protocols, is defined as the inferred financial dependency relationships derived from changes in Total Value Locked (TVL). We develop a token-to-protocol model using DefiLlama metadata to infer inter-protocol credit exposure from the token's stock dynamics, as reported by the protocols. Based on the curated dataset, we develop three benchmarks for machine learning research with financial applications: (1) graph clustering for global network measurement, tracking the structural evolution of credit exposure networks, (2) vector autoregression for sector-level credit exposure dynamics during major shocks (Terra and FTX), and (3) temporal graph neural networks for dynamic link prediction on temporal graphs. From the analysis, we observe (1) a rapid growth of network volume, (2) a trend of concentration to key protocols, (3) a decline of network density (the ratio of actual connections to possible connections), and (4) distinct shock propagation across sectors, such as lending platforms, trading exchanges, and asset management protocols. The DeXposure dataset and code have been released publicly. We envision they will help with research and practice in machine learning as well as financial risk monitoring, policy analysis, DeFi market modeling, amongst others. The dataset also contributes to machine learning research by offering benchmarks for graph clustering, vector autoregression, and temporal graph analysis.

cs.LG

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of comprehensive evaluation benchmarks that take into account both the human-oriented granular level and higher-dimensional causal reasoning ability. Such high-quality evaluation benchmarks face tough obstacles, given the physical complexity of the human body and the difficulty of annotating granular structures. In this paper, we propose Human-MME, a curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric scene understanding. Compared with other existing benchmarks, our work provides three key features: 1. Diversity in human scene, spanning 4 primary visual domains with 15 secondary domains and 43 sub-fields to ensure broad scenario coverage. 2. Progressive and diverse evaluation dimensions, evaluating the human-based activities progressively from the human-oriented granular perception to the higher-dimensional reasoning, consisting of eight dimensions with 19,945 real-world image question pairs and an evaluation suite. 3. High-quality annotations with rich data paradigms, constructing the automated annotation pipeline and human-annotation platform, supporting rigorous manual labeling to facilitate precise and reliable model assessment. Our benchmark extends the single-target understanding to the multi-person and multi-image mutual understanding by constructing the choice, short-answer, grounding, ranking and judgment question components, and complex questions of their combination. The extensive experiments on 17 state-of-the-art MLLMs effectively expose the limitations and guide future MLLMs research toward better human-centric image understanding. All data and code are available at https://github.com/Yuan-Hou/Human-MME.

cs.CV