SearcharxivSearch

arXiv subjects

Yinghao Zhu

Publications and source records attributed to Yinghao Zhu.

At least 19 recordsLinked to original sources

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design allows minimal baselines that process only the last four frames to match or surpass complex streaming models, while answer options also expose language shortcuts. We introduce StreamArena, a benchmark for hour-scale, interactive streaming video understanding. StreamArena contains 243 full-length videos averaging 88.8 minutes and 3,646 rigorously annotated, open-ended question-answer pairs that evaluate real-time perception, historical retrospection, proactive interaction, and multimodal tool utilization. Evaluation across diverse systems exposes a tension between continuous interaction and long-horizon multimodal comprehension. Methods that retain only recent frames cannot recover distant events, methods that convert past observations into text lose visual evidence, and methods that repeatedly compress visual memory struggle to preserve fine-grained details over time. We address this tension with StreamMind, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search. StreamMind outperforms existing streaming baselines across all four capabilities and reduces query-to-answer latency by reusing persistent state.

cs.CV

LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.

cs.CL

CodeSpec: Dual Executable Specifications for Agentic Long-Horizon Feature Development

LLM-based code agents have advanced repository-level software development through iterative interaction with codebases and tools. However, feature development requires integrating new behaviors into existing architectures through coherent cross-component functional chains. Existing agents typically derive such chains through free-form reasoning, often producing unreliable feature designs with incomplete functional chains. Moreover, textual designs are difficult to verify and enforce, making it challenging to maintain design-implementation consistency throughout long-horizon development. We propose CodeSpec, a dual executable specification method for repository-level feature development. It builds reliable functional chains from evidence pairing sub-requirement semantics with repository architectures, then compiles them into complementary architecture and behavior specifications that check chain completeness and correctness while preserving design-implementation consistency over long interactions. On FeatureBench, which targets feature development in existing repositories, CodeSpec achieves 70.7%, 55.0%, and 49.9% pass rates under DeepSeek-V4-Pro, outperforming representative baselines such as Claude Code. Results on the repository generation benchmark NL2Repo-Bench further demonstrate its generalizability.

cs.SE

MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair

Modern software ecosystems face a rapidly growing number of disclosed vulnerabilities, increasing the need for automated repair techniques that can operate reliably at repository scale. Although Large Language Model (LLM)-based agents have recently shown promise for automated vulnerability repair (AVR), most existing systems still treat repair as a single generation step over the currently visible code context. As a result, they lack a persistent mechanism for reusing prior fixes or learning from failed validation attempts, which limits their effectiveness on complex, multi-file repair tasks. We present MemRepair, a memory-augmented agentic framework that formulates vulnerability repair as an iterative, experience-driven process. MemRepair combines three complementary memory layers, i.e., History-Fix, Security-Pattern, and Refinement-Trajectory memories, with a dynamic feedback-driven refinement loop. This design allows the agent to retrieve repository-specific repair conventions, apply reusable security defenses, and exploit prior "failure-to-success" trajectories to revise semantically invalid patches based on runtime evidence. We evaluate MemRepair on three representative repository-level vulnerability repair benchmarks: SEC-Bench, PatchEval (Python, Go, JavaScript), and the C++ subset of Multi-SWE-bench. MemRepair achieves state-of-the-art resolution rates of 58.0%, 58.2%, and 30.58%, respectively, outperforming strong general-purpose agents such as OpenHands and SWE-agent, as well as the specialized AVR tool InfCode-C++, while maintaining competitive repair cost. These results show that persistent, hierarchical repair memory can substantially improve the reliability of agentic vulnerability repair across diverse languages and repository settings.

cs.SE

ContraFix: Skill-Enhanced Contrastive Runtime Analysis for Vulnerability Repair

As software systems grow increasingly complex, automated vulnerability repair (AVR) remains difficult because the materials available to a repair system are usually failure artifacts rather than repair guidance. Traditional analysis techniques can provide suspicious locations, reduced triggers, or constraints, but they are costly to configure across repositories and seldom directly actionable for patch generation. Recent LLM-based agents can edit and validate repository-level patches, and experience-based systems can reuse prior repair traces or demonstrations, but they still need current-instance evidence that turns a broad, symptom-level failure report into a concrete repair decision. We present ContraFix, an agentic AVR framework that constructs such evidence through contrastive runtime analysis. Starting from a failing witness, ContraFix generates nearby failing and non-failing variants, executes them through aligned probe sites, and compares their runtime states to infer the repair boundary and guide source-level patching. Each candidate patch is accepted only after build and validation. ContraFix also stores validated repair episodes in a dual-track skill base, reusing mutation skills to construct useful variants and correction skills to refine failed patches. On SEC-Bench, ContraFix with GPT-5-mini achieves resolution rate of 92.0% over three repeated runs and an average resolution rate of 91.8% +/- 0.8. On PatchEval, it resolves 73.8% of 225 Go, Python, and JavaScript instances. A semantic audit of benchmark-validated SEC-Bench patches shows that 58.2% of ContraFix's patches are semantically correct, compared with 31.3% for the strongest baseline, indicating that the proposed framework improves semantic correctness beyond benchmark validation.

cs.SE

Nature of magnetism in bilayer nickelate La3Ni2O7 single crystals

The recent discovery of high-temperature superconductivity in pressurized and thin film nickelates has generated intense interest, yet the nature of magnetism in their ambient-pressure parent phases remains poorly understood, despite its potentially crucial role in pairing. Here we use neutron scattering to resolve the spin order and dynamics of single-crystalline La3Ni2O7, an ambient-pressure parent of this class. Well defined spin excitations are observed at Q = (0, 0.5, 2.5), featuring a~5 meV spin gap and anisotropic in-plane dispersions, with zone-boundary softening along the transverse direction indicative of competing exchange interactions. The excitations exhibit pronounced out-of-plane modulations with bilayer periodicity, providing direct evidence for antiferromagnetic interlayer coupling. Their dispersion is well described by a bilayer Heisenberg Hamiltonian with strong interlayer exchange and competing in-plane couplings within a stripe-type magnetic order. Normalization of the spectra to absolute units reveals that, although the spin-wave bandwidth is only about 25% of that in cuprates, the local dynamic susceptibility at comparable energies is significantly enhanced, yielding a total fluctuating moment of comparable magnitude. These results highlight intense mid-energy spin excitations rooted in substantial electronic correlations as a defining feature of this family, establishing a magnetic framework distinct from cuprates and directly relevant to understanding superconductivity in this system.

cond-mat.str-el

Collective spin excitations in trilayer nickelate La$_4$Ni$_3$O$_{10}$

Ruddlesden-Popper (RP) nickelates have recently emerged as a new family of high-temperature superconductors. In bilayer RP nickelates, magnetic excitations with large exchange couplings have been observed, supporting a spin-mediated pairing mechanism. Whether comparable spin correlations persist in trilayer nickelates, however, remains unknown. Here, we present a Ni $L$-edge resonant inelastic X-ray scattering (RIXS) study of La$_4$Ni$_3$O$_{10}$ single crystals. While the orbital excitations remain similar to those of La$_3$Ni$_2$O$_{7}$, the collective spin excitations in La$_4$Ni$_3$O$_{10}$ exhibit a comparable bandwidth of about $60$ meV but substantially suppressed spectral weight, implying a weaker electronic correlation in the trilayer compounds. Our results underscore the three-dimensional and multi-orbital electronic character in La$_4$Ni$_3$O$_{10}$, highlighting important differences from the bilayer nickelates. These findings provide crucial insights into the evolution of magnetism across the RP nickelate family and its connection to superconductivity.

cond-mat.supr-con

Tilted and Twisted Magnetic Moments in the Kitaev Magnet $\alpha$-RuCl$_3$

The layered honeycomb magnet $\alpha$-RuCl$_3$ has attracted intense scrutiny as a prime candidate for realizing the Kitaev quantum spin liquid, yet a consensus on its microscopic Hamiltonian remains elusive due to the material's extreme sensitivity to structural details. Here, we report a comprehensive reexamination of the low-temperature crystallographic and magnetic structures of high-quality $\alpha$-RuCl$_3$ single crystals using unpolarized and polarized neutron diffraction. We confirm a sharp, first-order structural phase transition to the rhombohedral $R\bar{3}$ space group with a pronounced thermal hysteresis. Crucially, using both spherical and longitudinal neutron polarization analysis, we determine the 3D orientation of the ordered magnetic moment without the ambiguity typically arising from domain distributions. We find that the Ru$^{3+}$ magnetic moments in the zigzag phase are tilted by $15.7^\circ$ out of the hexagonal plane and, remarkably, exhibit an additional in-plane twist of $-13.8^\circ$. This "tilted and twisted" geometry differentiates the ground state from the previously reported models based on unpolarized neutron diffraction or resonant elastic X-ray scattering (REXS) analysis.

cond-mat.str-el

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open-access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata-driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transforming multiple small data silos into larger, more coherent resources. Building on MDFP, we release an interactive discovery portal that enables end-to-end, automated medical image dataset integration, and compile all surveyed datasets into a unified, structured table that clearly summarizes their key characteristics and provides reference links, offering the community an accessible and comprehensive repository. By charting the current terrain and offering a principled path to dataset consolidation, our survey provides a practical roadmap for scaling medical imaging corpora, supporting faster data discovery, more principled dataset creation, and more capable medical foundation models.

cs.CV

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiographic evidence) to determine complete referral plans. We present Dental-TriageBench, the first expert-annotated benchmark for reasoning-driven multimodal dental triage. Built from authentic outpatient workflows, it contains 246 de-identified cases annotated with expert-authored golden reasoning trajectories, together with hierarchical triage labels. We benchmark 19 proprietary, open-source, and medical-domain MLLMs against three junior dentists serving as the human baseline, and find a substantial human--model gap, on fine-grained treatment-level triage. Further analyses show that accurate triage requires both complaint and OPG information, and that model errors concentrate on cases with multiple referral domains, where MLLMs tend to produce overly narrow referral sets and omission-heavy errors. Dental-TriageBench provides a realistic testbed for developing multimodal clinical AI systems that are more clinically grounded, coverage-aware, and safer for downstream care.

cs.CL

On estimating superconducting shielding volume fraction from susceptibility in pressurized Ruddlesden-Popper nickelates: Response to arXiv:2602.19282

In a recent preprint (arXiv:2602.19282) [1], the authors questioned the procedure we used to evaluate the demagnetization-corrected superconducting shielding volume fraction in pressurized Ruddlesden-Popper nickelates [2-5]. They further claimed that this methodology has neither been derived nor used previously, and they proposed an alternative normalization scheme. Here we clarify that our evaluation follows directly from the standard magnetostatic self-consistency relation for finite samples and has been widely adopted in the superconductivity literature for decades. We also demonstrate that the discrepancies claimed in Ref. [1] stem from a fundamental flaw in their approach, namely, the assumption that the measured diamagnetic moment is linearly proportional to the superconducting shielding volume fraction in the presence of a finite demagnetization factor N. This assumption is not valid for strongly demagnetized, thin disk-like specimens, where the internal field and the measured moment are coupled self-consistently through the demagnetizing field.

cond-mat.supr-con

Contrasting Momentum-Selective Spin-Density-Wave Gaps in Bilayer and Trilayer Nickelates

Resolving where the density-wave gap opens in momentum space is essential for identifying the microscopic origin of the instability in layered nickelates. Using polarization-resolved electronic Raman scattering, we map the momentum selectivity of the spin-density-wave (SDW) gap in trilayer La4Ni3O10. We observe a SDW-induced redistribution of spectral weight on both the $\alpha$ pocket at the Brillouin-zone centre and a portion of the $\beta$ pocket near the zone boundary, characterized by gap energies of approximately 55~meV. In contrast, no comparable spectral weight suppression is observed along the diagonal region of $\beta$ pockets, implying little or no gap opening. This gap topology contrasts sharply with that in La3Ni2O7, where anisotropic SDW gaps open solely on the $\beta$ pocket. Our results establish a distinct momentum-space gap topology between bilayer and trilayer nickelates, placing new constraints on the ordering wave vector and the mechanism of the density-wave instability relevant to superconductivity.

cond-mat.supr-con

Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and Obstetrics

Clinician skepticism toward opaque AI hinders adoption in high-stakes healthcare. We present AICare, an interactive and interpretable AI copilot for collaborative clinical decision-making. By analyzing longitudinal electronic health records, AICare grounds dynamic risk predictions in scrutable visualizations and LLM-driven diagnostic recommendations. Through a within-subjects counterbalanced study with 16 clinicians across nephrology and obstetrics, we comprehensively evaluated AICare using objective measures (task completion time and error rate), subjective assessments (NASA-TLX, SUS, and confidence ratings), and semi-structured interviews. Our findings indicate AICare's reduced cognitive workload. Beyond performance metrics, qualitative analysis reveals that trust is actively constructed through verification, with interaction strategies diverging by expertise: junior clinicians used the system as cognitive scaffolding to structure their analysis, while experts engaged in adversarial verification to challenge the AI's logic. This work offers design implications for creating AI systems that function as transparent partners, accommodating diverse reasoning styles to augment rather than replace clinical judgment.

cs.HC

MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing

Medical text editing is essential for improving communication among diverse stakeholders in clinical settings. However, adapting LLM agents to this task remains challenging because expert supervision is often sparse, fragmented, and distributed across interacting quality dimensions. We identify that direct accumulation or retrieval of individual feedback is insufficient for effective adaptation, as fragmented evaluations do not directly translate into a coherent understanding of medical text quality. Based on this observation, we propose MedTextWeaver, a training-free framework that transforms fragmented evaluative evidence into global quality principles and actionable procedural knowledge for medical text editing. Across three clinical text datasets and a real-world validation experiment, MedTextWeaver consistently improves performance over strong LLM baselines and existing memory-based adaptation approaches. Further analysis demonstrates that the learned knowledge enables more effective adaptation under limited supervision while providing an explicit and interpretable interface between expert evaluations and LLM editing behavior.

cs.CL

SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation

Search agents have emerged as a pivotal paradigm for solving open-ended, knowledge-intensive reasoning tasks. However, training these agents via Reinforcement Learning (RL) faces a critical dilemma: interacting with live commercial Web APIs is prohibitively expensive, while relying on static data snapshots often introduces noise due to data misalignment. This misalignment generates corrupted reward signals that destabilize training by penalizing correct reasoning or rewarding hallucination. To address this, we propose SearchGym, a simulation environment designed to bootstrap robust search agents. SearchGym employs a rigorous generative pipeline to construct a verifiable knowledge graph and an aligned document corpus, ensuring that every reasoning task is factually grounded and strictly solvable. Building on this controllable environment, we introduce SearchGym-RL, a curriculum learning methodology that progressively optimizes agent policies through purified feedback, evolving from basic interactions to complex, long-horizon planning. Extensive experiments across the Llama and Qwen families demonstrate strong Sim-to-Real generalization. Notably, our Qwen2.5-7B-Base model trained within SearchGym surpasses the web-enhanced ASearcher baseline across nine diverse benchmarks by an average relative margin of 10.6%. Our results validate that high-fidelity simulation serves as a scalable and highly cost-effective methodology for developing capable search agents.

cs.CL

CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation

Large language models (LLMs) substantially enhance developer productivity in repository-level code generation through interactive collaboration. However, as interactions progress, repository context must be continuously preserved and updated to integrate newly validated information. Meanwhile, the expanding session history increases cognitive burden, often leading to forgetting and the reintroduction of previously resolved errors. Existing memory management approaches show promise but remain limited by natural language-centric representations. To overcome these limitations, we propose CodeMEM, an AST-guided dynamic memory management system tailored for repository-level iterative code generation. Specifically, CodeMEM introduces the Code Context Memory component that dynamically maintains and updates repository context through AST-guided LLM operations, along with the Code Session Memory that constructs a code-centric representation of interaction history and explicitly detects and mitigates forgetting through AST-based analysis. Experimental results on the instruction-following benchmark CodeIF-Bench and the code generation benchmark CoderEval demonstrate that CodeMEM achieves state-of-the-art performance, improving instruction following by 12.2% for the current turn and 11.5% for the session level, and reducing interaction rounds by 2-3, while maintaining competitive inference latency and token efficiency.

cs.SE

Nonthermal melting and density wave instability coupled to the lattice in La$_4$Ni$_3$O$_{10}$

The recent discovery of high-temperature superconductivity in pressurized nickelates has renewed interest in the broken-symmetry states of their ambient-pressure parent phases, where a density-wave (DW) order emerges and competes with superconductivity, but its microscopic origin remains unresolved. Using ultrafast optical spectroscopy, we track quasiparticle relaxation dynamics across the DW transition at $T_{\rm DW} \approx$ 136 K in trilayer nickelate {\LNO} single crystals, revealing the opening of an energy gap of $\sim$52 meV. Multiple coherent phonons, including $A_g$ modes near 3.88, 5.28, and 2.09 THz, display pronounced mode-selective anomalies across the transition, indicating that the DW is strongly coupled to lattice degrees of freedom and suggesting an important role of electron-phonon coupling. At higher excitation densities, the DW is nonthermally suppressed, producing a temperature-fluence phase diagram that parallels pressure-tuned behavior. These results establish the DW in {\LNO} as a lattice-entangled instability involving multiple phonon modes, and highlight ultrafast optical excitation as a nonequilibrium tuning parameter for suppressing density-wave order in nickelates.

cond-mat.str-el

Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities

The clinical translation of dermatological AI is hindered by opaque reasoning and systematic performance disparities across skin tones. Here we present SkinGPT-R1, a multimodal large language model that integrates chain-of-thought diagnostic reasoning with a fairness-aware mixture-of-experts architecture for interpretable and equitable skin disease diagnosis. Through parameter-efficient adaptation of a frozen reasoning backbone, SkinGPT-R1 generates structured diagnostic reports comprising visual findings, differential reasoning, and final diagnosis. Across seven external datasets spanning diverse pathologies and imaging conditions, SkinGPT-R1 achieves state-of-the-art accuracy on six benchmarks, including 82.50\% on a challenging 40-class long-tail classification task (+19.30\% over leading baselines). Blinded evaluation by five board-certified dermatologists on 1,000 phenotypically balanced cases yields a mean score of 3.6 out of 5, with the highest ratings in safety (3.8) and reasoning coherence (3.6), indicating that the generated rationales are clinically safe, logically grounded, and suitable for supporting diagnostic decision-making. Critically, SkinGPT-R1 mitigates algorithmic bias across the full Fitzpatrick spectrum, achieving a robust worst-group performance of 41.40\% on the Fitz17k benchmark and a five-fold relative improvement in lower-bound accuracy on the DDI dataset compared to standard multimodal baselines. These results establish a framework for trustworthy, fair, and explainable AI-assisted dermatological diagnosis.

cs.CV