SearcharxivSearch

arXiv subjects

Yuyang Ji

Publications and source records attributed to Yuyang Ji.

18 recordsLinked to original sources

Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning

Digital entities such as AI agents and humanoid robots increasingly operate alongside real humans, yet their identity infrastructure is based on credentials rather than embodied biometric identity. We introduce Biometric Identity Provisioning (BIP), a new problem and solution framework that addresses: given an enrollment gallery of real human identities, provision virtual identities that are non-colliding with every enrolled identity, maintain sufficient inter-class separability, and are realizable as high-fidelity face images. The key geometric insight is that real face identities occupy a low-dimensional subspace of the embedding hypersphere, leaving no residual subspace for virtual identities. Hence, virtual identities must instead be allocated as unclaimed gaps within the real face manifold itself. BIP is therefore a constrained packing problem: available gaps vastly exceed any foreseeable enrollment scale, and provisioned identities remain non-colliding even as new real identities are subsequently enrolled. Grounded in this geometry, our repulsion-based allocation is not bounded by any fixed provisioning count; we demonstrate 10M non-colliding virtual identity embeddings against a gallery of 360K real identities. Realizing these embeddings as face images requires a generator that operates outside the training distribution of real face images; we introduce GapGen, a gap-aware generator trained with a curriculum that progressively extends synthesis into non-colliding regions, validated at 1M photorealistic virtual face images. We further construct v-LFW, a virtual counterpart to LFW face dataset, with protocols for virtual face verification, cross-reality matching, real-vs-virtual detection, and unified recognition and detection.

cs.CV

From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching

We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exercise-specific degree-of-freedom selector that focuses analysis on salient joints; a structured biomechanical context that pairs individualized morphometrics with cycle and constraint analysis; and a vision--biomechanics conditioned feedback module that applies cross-attention to generate precise, actionable text. Using parameter-efficient training that freezes the vision and language backbones, BioCoach yields transparent, personalized reasoning rather than pattern matching. To enable learning and fair evaluation, we augment QEVD-fit-coach with biomechanics-oriented feedback to create QEVD-bio-fit-coach, and we introduce a biomechanics-aware LLM judge metric. BioCoach delivers clear gains on QEVD-bio-fit-coach across lexical and judgment metrics while maintaining temporal triggering; on the original QEVD-fit-coach, it improves text quality and correctness with near-parity timing, demonstrating that explicit kinematics and constraints are key to accurate, phase-aware coaching.

cs.CV

Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos

K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these interactions has made automated analysis elusive. Existing benchmarks for classroom discourse focus primarily on mathematics and rely solely on transcripts, overlooking the visual artifacts and model-based reasoning emphasized by the Next Generation Science Standards (NGSS). We address this gap with SciIBI, the first video benchmark for analyzing science classroom discourse, featuring 113 NGSS-aligned clips annotated with Core Instructional Practices (CIP) and sophistication levels. By evaluating eight state-of-the-art LLMs and Multimodal LLMs, we reveal fundamental limitations: current models struggle to distinguish pedagogically similar practices, suggesting that CIP coding requires instructional reasoning beyond surface pattern matching. Furthermore, adding video input yields inconsistent gains across architectures. Crucially, our evidence-based evaluation reveals that models often succeed through surface shortcuts rather than genuine pedagogical understanding. These findings establish science classroom discourse as a challenging frontier for multimodal AI and point toward human-AI collaboration, where models retrieve evidence to accelerate expert review rather than replace it.

cs.CY

IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition

Video-based person recognition achieves robust identification by integrating face, body, and gait. However, current systems waste computational resources by processing all modalities with fixed heavyweight ensembles regardless of input complexity. To address these limitations, we propose IDSelect, a reinforcement learning-based cost-aware selector that chooses one pre-trained model per modality per-sequence to optimize the accuracy-efficiency trade-off. Our key insight is that an input-conditioned selector can discover complementary model choices that surpass fixed ensembles while using substantially fewer resources. IDSelect trains a lightweight agent end-to-end using actor-critic reinforcement learning with budget-aware optimization. The reward balances recognition accuracy with computational cost, while entropy regularization prevents premature convergence. At inference, the policy selects the most probable model per modality and fuses modality-specific similarities for the final score. Extensive experiments on challenging video-based datasets demonstrate IDSelect's superior efficiency: on CCVID, it achieves 95.9% Rank-1 accuracy with 92.4% less computation than strong baselines while improving accuracy by 1.8%; on MEVID, it reduces computation by 41.3% while maintaining competitive performance.

cs.CV

BioGait-VLM: A Tri-Modal Vision-Language-Biomechanics Framework for Interpretable Clinical Gait Assessment

Video-based Clinical Gait Analysis often suffers from poor generalization as models overfit environmental biases instead of capturing pathological motion. To address this, we propose BioGait-VLM, a tri-modal Vision-Language-Biomechanics framework for interpretable clinical gait assessment. Unlike standard video encoders, our architecture incorporates a Temporal Evidence Distillation branch to capture rhythmic dynamics and a Biomechanical Tokenization branch that projects 3D skeleton sequences into language-aligned semantic tokens. This enables the model to explicitly reason about joint mechanics independent of visual shortcuts. To ensure rigorous benchmarking, we augment the public GAVD dataset with a high-fidelity Degenerative Cervical Myelopathy (DCM) cohort to form a unified 8-class taxonomy, establishing a strict subject-disjoint protocol to prevent data leakage. Under this setting, BioGait-VLM achieves state-of-the-art recognition accuracy. Furthermore, a blinded expert study confirms that biomechanical tokens significantly improve clinical plausibility and evidence grounding, offering a path toward transparent, privacy-enhanced gait assessment.

cs.CV

Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs

Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited context windows. In this work, we introduce VideoMindPalace, a new framework inspired by the "Mind Palace", which organizes critical video moments into a topologically structured semantic graph. VideoMindPalace organizes key information through (i) hand-object tracking and interaction, (ii) clustered activity zones representing specific areas of recurring activities, and (iii) environment layout mapping, allowing natural language parsing by LLMs to provide grounded insights on spatio-temporal and 3D context. In addition, we propose the Video MindPalace Benchmark (VMB), to assess human-like reasoning, including spatial localization, temporal reasoning, and layout-aware sequential understanding. Evaluated on VMB and established video QA datasets, including EgoSchema, NExT-QA, IntentQA, and the Active Memories Benchmark, VideoMindPalace demonstrates notable gains in spatio-temporal coherence and human-aligned reasoning, advancing long-form video analysis capabilities in VLMs.

cs.CV

Real-time time-dependent density functional theory simulations with range-separated hybrid functionals for periodic systems

Real-time time-dependent density functional theory (RT-TDDFT) is a powerful approach for investigating various ultrafast phenomena in materials. However, most existing RT-TDDFT studies rely on adiabatic local or semi-local approximations, which suffer from several shortcomings, including the inability to accurately capture excitonic effects in periodic systems. Combining RT-TDDFT with range-separated hybrid (RSH) functionals has emerged as an effective strategy to overcome these limitations. The RT-TDDFT-RSH implementation for periodic systems requires careful treatment of the Coulomb singularity and choosing proper gauges for the incorporation of external fields. We benchmark two schemes for treating the Coulomb singularity - the truncated Coulomb potential and the auxiliary-function correction - and find that the latter shows better convergence behavior and numerical stability for long-range corrected hybrid functions. Additionally, we assess the impact of gauge choice in simulations using numerical atomic orbitals and show that the recently proposed hybrid gauge incorporating position-dependent phases provides a more accurate description of excitonic absorption than the conventional velocity gauge. Our implementation significantly improves the accuracy of RT-TDDFT-RSH for modeling ultrafast excitonic dynamics in periodic systems.

cond-mat.mtrl-sci

ABACUS: An Electronic Structure Analysis Package for the AI Era

ABACUS (Atomic-orbital Based Ab-initio Computation at USTC) is an open-source software for first-principles electronic structure calculations and molecular dynamics simulations. It mainly features density functional theory (DFT) and molecular dynamics functions and is compatible with both plane-wave basis sets and numerical atomic orbital basis sets. ABACUS serves as a platform that facilitates the integration of various electronic structure methods, such as Kohn-Sham DFT, stochastic DFT, orbital-free DFT, and real-time time-dependent DFT, etc. In addition, with the aid of high-performance computing, ABACUS is designed to perform efficiently and provide massive amounts of first-principles data for generating general-purpose machine learning potentials, such as DPA models. Furthermore, ABACUS serves as an electronic structure platform that interfaces with several AI-assisted algorithms and packages, such as DeePKS-kit, DeePMD, DP-GEN, DeepH, DeePTB, HamGNN, etc.

cond-mat.mtrl-sci

IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Experts

Large language model (LLM) agents have emerged as a promising solution to automate the workflow of machine learning, but most existing methods share a common limitation: they attempt to optimize entire pipelines in a single step before evaluation, making it difficult to attribute improvements to specific changes. This lack of granularity leads to unstable optimization and slower convergence, limiting their effectiveness. To address this, we introduce Iterative Refinement, a novel strategy for LLM-driven ML pipeline design inspired by how human ML experts iteratively refine models, focusing on one component at a time rather than making sweeping changes all at once. By systematically updating individual components based on real training feedback, Iterative Refinement improves overall model performance. We also provide some theoretical edvience of the superior properties of this Iterative Refinement. Further, we implement this strategy in IMPROVE, an end-to-end LLM agent framework for automating and optimizing object classification pipelines. Through extensive evaluations across datasets of varying sizes and domains, we demonstrate that Iterative Refinement enables IMPROVE to consistently achieve better performance over existing zero-shot LLM-based approaches.

cs.CV

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning either rely on training-free prompting or large-scale fine-tuning; both lack active tool exploration and typically assume limited tool diversity, and fine-tuning methods additionally demand extensive human supervision. In contrast, VisTA leverages end-to-end reinforcement learning to iteratively refine sophisticated, query-specific tool selection strategies, using task outcomes as feedback signals. Through Group Relative Policy Optimization (GRPO), our framework enables an agent to autonomously discover effective tool-selection pathways without requiring explicit reasoning supervision. Experiments on the ChartQA, Geometry3K, and BlindTest benchmarks demonstrate that VisTA achieves substantial performance gains over training-free baselines, especially on out-of-distribution examples. These results highlight VisTA's ability to enhance generalization, adaptively utilize diverse tools, and pave the way for flexible, experience-driven visual reasoning systems.

cs.CV

Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding

Multimodal Large Language Models (MLLMs) have shown remarkable versatility but face challenges in demonstrating true visual understanding, particularly in chart reasoning tasks. Existing benchmarks like ChartQA reveal significant reliance on text-based shortcuts and probabilistic pattern-matching rather than genuine visual reasoning. To rigorously evaluate visual reasoning, we introduce a more challenging test scenario by removing textual labels and introducing chart perturbations in the ChartQA dataset. Under these conditions, models like GPT-4o and Gemini-2.0 Pro experience up to a 30% performance drop, underscoring their limitations. To address these challenges, we propose Socratic Chart, a new framework that transforms chart images into Scalable Vector Graphics (SVG) representations, enabling MLLMs to integrate textual and visual modalities for enhanced chart understanding. Socratic Chart employs a multi-agent pipeline with specialized agent-generators to extract primitive chart attributes (e.g., bar heights, line coordinates) and an agent-critic to validate results, ensuring high-fidelity symbolic representations. Our framework surpasses state-of-the-art models in accurately capturing chart primitives and improving reasoning performance, establishing a robust pathway for advancing MLLM visual understanding.

cs.CV

Do Vision Models Develop Human-Like Progressive Difficulty Understanding?

When a human undertakes a test, their responses likely follow a pattern: if they answered an easy question $(2 \times 3)$ incorrectly, they would likely answer a more difficult one $(2 \times 3 \times 4)$ incorrectly; and if they answered a difficult question correctly, they would likely answer the easy one correctly. Anything else hints at memorization. Do current visual recognition models exhibit a similarly structured learning capacity? In this work, we consider the task of image classification and study if those models' responses follow that pattern. Since real images aren't labeled with difficulty, we first create a dataset of 100 categories, 10 attributes, and 3 difficulty levels using recent generative models: for each category (e.g., dog) and attribute (e.g., occlusion), we generate images of increasing difficulty (e.g., a dog without occlusion, a dog only partly visible). We find that most of the models do in fact behave similarly to the aforementioned pattern around 80-90% of the time. Using this property, we then explore a new way to evaluate those models. Instead of testing the model on every possible test image, we create an adaptive test akin to GRE, in which the model's performance on the current round of images determines the test images in the next round. This allows the model to skip over questions too easy/hard for itself, and helps us get its overall performance in fewer steps.

cs.CV

Spin density wave in the bilayered nickelate La$_3$Ni$_2$O$_{7-δ}$ at ambient pressure

The recent discovery of high-temperature superconductivity in high-pressurized La$_3$Ni$_2$O$_{7-δ}$ has garnered significant attention. Using density functional theory, we investigate the magnetic properties of La$_3$Ni$_2$O$_{7-δ}$ at ambient pressure. Our calculations suggest that with $δ=0$, the double spin stripe phase is favored as the magnetic ground state. Oxygen vacancies may effectively turn nearest Ni spins into \textit{charge} sites. Consequently, with moderate $δ$ values, our theoretical magnetic ground state exhibits characteristics of both double spin stripe and spin-charge stripe configurations, providing a natural explanation to reconcile the seemingly contradictory experimental findings that suggest both the configurations as candidates for the spin-density-wave phase. With higher $δ$ values, we anticipate the ground state to become a spin-glass-like noncollinear magnetic phase with only short-range order. The oxygen vacancies are expected to significantly impact the magnetic excitations and the transition temperatures $T_{SDW}$. Notably, the magnetic ordering also induces concomitant charge ordering and orbital ordering, driven by spin-lattice coupling under the low symmetry magnetic order. We further offer a plausible explanation for the experimental observations that the measured $T_{SDW}$ appears insensitive to the variation of samples and the lack of direct evidence for long-range magnetic ordering.

cond-mat.mtrl-sci

Efficient hybrid-functional-based force and stress calculations for periodic systems with thousands of atoms

We present an efficient linear-scaling algorithm for evaluating the analytical force and stress contributions derived from the exact-exchange energy, a key component in hybrid functional calculations. The algorithm, working equally well for molecular and periodic systems, is formulated within the framework of numerical atomic orbital (NAO) basis sets and takes advantage of the localized resolution-of-identity (LRI) technique for treating the two-electron Coulomb repulsion integrals. The linear-scaling behavior is realized by fully exploiting the sparsity of the expansion coefficients resulting from the strict locality of the NAOs and the LRI ansatz. Our implementation is massively parallel, and enables efficient structural relaxation based on hybrid density functionals for bulk materials containing thousands of atoms. In this work, we will present a detailed description of our algorithm and benchmark the performance of our implementation using illustrating examples. By optimizing the structures of the pristine and doped halide perovskite material CsSnI$_3$ with different functionals, we find that in the presence of lattice strain, hybrid functionals provide a more accurate description of the stereochemical expression of the lone pair.

physics.comp-ph

AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation

Pre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks. However, we often fail to fully unleash their potential when adapting them for new concept understanding due to limited information on new classes. To address this limitation, we introduce a novel adaptation framework, AWT (Augment, Weight, then Transport). AWT comprises three key components: augmenting inputs with diverse visual perspectives and enriched class descriptions through image transformations and language models; dynamically weighting inputs based on the prediction entropy; and employing optimal transport to mine semantic correlations in the vision-language space. AWT can be seamlessly integrated into various VLMs, enhancing their zero-shot capabilities without additional training and facilitating few-shot learning through an integrated multimodal adapter module. We verify AWT in multiple challenging scenarios, including zero-shot and few-shot image classification, zero-shot video action recognition, and out-of-distribution generalization. AWT consistently outperforms the state-of-the-art methods in each setting. In addition, our extensive studies further demonstrate AWT's effectiveness and adaptability across different VLMs, architectures, and scales.

cs.CV

PYATB: An Efficient Python Package for Electronic Structure Calculations Using Ab Initio Tight-Binding Model

We present PYATB, a Python package designed for computing band structures and related properties of materials using the ab initio tight-binding Hamiltonian. The Hamiltonian is directly obtained after conducting self-consistent calculations with first-principles packages using numerical atomic orbital (NAO) bases, such as ABACUS. The package comprises three modules: Bands, Geometric, and Optical. In the Bands module, one can calculate essential properties of band structures, including the partial density of states (PDOS), fat bands, Fermi surfaces, and Weyl/Dirac points. The band unfolding method is utilized to obtain the energy band spectra of a supercell by projecting the electronic structure of the supercell onto the Brillouin zone of the primitive cell. With the Geometric module, one can compute the Berry phase and Berry curvature-related quantities, such as electric polarization, Wilson loops, Chern numbers, and anomalous Hall conductivities. The Optical module offers a range of optical property calculations, including optical conductivity and nonlinear optical responses, such as shift current and Berry curvature dipole.

cond-mat.mtrl-sci

Reproducibility of Hybrid Density Functional Calculations for Equation-of-State Properties and Band Gaps

Hybrid density functional (HDF) approximations usually deliver higher accuracy than local and semilocal approximations to the exchange-correlation functional, but this comes with drastically increased computational cost. Practical implementations of HDFs inevitably involve numerical approximations -- even more so than their local and semilocal counterparts due to the additional numerical complexity arising from treating the exact-exchange component. This raises the question regarding the reproducibility of the HDF results yielded by different implementations. In this work, we benchmark the numerical precision of four independent implementations of the popular Heyd-Scuseria-Ernzerhof (HSE) range-separated HDF on describing key materials' properties, including both properties derived from equations of states (EOS) and band gaps of 20 crystalline solids. We find that the energy band gaps obtained by the four codes agree with each other rather satisfactorily. However, for lattice constants and bulk moduli, the deviations between the results computed by different codes are of the same order of magnitude as the deviations between the computational and experimental results. On the one hand, this means that the HSE functional is rather accurate for describing the cohesive properties of simple insulating solids. On the other hand, this also suggests that the numerical precision achieved with current major HSE implementation is not sufficiently high to unambiguously assess the physical accuracy of HDFs. It is found that the pseudopotential treatment of the core electrons is a major factor that contributes to this uncertainty.

cond-mat.mtrl-sci

Importance of exact exchange to the geometric and electronic structures of Cs$_2$$B$$B'$$X_6$ double perovskites

We investigate the lead-free halide double perovskites (HDPs) Cs$ _2BB'X_6$ ($B$=Ag, Na; $B'$=In, Bi; $X$=Cl, Br) via first-principles calculations. We find that both the geometric and electric structures of the HDPs obtained by the Heyd-Scuseria-Ernzerhof (HSE) hybrid functional are much better than those of the Perdew-Burke-Ernzerhof (PBE) functional. Importantly, we find that the electronic structures of DHPs are very sensitive to their geometries, especially the $B$-$X$ bond lengths. As a consequence, the electronic structures calculated by the HSE functional using the PBE optimized geometries may still significantly underestimate the band gaps, whereas the calculations on the HSE optimized geometries provide much more satisfactory results. The sensitivity of the band gaps of the DHPs to their geometries opens a promising path for the band structure engineering via doping and alloying. This work therefore provides an useful guideline for further improvement of HDPs materials.

cond-mat.mtrl-sci