SearcharxivSearch

arXiv subjects

Hao Hao

Publications and source records attributed to Hao Hao.

At least 19 recordsLinked to original sources

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.

cs.CL

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depend on manually designed rubrics that scale poorly to long-tail pedagogical scenarios. We introduce Elmes*, an end-to-end framework for constructing, refining, and applying fine-grained scenario-specific rubrics. Elmes* combines a declarative multi-agent engine for teacher--student--judge interactions with SceneGen, a self-evolving module that co-optimizes evaluation criteria and test data from expert-defined pedagogical dimensions. Using Elmes*, we build Edu-330, covering 330 scenarios across 11 subjects, 3 grade bands, and 10 task types, with over 1{,}000 second-level indicators. Experiments on Edu-330 and four expert-authored gold-standard scenarios show that educational capability is multidimensional: top-tier LLMs differ mainly in creativity and values integration, knowledge-strong models may fail at Socratic scaffolding, and the education-specialized InnoSpark achieves the best human-evaluated average score. LLM judges preserve human-comparable rankings with much lower scoring variance, but exhibit judge-specific biases such as self-preference. Ablations show that expert-scored few-shot anchoring improves human--LLM alignment, while reasoning enforcement and greedy decoding are model-dependent. Elmes* thus provides scalable diagnostic infrastructure for pedagogically grounded LLM evaluation.

cs.LG

Relation Reasoning with LLMs in Expensive Optimization

Expensive optimization problems (EOPs) are black-box tasks with costly objective evaluations and no gradient access, making the evaluation budget the key bottleneck. Surrogate-assisted evolutionary algorithms (SAEAs) reduce evaluations via surrogate predictions, but conventional surrogates often require frequent retraining as populations evolve, incurring overhead. This paper proposes R2SAEA, a reinforcement-trained relation-based large language model (LLM) surrogate assisted evolutionary algorithm. We cast relation-based surrogate modeling as an in-context pairwise reasoning task. To enable efficient inference in evolutionary loops, we develop an anchor-based iterative context construction strategy that reduces prompt complexity from quadratic to linear in population size, and a voting-based aggregation scheme that converts predicted relations into scores for offspring selection. We further build an RL pipeline from evolutionary trajectories and fine-tune Qwen2.5 with GRPO. Experiments on single- and multi-objective benchmarks show improved relation prediction and state-of-the-art optimization performance over strong SAEA baselines and general LLMs. Quantization also enables efficient edge deployment, supporting a zero-shot surrogate paradigm without per-generation retraining. Code and models are available at https://github.com/Septend9/R2SAEA.

cs.NE

Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction

The transition from monolithic large language models (LLMs) to modular, skill-equipped agents represents a fundamental architectural shift in artificial intelligence deployment. While general-purpose models demonstrate remarkable breadth in declarative knowledge, their utility in autonomous workflows is frequently constrained by insufficient specialized procedural expertise. This report investigates a systematic framework for automated acquisition of high-quality agent skills through mining of open-source repositories on platforms such as GitHub. We focus on the extraction of visualization and educational capabilities from state-of-the-art systems including TheoremExplainAgent and Code2Video, both utilizing the Manim mathematical animation engine. The framework encompasses repository structural analysis, semantic skill identification through dense retrieval, and translation to the standardized SKILL.md format. We demonstrate that systematic extraction from agentic repositories, combined with rigorous security governance and multi-dimensional evaluation metrics, enables scalable acquisition of procedural knowledge that augments LLM capabilities without requiring model retraining. Our analysis reveals that agent-generated educational content can achieve 40\% gains in knowledge transfer efficiency while maintaining pedagogical quality comparable to human-crafted tutorials.

cs.AI

Scaling Laws for Educational AI Agents

While scaling laws for Large Language Models (LLMs) have been extensively studied along dimensions of model parameters, training data, and compute, the scaling behavior of LLM-based educational agents remains unexplored. We propose that educational agent capability scales not merely with the underlying model size, but through structured dimensions that we collectively term the Agent Scaling Law: role definition clarity, skill depth, tool completeness, runtime capability, and educator expertise injection. Central to this framework is AgentProfile, a structured JSON-based specification that serves as the mechanism enabling systematic capability growth of educational agents. We present EduClaw, a profile-driven multi-agent platform that operationalizes this scaling law, demonstrating its effectiveness through the construction and deployment of 330+ educational agent profiles encompassing 1,100+ skill modules across K-12 subjects. Our empirical observations suggest that educational agent performance scales predictably with profile structural richness. We identify two complementary scaling axes -- Tool Scaling and Skill Scaling -- as future directions, arguing that the path to more capable educational AI lies not solely in larger models, but in stronger structured capability systems.

cs.AI

An analytical-numerical coupled model of liquid droplet impact on solid material surfaces

Impacts of liquid droplets on wind turbine blade surfaces, for example sea sprays, can result in material damage through erosion. In this study, we derive an explicit, closed-form analytical approximation for droplet impact and subsequent spreading on a solid surface in inertia-dominated regimes of large Reynolds and Weber numbers. The formulation extends an existing theoretical framework based on inviscid potential flow for a rising expanding disk in an infinite liquid domain. The modified solution provides full spatio-temporal pressure distributions and impact force histories on the impact surface over the entire impact duration, capturing both the early-time self-similar flow and the inertia-driven lamella spreading following the peak impact force. The predicted pressure and force profiles show good agreement with analytical, numerical and experimental results reported in the literature, including accurate reproduction of the well-known ring-shaped pressure distribution. Key quantities, such as the radial location and magnitude of peak pressure, as well as the timing and magnitude of the peak impact force, are predicted analytically with reasonable accuracy. To enable solid material erosion analysis, the analytical liquid-phase solution is coupled with a finite-element (FE) simulation for the solid response. This analytical-numerical coupled method (ANCM) eliminates the need to explicitly simulate droplet fluid dynamics, which is conventionally performed using smoothed particle hydrodynamics (SPH). As a result, for the purpose of material response analysis, the proposed approach achieves grid independence at substantially lower mesh resolutions and reduces computational cost by more than 97% compared to SPH-based simulations, while maintaining or improving numerical accuracy.

physics.flu-dyn

Evaluation of the performance of an analytical-numerical coupled method for droplet impacts on soft material surfaces

Impacts between droplets and solid surfaces can commonly cause erosion problem in Engineering applications, including aircraft surface erosion, wind blade leading-edge erosion and steam turbine blade erosion. In practice, the impacted solid surfaces have varied material softness, ranging from stiff metallic coatings to soft materials. An analytical-numerical coupled model (ANCM) for simulating droplet impacts on surfaces, and corresponding material analysis, has been developed in the literature. However, the analytical impact pressure solution of the ANCM model has been derived assuming rigid solid surface. In the current study, we investigate the performance of the ANCM model for droplet impacts on soft materials made of urethane gel phantom, by comparing the ANCM computations to lab-based experiments and numerical simulations based on Smoothed Particle Hydrodynamics (SPH). Parametric studies explore the applicability limit of the ANCM model for droplet impacts on very soft materials at low Young's modulus. It was found that for materials at Young's modulus of $47,400$ Pa or stiffer, which covers most engineering applications, the developed ANCM model performs as expected for an assumed rigid surface. For softer solid materials, the SPH-modeled liquid interacts with the evolving surface geometry and mitigates impact intensity as deformation occurs. The analytical impact loads estimated by ANCM are independent of surface geometry, and hence provide conserved impact impulse in a non-physical way. Results show a critical value of Young's modulus at $E=10,000$ pa for the ANCM model, below which the model exhibits overshoot in total contact force and surface deformation, leading to the formation of steep wall craters.

physics.flu-dyn

See and Remember: A Multimodal Agent for Web Traversal

Autonomous web navigation requires agents to perceive complex visual environments and maintain long-term context, yet current Large Language Model (LLM) based agents often struggle with spatial disorientation and navigation loops. In this paper, we propose generally applicable V-GEMS(Visual Grounding and Explicit Memory System), a robust multimodal agent architecture designed for precise and resilient web traversal. Our agent integrates visual grounding to resolve ambiguous interactive elements and introduces an explicit memory stack with state tracking. This dual mechanism allows the agent to maintain a structured map of its traversal path, enabling valid backtracking and preventing cyclical failures in deep navigation tasks. We also introduce an updatable dynamic benchmark to rigorously evaluate adaptability. Experiments show V-GEMS significantly dominates the WebWalker baseline, achieving a substantial 28.7% performance gain. Code is available at https://github.com/Vaultttttttttttt/V-GEMS.

cs.AI

EduResearchBench: A Hierarchical Atomic Task Decomposition Benchmark for Full-Lifecycle Educational Research

While Large Language Models (LLMs) are reshaping the paradigm of AI for Social Science (AI4SS), rigorously evaluating their capabilities in scholarly writing remains a major challenge. Existing benchmarks largely emphasize single-shot, monolithic generation and thus lack the fine-grained assessments required to reflect complex academic research workflows. To fill this gap, we introduce EduResearchBench, the first comprehensive evaluation platform dedicated to educational academic writing. EduResearchBench is built upon our Hierarchical Atomic Task Decomposition (HATD) framework, which decomposes an end-to-end research workflow into six specialized research modules (e.g., Quantitative Analysis, Qualitative Research, and Policy Research) spanning 24 fine-grained atomic tasks. This taxonomy enables an automated evaluation pipeline that mitigates a key limitation of holistic scoring, where aggregate scores often obscure specific capability bottlenecks, and instead provides fine-grained, diagnostic feedback on concrete deficiencies. Moreover, recognizing the high cognitive load inherent in scholarly writing, we propose a curriculum learning strategy that progressively builds competence from foundational skills to complex methodological reasoning and argumentation. Leveraging 55K raw academic samples, we curate 11K high-quality instruction pairs to train EduWrite, a specialized educational scholarly writing model. Experiments show that EduWrite (30B) substantially outperforms larger general-purpose models (72B) on multiple core metrics, demonstrating that in vertical domains, data quality density and hierarchically staged training curricula are more decisive than parameter scale.

cs.CL

IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

Learning Path Recommendation (LPR) aims to generate personalized sequences of learning items that maximize long-term learning effect while respecting pedagogical principles and operational constraints. Although large language models (LLMs) offer rich semantic understanding for free-form recommendation, applying them to long-horizon LPR is challenging due to (i) misalignment with pedagogical objectives such as the Zone of Proximal Development (ZPD) under sparse, delayed feedback, (ii) scarce and costly expert demonstrations, and (iii) multi-objective interactions among learning effect, difficulty scheduling, length controllability, and trajectory diversity. To address these issues, we propose IB-GRPO (Indicator-Based Group Relative Policy Optimization), an indicator-guided alignment approach for LLM-based LPR. To mitigate data scarcity, we construct hybrid expert demonstrations via Genetic Algorithm search and teacher RL agents and warm-start the LLM with supervised fine-tuning. Building on this warm-start, we design a within-session ZPD alignment score for difficulty scheduling. IB-GRPO then uses the $I_{ε+}$ dominance indicator to compute group-relative advantages over multiple objectives, avoiding manual scalarization and improving Pareto trade-offs. Experiments on ASSIST09 and Junyi using the KES simulator with a Qwen2.5-7B backbone show consistent improvements over representative RL and LLM baselines.

cs.AI

On-Chip Generation of Co-Polarized and Spectrally Separable Photon Pairs

On-chip generation of high-purity single photons is essential for scalable photonic quantum technologies. Spontaneous parametric down-conversion (SPDC) is widely used to generate photon pairs for heralded single-photon sources, but intrinsic spectral correlations of the pairs often limit the purity and interference visibility of the heralded photons. Existing approaches to suppress these correlations rely on narrowband spectral filtering, which introduces loss, or exploiting different polarizations, which complicates on-chip integration. Here, we demonstrate a new strategy for generating spectrally separable photon pairs in thin-film lithium niobate nanophotonic circuits by harnessing higher-order spatial modes, with all interacting fields residing in the same polarization. Spectral separability is achieved by engineering group-velocity matching using higher-order transverse-electric modes, combined with a Gaussian-apodized poling profile to further suppress residual correlations inherent to standard periodic poling. Subsequent on-chip mode conversion with efficiency exceeding 95\% maps the higher-order mode to the fundamental mode and routes the photons into distinct output channels. The resulting heralded photons exhibit spectral purities exceeding 94\% inferred from joint-spectral intensity and 89\% from unheralded $g^{(2)}$ measurement. This approach enables flexible spectral and temporal engineering of on-chip quantum light sources for quantum computing and quantum networking.

quant-ph

Integrated polarization-entangled photon source for wavelength-multiplexed quantum networks

Entangled photons are fundamental resources for quantum communication, computing, and networking. Among them, polarization-entangled photon pairs play an important role due to their straightforward state manipulation and direct use in quantum key distribution, teleportation, and network protocols. However, realizing compact, efficient, and scalable polarization-entangled sources that meet the requirements of practical deployment remains a major challenge. Here, we present a simple yet high-performance on-chip polarization-entangled photon-pair source on thin-film lithium niobate (TFLN). Our device employs dual quasi-phase matching (D-QPM) that sequentially supports type-0 and type-I spontaneous parametric down-conversion in a single nanophotonic waveguide, eliminating the need for interferometers, polarization rotators, or other complex circuits. The source directly produces high-fidelity Bell states with broad bandwidth, high brightness, and low noise. Using this integrated platform, we realize wavelength-multiplexed entanglement distribution in a four-user quantum network deployed over metropolitan fiber links up to 50 km. These results establish a robust and scalable pathway toward practical quantum communication systems and multi-user quantum mesh networks based on integrated photonics.

physics.optics

AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search

Supervised fine-tuning (SFT) of large language models (LLMs) for specialized tasks requires high-quality datasets, but manual curation is prohibitively expensive. Synthetic data generation offers scalability, but its effectiveness relies on complex, multi-stage workflows, integrating prompt engineering and model orchestration. Existing automated workflow methods face a cold start problem: they require labeled datasets for reward modeling, which is especially problematic for subjective, open-ended tasks with no objective ground truth. We introduce AutoSynth, a framework that automates workflow discovery and optimization without reference datasets by reframing the problem as a Monte Carlo Tree Search guided by a novel dataset-free hybrid reward. This reward enables meta-learning through two LLM-as-judge components: one evaluates sample quality using dynamically generated task-specific metrics, and another assesses workflow code and prompt quality. Experiments on subjective educational tasks show that while expert-designed workflows achieve higher human preference rates (96-99% win rates vs. AutoSynth's 40-51%), models trained on AutoSynth-generated data dramatically outperform baselines (40-51% vs. 2-5%) and match or surpass expert workflows on certain metrics, suggesting discovery of quality dimensions beyond human intuition. These results are achieved while reducing human effort from 5-7 hours to just 30 minutes (>90% reduction). AutoSynth tackles the cold start issue in data-centric AI, offering a scalable, cost-effective method for subjective LLM tasks. Code: https://github.com/bisz9918-maker/AutoSynth.

cs.LG

Squeezed Light Generation in Periodically Poled Thin-Film Lithium Niobate Waveguides

Squeezed states of light play a key role in quantum-enhanced sensing and continuous-variable quantum information processing. Realizing integrated squeezed light sources is crucial for developing compact and scalable photonic quantum systems. In this work, we demonstrate on-chip broadband vacuum squeezing at telecommunication wavelengths on the thin-film lithium niobate (TFLN) platform. Our device integrates periodically poled lithium niobate (PPLN) nanophotonic waveguides with low-loss edge couplers, comprising bilayer inverse tapers and an SU-8 polymer waveguide. This configuration achieves a fiber-to-chip coupling loss of 1.4 dB and a total homodyne detection loss of 4 dB, enabling a measured squeezing level of 1.4 dB. Additional measurements in a more efficient PPLN waveguide (without low-loss couplers) infer an on-chip squeezing level of over 10 dB at a pump power of 62 mW. These results underscore the potential of TFLN platform for efficient and scalable squeezed light generation.

physics.optics

EA4LLM: A Gradient-Free Approach to Large Language Model Optimization via Evolutionary Algorithms

In recent years, large language models (LLMs) have made remarkable progress, with model optimization primarily relying on gradient-based optimizers such as Adam. However, these gradient-based methods impose stringent hardware requirements, demanding high-concurrency, high-memory GPUs. Moreover, they require all neural network operations to be differentiable, thereby excluding many promising non-differentiable architectures from practical use. To address these limitations, we propose EA4LLM, an evolutionary algorithm for optimizing LLMs, and, for the first time, empirically verify full-parameter optimization from the pretraining stage across model sizes ranging from 0.5B to 32B. We conduct extensive experiments and provide key insights into how evolutionary algorithms can effectively optimize neural networks. Our work challenges the prevailing assumption that gradient-based optimization is the only viable approach for training neural networks. It also holds significant potential to reduce the computational cost of training large language models, thereby enabling groups with limited computational resources to participate in deep learning research.

cs.AI

Lyte Quorum: Off-Chain Ready Smart Contract Hosted with Choice

This paper introduces Lyquor, a decentralized platform that reimagines blockchain infrastructure through a service-centric model where nodes selectively host smart contracts (called Lyquids) while preserving global composability. We present three key innovations: (1) Fate-Constrained Ordering (FCO), which decouples consensus from execution to enable selective hosting without sacrificing Layer-1 grade composability; (2) Direct Memory Architecture (DMA), which eliminates state access bottlenecks by providing each contract with persistent, byte-addressable virtual memory; and (3) Universal Procedure Call (UPC), which enables fault-tolerant, programmable coordination across distributed off-chain computation. Together, these components are powered by a Rust-macroed unified programming model where on-chain and off-chain logic coexist seamlessly, supporting both traditional smart contract patterns and novel distributed applications. Lyquor addresses critical limitations in existing systems while maintaining compatibility with Ethereum APIs, offering a path toward truly scalable decentralized computation.

cs.DC

Robust Paths: Geometry and Computation

Applying robust optimization often requires selecting an appropriate uncertainty set both in shape and size, a choice that directly affects the trade-off between average-case and worst-case performances. In practice, this calibration is usually done via trial-and-error: solving the robust optimization problem many times with different uncertainty set shapes and sizes, and examining their performance trade-off. This process is computationally expensive and ad hoc. In this work, we take a principled approach to study this issue for robust optimization problems with linear objective functions, convex feasible regions, and convex uncertainty sets. We introduce and study what we define as the robust path: a set of robust solutions obtained by varying the uncertainty set's parameters. Our central geometric insight is that a robust path can be characterized as a Bregman projection of a curve (whose geometry is defined by the uncertainty set) onto the feasible region. This leads to a surprising discovery that the robust path can be approximated via the trajectories of standard optimization algorithms, such as the proximal point method, of the deterministic counterpart problem. We give a sharp approximation error bound and show it depends on the geometry of the feasible region and the uncertainty set. We also illustrate two special cases where the approximation error is zero: the feasible region is polyhedrally monotone (e.g., a simplex feasible region under an ellipsoidal uncertainty set), or the feasible region and the uncertainty set follow a dual relationship. We demonstrate the practical impact of this approach in two settings: portfolio optimization and adversarial deep learning.

math.OC

SID: Benchmarking Guided Instruction Capabilities in STEM Education with a Socratic Interdisciplinary Dialogues Dataset

Fostering students' abilities for knowledge integration and transfer in complex problem-solving scenarios is a core objective of modern education, and interdisciplinary STEM is a key pathway to achieve this, yet it requires expert guidance that is difficult to scale. While LLMs offer potential in this regard, their true capability for guided instruction remains unclear due to the lack of an effective evaluation benchmark. To address this, we introduce SID, the first benchmark designed to systematically evaluate the higher-order guidance capabilities of LLMs in multi-turn, interdisciplinary Socratic dialogues. Our contributions include a large-scale dataset of 10,000 dialogue turns across 48 complex STEM projects, a novel annotation schema for capturing deep pedagogical features, and a new suite of evaluation metrics (e.g., X-SRG). Baseline experiments confirm that even state-of-the-art LLMs struggle to execute effective guided dialogues that lead students to achieve knowledge integration and transfer. This highlights the critical value of our benchmark in driving the development of more pedagogically-aware LLMs.

cs.AI