Searcharxiv⌕ Search

arXiv subjects

Peng He

Publications and source records attributed to Peng He.

At least 37 records · Page 2Linked to original sources

How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment

The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptability to diverse question types and flexibility in output formats, they also introduce new challenges related to output uncertainty, stemming from the inherently probabilistic nature of LLMs. Output uncertainty is an inescapable challenge in automatic assessment, as assessment results often play a critical role in informing subsequent pedagogical actions, such as providing feedback to students or guiding instructional decisions. Unreliable or poorly calibrated uncertainty estimates can lead to unstable downstream interventions, potentially disrupting students' learning processes and resulting in unintended negative consequences. To systematically understand this challenge and inform future research, we benchmark a broad range of uncertainty quantification methods in the context of LLM-based automatic assessment. Although the effectiveness of these methods has been demonstrated in many tasks across other domains, their applicability and reliability in educational settings, particularly for automatic grading, remain underexplored. Through comprehensive analyses of uncertainty behaviors across multiple assessment datasets, LLM families, and generation control settings, we characterize the uncertainty patterns exhibited by LLMs in grading scenarios. Based on these findings, we evaluate the strengths and limitations of different uncertainty metrics and analyze the influence of key factors, including model families, assessment tasks, and decoding strategies, on uncertainty estimates. Our study provides actionable insights into the characteristics of uncertainty in LLM-based automatic assessment and lays the groundwork for developing more reliable and effective uncertainty-aware grading systems in the future.

cs.AI↗

Negative-Aware Diffusion Process for Temporal Knowledge Graph Extrapolation

Temporal Knowledge Graph (TKG) reasoning seeks to predict future missing facts from historical evidence. While diffusion models (DM) have recently gained attention for their ability to capture complex predictive distributions, two gaps remain: (i) the generative path is conditioned only on positive evidence, overlooking informative negative context, and (ii) training objectives are dominated by cross-entropy ranking, which improves candidate ordering but provides little supervision over the calibration of the denoised embedding. To bridge this gap, we introduce Negative-Aware Diffusion model for TKG Extrapolation (NADEx). Specifically, NADEx encodes subject-centric histories of entities, relations and temporal intervals into sequential embeddings. NADEx perturbs the query object in the forward process and reconstructs it in reverse with a Transformer denoiser conditioned on the temporal-relational context. We further derive a cosine-alignment regularizer derived from batch-wise negative prototypes, which tightens the decision boundary against implausible candidates. Comprehensive experiments on four public TKG benchmarks demonstrate that NADEx delivers state-of-the-art performance.

cs.AI↗

Multi-View Adaptive Contrastive Learning for Information Retrieval Based Fault Localization

Most studies focused on information retrieval-based techniques for fault localization, which built representations for bug reports and source code files and matched their semantic vectors through similarity measurement. However, such approaches often ignore some useful information that might help improve localization performance, such as 1) the interaction relationship between bug reports and source code files; 2) the similarity relationship between bug reports; and 3) the co-citation relationship between source code files. In this paper, we propose a novel approach named Multi-View Adaptive Contrastive Learning for Information Retrieval Fault Localization (MACL-IRFL) to learn the above-mentioned relationships for software fault localization. Specifically, we first generate data augmentations from report-code interaction view, report-report similarity view and code-code co-citation view separately, and adopt graph neural network to aggregate the information of bug reports or source code files from the three views in the embedding process. Moreover, we perform contrastive learning across these views. Our design of contrastive learning task will force the bug report representations to encode information shared by report-report and report-code views,and the source code file representations shared by code-code and report-code views, thereby alleviating the noise from auxiliary information. Finally, to evaluate the performance of our approach, we conduct extensive experiments on five open-source Java projects. The results show that our model can improve over the best baseline up to 28.93%, 25.57% and 20.35% on Accuracy@1, MAP and MRR, respectively.

cs.SE↗

DrawSim-PD: Simulating Student Science Drawings to Support NGSS-Aligned Teacher Diagnostic Reasoning

Developing expertise in diagnostic reasoning requires practice with diverse student artifacts, yet privacy regulations prohibit sharing authentic student work for teacher professional development (PD) at scale. We present DrawSim-PD, the first generative framework that simulates NGSS-aligned, student-like science drawings exhibiting controllable pedagogical imperfections to support teacher training. Central to our approach are apability profiles--structured cognitive states encoding what students at each performance level can and cannot yet demonstrate. These profiles ensure cross-modal coherence across generated outputs: (i) a student-like drawing, (ii) a first-person reasoning narrative, and (iii) a teacher-facing diagnostic concept map. Using 100 curated NGSS topics spanning K-12, we construct a corpus of 10,000 systematically structured artifacts. Through an expert-based feasibility evaluation, K--12 science educators verified the artifacts' alignment with NGSS expectations (>84% positive on core items) and utility for interpreting student thinking, while identifying refinement opportunities for grade-band extremes. We release this open infrastructure to overcome data scarcity barriers in visual assessment research.

cs.CY↗

Floquet quantum geometry in periodically driven topological insulators

Quantum geometry plays a fundamental role across many branches of modern physics, yet its full characterization in nonequilibrium systems remains a challenge. Here, we propose a framework for quantum geometry in Floquet topological insulators by introducing a time-resolved quantum metric tensor, defined via the trace distance between micromotion operators in momentum-time space. For class A in two spatial dimensions, we find a general inequality linking the Floquet quantum metric tensor and the Floquet topology: the associated quantum volume is bounded below by the Floquet topological invariant. This relation is found to also hold in class AIII in one dimension, where the Floquet geometric tensor may be notably reduced due to time-reflection symmetry. This work will be useful in digesting the general aspects of quantum geometry in periodically driven systems in connection with their topological characterization.

cond-mat.mes-hall↗

Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials

Designing high-quality, standards-aligned instructional materials for K--12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns in where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K--12 science education.

cs.CY↗

Constructing left-continuous triangular norms on complete lattices

This article focuses on the construction of left-continuous t-norms on complete lattices. The concepts of $\mathfrak{f}$-mappings and weak $\mathfrak{f}$-mappings on complete lattices are first introduced, respectively. They are then applied to establish the following key results: weak $\mathfrak{f}$-mappings are used to induce left-continuous t-subnorms; $\mathfrak{f}$-mappings are used to generate left-continuous t-norms whenever the top element $1$ of the complete lattice is a completely join-irreducible element. Finally, some necessary and sufficient conditions are provided for an operator constructed by the ordinal sum of a series of annihilating binary operators being a left-continuous t-norm on a complete lattice.

math.GM↗

HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation

We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion Transformer (DiT)-based flow matching models to the billion-parameter scale within the motion generation domain, delivering instruction-following capabilities that significantly outperform current open-source benchmarks. Uniquely, we introduce a comprehensive, full-stage training paradigm -- including large-scale pretraining on over 3,000 hours of motion data, high-quality fine-tuning on 400 hours of curated data, and reinforcement learning from both human feedback and reward models -- to ensure precise alignment with the text instruction and high motion quality. This framework is supported by our meticulous data processing pipeline, which performs rigorous motion cleaning and captioning. Consequently, our model achieves the most extensive coverage, spanning over 200 motion categories across 6 major classes. We release HY-Motion 1.0 to the open-source community to foster future research and accelerate the transition of 3D human motion generation models towards commercial maturity.

cs.CV↗

Constraining the Hubble Constant with a Simulated Full Covariance Matrix Using Neural Networks

The Hubble parameter, $H(z)$, plays a crucial role in understanding the expansion history of the universe and constraining the Hubble constant, $\mathrm{H}_0$. The Cosmic Chronometers (CC) method provides an independent approach to measuring $H(z)$, but existing studies either neglect off-diagonal elements in the covariance matrix or use an incomplete covariance matrix, limiting the accuracy of $\mathrm{H}_0$ constraints. To address this, we use a Positive-Definite Covariance Network (PD-CovNet) to simulate the full $33 \times 33$ covariance matrix based on a previously published $15 \times 15$ covariance matrix. Hyperparameters are chosen via leave-one-z-out validation, and performance is benchmarked against a Gaussian-process (GP) baseline. Under identical five-fold cross-validation over redshift groups, we prove that PD-CovNet is a reliable generator of the full covariance compared to the GP baseline. Using this full PD-CovNet-simulated covariance alongside three comparators with different covariance specifications, we constrain $\mathrm{H}_0$ with two independent methods (EMCEE and GP). Across all covariance specifications and both constraint methods, standardized differences and two-sided p-values show no statistically meaningful shift in the central value of the constrained $\mathrm{H}_0$. However, the precision of the constrained $\mathrm{H}_0$ depends on both covariance and method: EMCEE is uniformly more precise than GP once covariance is modeled; within a fixed method, incorporating more covariance reduces precision; and PD-CovNet hyperparameters have a modest effect on uncertainty. These results indicate the importance of accurate covariance modeling in CC-based $\mathrm{H}_0$ constraints.

astro-ph.CO↗

Exploiting Inter-Session Information with Frequency-enhanced Dual-Path Networks for Sequential Recommendation

Sequential recommendation (SR) aims to predict a user's next item preference by modeling historical interaction sequences. Recent advances often integrate frequency-domain modules to compensate for self-attention's low-pass nature by restoring the high-frequency signals critical for personalized recommendations. Nevertheless, existing frequency-aware solutions process each session in isolation and optimize exclusively with time-domain objectives. Consequently, they overlook cross-session spectral dependencies and fail to enforce alignment between predicted and actual spectral signatures, leaving valuable frequency information under-exploited. To this end, we propose FreqRec, a Frequency-Enhanced Dual-Path Network for sequential Recommendation that jointly captures inter-session and intra-session behaviors via a learnable Frequency-domain Multi-layer Perceptrons. Moreover, FreqRec is optimized under a composite objective that combines cross entropy with a frequency-domain consistency loss, explicitly aligning predicted and true spectral signatures. Extensive experiments on three benchmarks show that FreqRec surpasses strong baselines and remains robust under data sparsity and noisy-log conditions.

cs.IR↗

RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging

We unveil that internal representations in large language models (LLMs) serve as reliable proxies of learned knowledge, and propose RECALL, a novel representation-aware model merging framework for continual learning without access to historical data. RECALL computes inter-model similarity from layer-wise hidden representations over clustered typical samples, and performs adaptive, hierarchical parameter fusion to align knowledge across models. This design enables the preservation of domain-general features in shallow layers while allowing task-specific adaptation in deeper layers. Unlike prior methods that require task labels or incur performance trade-offs, RECALL achieves seamless multi-domain integration and strong resistance to catastrophic forgetting. Extensive experiments across five NLP tasks and multiple continual learning scenarios show that RECALL outperforms baselines in both knowledge retention and generalization, providing a scalable and data-free solution for evolving LLMs.

cs.CL↗

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV↗

Theoretical Radio Signals from Radio-Band Gravitational Waves Converted from the Neutron Star Magnetic Field

Gravitational waves (GWs) can convert into electromagnetic waves in the presence of a magnetic field via the Gertsenshtein-Zeldovich (GZ) effect. The characteristics of the magnetic field substantially affect this conversion probability. This paper confirms that strong magnetic fields in neutron stars significantly enhance the conversion probability, facilitating detectable radio signatures of very high-frequency (VHF, $\left(10^6-10^{11}\mathrm{~Hz}\right)$) gravitational waves. We theoretically identify two distinct signatures using single-dish telescopes (FAST, TMRT, QTT, GBT) and interferometers (SKA1/2-MID): transient signals from burst-like gravitational wave sources and persistent signals from cosmological background gravitational wave sources. These signatures are mapped to graviton spectral lines derived from quantum field theory by incorporating spin-2 and mass constraints, resulting in smooth, featureless profiles that are critical for distinguishing gravitational wave signals from astrophysical foregrounds. FAST attains a characteristic strain bound of $h_c<10^{-23}$, approaching $10^{-24}$ in the frequency range of $1-3\mathrm{~GHz}$ with a 6-hour observation period. This performance exceeds the $5 σ$ detection thresholds for GWs originating from primordial black holes (PBHs) and nears the limits set by Big Bang nucleosynthesis. Additionally, projections for SKA2-MID indicate even greater sensitivity. Detecting such gravitational waves would improve our comprehension of cosmological models, refine the parameter spaces for primordial black holes, and function as a test for quantum field theory. This approach addresses significant deficiencies in VHF GW research, improving detection sensitivity and facilitating the advancement of next-generation radio telescopes such as FASTA and SKA, which feature larger fields of view and enhanced gain.

astro-ph.HE↗

Estimating Cosmological Parameters and Reconstructing Hubble Constant with Artificial Neural Networks: A Test with covariance matrix and mock H(z)

In this work, we reconstruct the H(z) based on observational Hubble data with Artificial Neural Network, then estimate the cosmological parameters and the Hubble constant. The training data we used are covariance matrix and mock H(z), which are generated based on the real OHD data and Gaussian Process(GP). The use of the covariance matrix propagates the correlated uncertainties and improves training efficiency. Using the reconstructed H(z) data, we first determine the Hubble constant and compare it with CMB-based measurements. To constrain cosmological parameters, we sample on the reconstructed data and calculate the corresponding posterior distributions with Markov Chain Monte Carlo (MCMC). Through comprehensive statistical comparisons, we demonstrate that the parameter estimation using reconstructed samples achieves comparable statistical accuracy to the result derived from real OHD data.

astro-ph.CO↗

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360° immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360° world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

cs.CV↗

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation, the field remains largely accessible only to researchers, developers, and designers due to the complexities involved in collecting, processing, and training 3D models. To address these challenges, we introduce Hunyuan3D 2.1 as a case study in this tutorial. This tutorial offers a comprehensive, step-by-step guide on processing 3D data, training a 3D generative model, and evaluating its performance using Hunyuan3D 2.1, an advanced system for producing high-resolution, textured 3D assets. The system comprises two core components: the Hunyuan3D-DiT for shape generation and the Hunyuan3D-Paint for texture synthesis. We will explore the entire workflow, including data preparation, model architecture, training strategies, evaluation metrics, and deployment. By the conclusion of this tutorial, you will have the knowledge to finetune or develop a robust 3D generative model suitable for applications in gaming, virtual reality, and industrial design.

cs.CV↗

IssueCourier: Multi-Relational Heterogeneous Temporal Graph Neural Network for Open-Source Issue Assignment

Issue assignment plays a critical role in open-source software (OSS) maintenance, which involves recommending the most suitable developers to address the reported issues. Given the high volume of issue reports in large-scale projects, manually assigning issues is tedious and costly. Previous studies have proposed automated issue assignment approaches that primarily focus on modeling issue report textual information, developers' expertise, or interactions between issues and developers based on historical issue-fixing records. However, these approaches often suffer from performance limitations due to the presence of incorrect and missing labels in OSS datasets, as well as the long tail of developer contributions and the changes of developer activity as the project evolves. To address these challenges, we propose IssueCourier, a novel Multi-Relational Heterogeneous Temporal Graph Neural Network approach for issue assignment. Specifically, we formalize five key relationships among issues, developers, and source code files to construct a heterogeneous graph. Then, we further adopt a temporal slicing technique that partitions the graph into a sequence of time-based subgraphs to learn stage-specific patterns. Furthermore, we provide a benchmark dataset with relabeled ground truth to address the problem of incorrect and missing labels in existing OSS datasets. Finally, to evaluate the performance of IssueCourier, we conduct extensive experiments on our benchmark dataset. The results show that IssueCourier can improve over the best baseline up to 45.49% in top-1 and 31.97% in MRR.

cs.SE↗

Left-continuous pseudo-t-norms on modular lattices

This article focuses on the relationship between pseudo-t-norms and the structure of lattices. First, we establish a necessary and sufficient condition for the existence of a left-continuous t-norm on the ordinal sum of two disjoint complete lattices. Then, we define the $1$-distributivity of a lattice, which is applied for characterizing a complete atomistic lattice that has a left-continuous pseudo-t-norm. We also describe the forbidden structures of a finite modular lattice that is a $1$-distributive lattice, which is used for representing a kind of finite planar modular lattices that have left-continuous pseudo-t-norms.

math.RT↗