SearcharxivSearch

arXiv subjects

Hong Yang

Publications and source records attributed to Hong Yang.

At least 19 recordsLinked to original sources

Spectral Approximation and Ergodic-Capacity Convergence of HMIMO Channels under Spatial-Wavenumber Domain Mismatch

We establish quantitative results on finite-dimensional spectral approximation and ergodic-capacity convergence for continuous Holographic Multiple-Input Multiple-Output (HMIMO) channels with square apertures and physically prescribed circular wavenumber support. The resulting spatial-wavenumber domain mismatch leads to a non-separable square-disk concentration problem for which the classical separable construction based on prolate spheroidal wave functions (PSWFs) cannot be directly applied. We project the continuous operator onto a tensor-product subspace of one-dimensional (1D) PSWFs while preserving the circular wavenumber support, yielding a generally non-diagonal but highly sparse finite-dimensional matrix. We show that the whole-spectrum approximation error, accounting for retained-eigenvalue perturbations and the residual spectral tail, remains controlled by a 1D PSWF eigenvalue-tail envelope despite the loss of separability and induced off-diagonal coupling. Beyond an explicit 1D truncation threshold, this error decays super-exponentially. This analysis further yields an asymptotic upper envelope for the eigenspectrum under the flattened two-dimensional eigenvalue ordering. We further establish a non-asymptotic upper bound on the gap between the actual ergodic capacities of the continuous and tensor-PSWF-truncated channels under their respective transmit-covariance optimizations. Combined with the spectral result, this capacity-gap bound inherits the same super-exponential dependence on the truncation order. Finally, quadrature rules with explicit radial and angular node thresholds are developed for evaluating the projected matrix. Numerical results show that conventional truncation based on spatial degrees of freedom can omit performance-relevant modes, particularly for compact apertures.

cs.IT

Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations. This setting remains underexplored, with limited benchmarks and few methods designed for the privacy, efficiency, and generalization demands of real-world deployment. We introduce a novel hybrid benchmark combining generative CCTV-style videos with real-world classroom pose data, and propose a lightweight, but robust motion-reasoning framework motivated by the observation that many incidents differ more in motion direction, speed, acceleration, and intensity than in pose alone. To that end, our method first constructs hierarchical kinematic representations of human actions. Our method then distills hierarchical, multi-order kinematic reasoning from a large teacher into a much smaller single-order student, enabling efficient per-person inference while preserving expressive motion understanding. Experiments show that our model outperforms substantially larger baselines at less than one-tenth of their computational cost, while also demonstrating stronger out-of-domain motion reasoning and zero-shot synthetic-to-real generalization. We will publicly release the benchmark, codebase, and supporting tools to facilitate further research in privacy-aware classroom safety.

cs.CV

Computational and Effective Degrees of Freedom for Spatially Stationary HMIMO Channel Modeling

This paper establishes a comprehensive theoretical framework for the continuous-to-discrete modeling of spatially stationary holographic MIMO (HMIMO) channels utilizing the Nystrom method with Gauss-Legendre quadrature (NGLQ). Starting with an operator-theoretic analysis of the NGLQ method, we prove that its quadrature error exhibits a super-exponential decay. Furthermore, we derive a spatial sampling threshold, termed computational degrees of freedom (cDoF), which reveals a {\pi}/2 oversampling penalty over the physical DoF for 1D arrays, compounding to a 68% computational redundancy for 2D separable grids. To address the ill-conditioning of the eigenvalue decomposition (EVD) problem inherent to the Nystrom discretization, we invoke the multidimensional Szego-Widom asymptotic expansion. This analysis yields a physically grounded semi-analytical expression for the effective DoF (eDoF) of 2D rectangular apertures, capturing the anisotropic boundary truncation effects to guide partial EVD and reduce computational complexity. Numerical evaluations confirm the tightness of the cDoF threshold under worst-case end-fire conditions. Moreover, simulations utilizing closed-form kernels for isotropic scattering verify that the derived eDoF acts as an accurate asymptotic approximation. Finally, by deploying the exact non-uniform discrete Fourier transform to eliminate interpolation error floors, we demonstrate spectral convergence down to the machine-precision level for non-isotropic scattering environments.

eess.SP

Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks

User modeling is a critical task in a variety of personalized systems. Recognizing their effectiveness in learning from graph-structured data, Graph Neural Networks (GNNs), particularly Graph Convolutional Networks (GCNs), are increasingly employed for user modeling. However, existing approaches typically treat different relation types in a graph as homogeneous, limiting their ability to capture richer semantics and construct more informative user models. While multi-relational GNNs (MR-GNNs) have been adopted for representation learning and recommendation, their application for user modeling remains unexplored. Moreover, existing GNN-based user modeling approaches ignore the user interaction sequence. To address these research gaps, in this work we propose MR-ConceptGCN, a novel fully unsupervised approach focused on concept-based sequential learner modeling using multi-relational GCNs (MR-GCNs). MR-ConceptGCN effecively combines Personal Knowledge Graphs (PKGs), MR-GCNs, and the pre-trained language model SBERT to obtain enhanced relation- and semantic-aware representations of the PKG items. The enriched embeddings of the knowledge concepts that a learner did not understand when interacting with learning materials in CourseMapper are then used to construct a sequential learner model that combines long-term and short-term learner interactions. We report the results of an online user study (n = 31), demonstrating the benefits of MR-ConceptGCN in terms of several important user-centric aspects including accuracy, usefulness, diversity, and satisfaction with an educational recommender system.

cs.AI

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation

Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters. We run the missing crossed comparison: an integrity-clean three-arm ablation (baseline / bash_only / code_only) on synthetic computation tasks and SWE-bench Mini modification tasks, holding model, harness, and prompts fixed, with two agents (Claude Code, OpenAI Codex CLI) so the comparison spans both regime and agent-design axes. Across the four resulting (regime, agent) cells, restricting the agent to a single execute_code MCP tool is cheaper than -- or statistically tied with -- its cheapest tool-rich rival in three cells (significantly on Artifact/Claude and SWE-bench/Codex; directionally on Artifact/Codex), with pass rates statistically tied within each cell. The lone exception is SWE-bench/Claude, where code_only is directionally costlier (+14.4%, not significant); a conditional-cost analysis localizes that gap to failure-cost on doomed-run trajectories, not a per-edit tax on successful runs. Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes we evaluate. The benchmark harness, task suite, and analysis code are available at https://github.com/hyang0129/onlycodes.

cs.SE

List-Decoding Counterexamples Yield Lower Bounds on Mutual Correlated Agreement Error

Mutual correlated agreement captures whether a random linear combination of received words can create a new large agreement with a code, a property relevant to the soundness of batched proximity testing. We show constructively that list-decoding counterexamples yield lower bounds on the mutual correlated agreement error. Given an explicit counterexample to the $(p,L)$-list-decodability of a linear code over $\mathbb{F}_q$, we construct a related code $C'$ of the same length and dimension such that $\operatorname{err}_{\mathrm{MCA}}(C',p)\ge\frac{1}{q}\left\lceil\frac{(L+1)q}{q+L}\right\rceil$, while decreasing its minimum distance by at most one. The construction also produces an explicit pair of words witnessing this error. We further give a structure-preserving version for code families whose coordinates are indexed by a finite set $\Omega$, with each index determining a generator-matrix column through a map $v:\Omega\to\mathbb{F}_q^k$. The construction changes at most one coordinate index and ensures that the output code remains in the same indexed family. As applications, we instantiate this principle for algebraic-geometry (AG) evaluation codes and Reed--Solomon codes. For AG codes, if $G$ is the divisor defining the underlying Riemann--Roch space and $N$ is the number of rational places outside $\operatorname{supp}(G)$ available for evaluation, the resulting code remains over the same function field and Riemann--Roch space, with a modified set of evaluation places. Its mutual correlated agreement error is at least $\frac{1}{q}\left\lceil\frac{(L+1)N}{N+L\mathrm{deg} G}\right\rceil$. The Reed--Solomon conclusion follows as the Vandermonde-column specialization.

cs.IT

ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations. Yet existing visual and embodied question answering benchmarks often provide limited control over the reasoning dependencies being tested, making it difficult to distinguish grounded embodied reasoning from shortcut-driven visual or linguistic pattern matching. We present ERQA-Plus, a diagnostic benchmark for reasoning in embodied AI. ERQA-Plus contains 1,766 question-answer instances grounded in 711 robot-centric images and organized according to a structured taxonomy spanning perceptual, action-centric, social-interaction, navigation-environmental, and contextual commonsense reasoning. The dataset is constructed using a multi-stage generation and validation pipeline that combines taxonomy-guided question generation, automatic quality judging, iterative revision, and human assessment to improve visual grounding, answer validity, and reasoning quality. We benchmark representative general-purpose vision-language models and embodied models, including LLaVA-NeXT-8B, Prismatic-7B, MiniCPM-V-4.5-8B, Qwen3-VL, RoboRefer-8B, and RoboBrain2.5-8B. Although the strongest model, Qwen3-VL-32B, achieves 83.4% overall accuracy and 61.4 SBERT score, category-level results reveal persistent weaknesses in spatial reasoning, procedural reasoning, event prediction, and intention inference. ERQA-Plus therefore provides a fine-grained evaluation framework for measuring not only whether embodied agents answer correctly, but also which forms of embodied reasoning they can and cannot perform reliably. The dataset is available https://huggingface.co/datasets/huggingdas/erqa-plus and the project page at https://github.com/LUNAProject22/erqa-plus.

cs.RO

IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction

Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in areas that lack applicable laws. We propose the term IUU+ to capture a broader suite of fisheries sector environmental and associated supply chain trade-related crimes and behaviors. Although IUU+ activity is widely recognized as a serious threat to marine ecosystems, markets, and livelihoods, a quantitative understanding of these incidents, e.g., their frequency, geography, species, actors, and patterns in the type of illicit activity, remains difficult to obtain. We propose IUU+DB, a large language model driven system for building a global incident database of IUU+ activity. The system ingests heterogeneous documents, classifies whether they describe relevant incidents, extracts key data elements such as actors, locations, species, vessels, violations, and enforcement outcomes, and supports deduplication and trend analysis. Case studies and validation results show that IUU+DB can help organize fragmented evidence, surface geographic and behavioral hotspots, support fisheries-domain specific research in academia and non-government organizations, assist source and species risk assessments for industry, and provide support for policy implementation and targeted enforcement efforts to government agencies.

cs.IR

Chiral Quantum Entanglement Transfer with Giant Atoms

We investigate entanglement transfer in a multi-giant-atom waveguide system. By tailoring chiral spontaneous emission and exploiting dark-state dynamics, the setup enables perfect, unidirectional sequential or selective transfer of quantum states and their associated entanglement. The distance between two entangled atoms, i.e., the entanglement length, can be dynamically adjusted, allowing robust conversion between long-range and short-range entanglement during propagation. The system inherently converges to a dark state, guaranteeing high-fidelity directional transfer. When the additional phase is modulated as a periodic piecewise function, spatially separated giant atoms exhibit stable, nearly lossless state exchange and maintain steady entanglement even under non-Markovian conditions. This behavior mimics conventional braided architectures without suffering from propagation delays or spatial restrictions. Our proposal offers a scalable pathway for continuous long-distance entanglement transport and resilient state exchange in quantum networks.

quant-ph

Superdirectivity as Boundary Concentration under Spectral Collision

Array superdirectivity is traditionally derived through singular optimization of densely spaced antenna arrays. In this paper, we show that the phenomenon admits a geometric interpretation as a concentration effect induced by spectral collision. As the spacing of an $M$-element linear array tends to zero, the exponential family generated by a linear array undergoes a spectral collision, and the associated finite-dimensional subspaces converge in reproducing kernel to a polynomial jet space. The maximum achievable array gain equals the diagonal evaluation of the reproducing kernel, and is therefore governed by the reciprocal Christoffel function. For the classical flat $L^2([-1,1])$ geometry, the Christoffel--Darboux kernel exhibits boundary concentration, yielding the quadratic $M^2$ superdirective law as a direct consequence of kernel asymptotics. This viewpoint separates intrinsic gain limits from numerical conditioning and identifies superdirectivity as a manifestation of a more general concentration mechanism. The framework further shows that the classical $M^2$ scaling is not universal: alternative spectral geometries produce different concentration laws through their associated Christoffel asymptotics. The results establish a direct connection between superdirectivity, reproducing kernels, orthogonal polynomials, and concentration phenomena arising from singular spectral limits.

cs.IT

Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle types in safety-critical interactions is essential for the safe deployment of automated driving technologies. This study extracts safety-critical pedestrian-vehicle interactions from the Argoverse 2 dataset to capture real-world crash avoidance behaviors in encounters involving AVs and HDVs. To model vehicle-type-specific pedestrian crash avoidance behavior, we develop a Smooth-Mamba Deep Deterministic Policy Gradient framework, termed SMamba-DDPG, which integrates smooth action constraints with efficient temporal representation learning. To quantify pedestrian behavioral differences, the framework trains separate crash avoidance policies for pedestrian interactions with AVs and HDVs. Results show that SMamba-DDPG outperforms baseline reinforcement learning and supervised learning models in reproducing pedestrian crash avoidance behaviors. Reconstructed trajectories demonstrate strong behavioral realism, accurately reproducing crash avoidance kinematics in both AV and HDV scenarios. Reaction time analysis shows that the model captures human-like response delays and reveals that pedestrians respond more quickly to AVs than to HDVs. Counterfactual analysis further indicates that pedestrians adopt lower crossing speeds when interacting with AVs. Large-scale safety analysis of model-generated data revealed that pedestrian-AV interactions consistently yielded lower conflict rates and higher pedestrian yielding rates compared to pedestrian-HDV interactions. The findings highlight the importance of incorporating vehicle-type-specific pedestrian behavioral models for safer automated driving system design and more realistic traffic simulations in mixed-traffic environments.

cs.AI

Robust Multi-Stream Massive MIMO Satellite Systems Based on Statistical CSI

This paper investigates multi-stream downlink precoding for massive multiple-input multiple-output low-Earthorbit satellite (SAT) communication systems. We adopt a delay and Doppler precompensation approach to achieve coherent transmission. Under this setting, we formulate a signal transmission model that incorporates the near-independent properties of inter-SAT interference and compensation errors. We then demonstrate that moving beyond single-stream transmission requires both multi-SAT cooperation and multi-antenna UTs. Based on this configuration and the established signal transmission model, we derive the first- and second-order statistical channel characteristics and utilize them to design locally optimal precoding algorithms for both total power constraint (TPC) and per-antenna power constraint (PAPC) conditions, which rely only on statistical channel state information (sCSI). In particular, the designed PAPC algorithm achieves linear complexity with respect to the number of antennas on the cooperative SATs. To reduce the computational complexity of the locally optimal precoder under TPC, we propose a low-complexity and robust precoding scheme optimized for both minimum mean squared error and sum-rate maximization objectives. Using majorization theory, we also provide a rigorous theoretical analysis of the optimal precoding structure under TPC. Moreover, the Lanczos algorithm is adopted to further reduce the complexity of the proposed robust designs. Simulation results show that when each SAT is equipped with a sufficiently large number of antennas, the proposed sCSI-based designs achieve performance comparable to that of instantaneous CSI-based designs.

eess.SP

Controlling entanglement by phase engineering in giant-atom waveguide

We investigate the entanglement dynamics of two giant atoms coupled to a common waveguide. By introducing additional phase modulation at each coupling point, every photon propagation path is jointly controlled by two distinct coupling phases, enabling precise and flexible manipulation of the entanglement evolution. This phase engineering induces destructive interference among different paths, leading to entanglement dynamics in nested giant atoms that become equivalent to those of small atoms, as well as dynamical equivalence between separated and braided configurations. Furthermore, the proposed scheme significantly enhances the robustness of entanglement against variations in the phase shift, offering a practical route to generate stable entanglement and enabling quantum devices with programmable propagation and controllable memory effects.

quant-ph

Spin crossover in FeO under shock compression

FeO (w\"ustite), which exhibits complex electronic and structural properties with increasing pressure and temperature, is a key mineralogical phase for understanding deep planetary interiors. However, direct measurements of its spin state at high-pressure and temperature remain challenging in static compression experiments. Here, we employ laser-driven shock compression to extend the FeO principal Hugoniot up to $\sim$900 GPa and perform in situ X-ray diffraction and X-ray emission spectroscopy up to 250 GPa, probing FeO's crystal structure and spin state. We demonstrate a continuous spin crossover of iron in FeO over a broad pressure range, with the high-spin state persisting beyond Earth's core-mantle boundary (CMB) conditions. These observations provide new experimental constraints on iron spin state at extreme conditions essential for geophysical models of (exo)planetary interiors.

cond-mat.mtrl-sci

Beyond the Class Subspace: Teacher-Guided Training for Reliable Out-of-Distribution Detection in Single-Domain Models

Out-of-distribution (OOD) detection methods perform well on multi-domain benchmarks, yet many practical systems are trained on single-domain data. We show that this regime induces a geometric failure mode, Domain-Sensitivity Collapse (DSC): supervised training compresses features into a low-rank class subspace and suppresses directions that carry domain-shift signal. We provide theory showing that, under DSC, distance- and logit-based OOD scores lose sensitivity to domain shift. We then introduce Teacher-Guided Training (TGT), which distills class-suppressed residual structure from a frozen multi-domain teacher (DINOv2) into the student during training. The teacher and auxiliary head are discarded after training, adding no inference overhead. Across eight single-domain benchmarks, TGT yields large far-OOD FPR@95 reductions for distance-based scorers: MDS improves by 11.61 pp, ViM by 10.78 pp, and kNN by 12.87 pp (ResNet-50 average), while maintaining or slightly improving in-domain OOD and classification accuracy.

cs.LG

A Vision-and-Knowledge Enhanced Large Language Model for Generalizable Pedestrian Crossing Behavior Inference

Existing paradigms for inferring pedestrian crossing behavior, ranging from statistical models to supervised learning methods, demonstrate limited generalizability and perform inadequately on new sites. Recent advances in Large Language Models (LLMs) offer a shift from numerical pattern fitting to semantic, context-aware behavioral reasoning, yet existing LLM applications lack domain-specific adaptation and visual context. This study introduces Pedestrian Crossing LLM (PedX-LLM), a vision-and-knowledge enhanced framework designed to transform pedestrian crossing inference from site-specific pattern recognition to generalizable behavioral reasoning. By integrating LLaVA-extracted visual features with textual data and transportation domain knowledge, PedX-LLM fine-tunes a LLaMA-2-7B foundation model via Low-Rank Adaptation (LoRA) to infer crossing decisions. PedX-LLM achieves 82.0% balanced accuracy, outperforming the best statistical and supervised learning methods. Results demonstrate that the vision-augmented module contributes a 2.9% performance gain by capturing the built environment and integrating domain knowledge yields an additional 4.1% improvement. To evaluate generalizability across unseen environments, cross-site validation was conducted using site-based partitioning. The zero-shot PedX-LLM configuration achieves 66.9% balanced accuracy on five unseen test sites, outperforming the baseline data-driven methods by at least 18 percentage points. Incorporating just five validation examples via few-shot learning to PedX-LLM further elevates the balanced accuracy to 72.2%. PedX-LLM demonstrates strong generalizability to unseen scenarios, confirming that vision-and-knowledge-enhanced reasoning enables the model to mimic human-like decision logic and overcome the limitations of purely data-driven methods.

cs.AI

Unidirectional reflection lasing based on destructive interference and Bragg scattering modulation in defective atomic lattice

The novel and ingenious scheme we propose for achieving unidirectional reflection lasing (URL) involves integrating a one-dimensional (1D) defective atomic lattice with a coherent gain atomic system. Its physical essence lies in the fact that the right-side reflectivity is drastically reduced due to the destructive interference between primary and secondary reflections, whereas on the left-side primary reflection is effectively suppressed and the secondary reflection is efficiently enhanced, ultimately reaching the lasing threshold. Through numerical results and further analyses, we have elucidated how to precisely tailor the lattice parameters and coupling fields to control destructive interference point (DIP), thereby realizing URL and enabling its active modulation. Our scheme is experimentally feasible and not only effectively circumvents the stringent conditions faced in directly realizing URL, providing a new pathway, but also beneficial for integrating active photonic devices into compact quantum networks and may improve the efficiency of optical information transmission.

physics.optics

VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance

Visual speech recognition (VSR) aims to transcribe spoken content from silent lip-motion videos and is particularly challenging in Mandarin due to severe viseme ambiguity and pervasive homophones. We propose VALLR-Pin, a two-stage Mandarin VSR framework that extends the VALLR architecture by explicitly incorporating Pinyin as an intermediate representation. In the first stage, a shared visual encoder feeds dual decoders that jointly predict Mandarin characters and their corresponding Pinyin sequences, encouraging more robust visual-linguistic representations. In the second stage, an LLM-based refinement module takes the predicted Pinyin sequence together with an N-best list of character hypotheses to resolve homophone-induced ambiguities. To further adapt the LLM to visual recognition errors, we fine-tune it on synthetic instruction data constructed from model-generated Pinyin-text pairs, enabling error-aware correction. Experiments on public Mandarin VSR benchmarks demonstrate that VALLR-Pin consistently improves transcription accuracy under multi-speaker conditions, highlighting the effectiveness of combining phonetic guidance with lightweight LLM refinement.

cs.CV