SearcharxivSearch

arXiv subjects

Die Hu

Publications and source records attributed to Die Hu.

18 recordsLinked to original sources

ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation

Volumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts. To address these challenges, we propose \ourname, a Gaussian splatting based framework for volumetric video reconstruction. Our key insight is that explicitly tracking long term complex motion with individual Gaussian primitives is inherently unstable. Instead, we organize Gaussians around time conditioned anchors that localize their spatial and temporal support, thereby reducing long range motion complexity. We further introduce a temporal windowing strategy to activate only anchors relevant to the queried time, which improves scalability and temporal coherence. In addition, to ensure spatial and temporal stability, we design a compact set of multi level anchor features that encode global features, local spatial features, and local temporal features, jointly constraining Gaussian generation. Extensive experiments demonstrate that \ourname \ consistently outperforms prior methods on long sequence volumetric videos with complex motions. Project page: https://github.com/WuJH2001/ATGS.

cs.CV

ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models

Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant entrances. We propose ParkSense, a framework that repurposes idle compute during low-risk AV states -- queuing at red lights, traffic congestion, parking-lot crawl -- to run a Vision-Language Model (VLM) on pre-cached satellite and street view imagery, identifying entrances and legal parking zones. We formalize the Delivery-Aware Precision Parking (DAPP) problem, show that a quantized 7B VLM completes inference in 4-8 seconds on HW4-class hardware, and estimate annual per-driver income gains of 3,000-8,000 USD in the U.S. Five open research directions are identified at this unexplored intersection of autonomous driving, computer vision, and last-mile logistics.

cs.CV

Pentest-R1: Towards Autonomous Penetration Testing Reasoning Optimized via Two-Stage Reinforcement Learning

Automating penetration testing is crucial for enhancing cybersecurity, yet current Large Language Models (LLMs) face significant limitations in this domain, including poor error handling, inefficient reasoning, and an inability to perform complex end-to-end tasks autonomously. To address these challenges, we introduce Pentest-R1, a novel framework designed to optimize LLM reasoning capabilities for this task through a two-stage reinforcement learning pipeline. We first construct a dataset of over 500 real-world, multi-step walkthroughs, which Pentest-R1 leverages for offline reinforcement learning (RL) to instill foundational attack logic. Subsequently, the LLM is fine-tuned via online RL in an interactive Capture The Flag (CTF) environment, where it learns directly from environmental feedback to develop robust error self-correction and adaptive strategies. Our extensive experiments on the Cybench and AutoPenBench benchmarks demonstrate the framework's effectiveness. On AutoPenBench, Pentest-R1 achieves a 24.2\% success rate, surpassing most state-of-the-art models and ranking second only to Gemini 2.5 Flash. On Cybench, it attains a 15.0\% success rate in unguided tasks, establishing a new state-of-the-art for open-source LLMs and matching the performance of top proprietary models. Ablation studies confirm that the synergy of both training stages is critical to its success.

cs.AI

Regret Minimization in Population Network Games: Vanishing Heterogeneity and Convergence to Equilibria

Understanding and predicting the behavior of large-scale multi-agents in games remains a fundamental challenge in multi-agent systems. This paper examines the role of heterogeneity in equilibrium formation by analyzing how smooth regret-matching drives a large number of heterogeneous agents with diverse initial policies toward unified behavior. By modeling the system state as a probability distribution of regrets and analyzing its evolution through the continuity equation, we uncover a key phenomenon in diverse multi-agent settings: the variance of the regret distribution diminishes over time, leading to the disappearance of heterogeneity and the emergence of consensus among agents. This universal result enables us to prove convergence to quantal response equilibria in both competitive and cooperative multi-agent settings. Our work advances the theoretical understanding of multi-agent learning and offers a novel perspective on equilibrium selection in diverse game-theoretic scenarios.

cs.GT

VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework

Penetration testing is a vital practice for identifying and mitigating vulnerabilities in cybersecurity systems, but its manual execution is labor-intensive and time-consuming. Existing large language model (LLM)-assisted or automated penetration testing approaches often suffer from inefficiencies, such as a lack of contextual understanding and excessive, unstructured data generation. This paper presents VulnBot, an automated penetration testing framework that leverages LLMs to simulate the collaborative workflow of human penetration testing teams through a multi-agent system. To address the inefficiencies and reliance on manual intervention in traditional penetration testing methods, VulnBot decomposes complex tasks into three specialized phases: reconnaissance, scanning, and exploitation. These phases are guided by a penetration task graph (PTG) to ensure logical task execution. Key design features include role specialization, penetration path planning, inter-agent communication, and generative penetration behavior. Experimental results demonstrate that VulnBot outperforms baseline models such as GPT-4 and Llama3 in automated penetration testing tasks, particularly showcasing its potential in fully autonomous testing on real-world machines.

cs.SE

Spin correlations in the parent phase of Li$_{1-x}$Fe$_x$ODFeSe

Elucidating spin correlations in the parent compounds of high-temperature superconductors is crucial for understanding superconductivity. We used neutron scattering to study spin correlations in Li$_{1-x}$Fe$_x$ODFeSe, an insulating material with reduced electron carriers compared to its superconducting counterpart ($T_c$ = 41 K), serving as the undoped parent compound. Our findings show a reduced total fluctuating moment in this insulator relative to FeSe and 122 iron pnictides, likely due to increased interlayer distances from intercalation, which enhance fluctuations and reduce the intensity of spin excitations. Moreover, we observed a V-shaped spin wave-like excitation dispersion, contrasting with the twisted hourglass pattern in the superconducting counterpart. Electron doping shifts spin excitation from ($\pi$, 0) point to an incommensurate position towards ($\pi$, $\pi$) direction below 65 meV. This transition from V-shaped to hourglass-like dispersion, akin to behaviors in hole-doped cuprates, suggests a potential shared mechanism in magnetism and superconductivity across these diverse systems.

cond-mat.supr-con

VQ-DeepVSC: A Dual-Stage Vector Quantization Framework for Video Semantic Communication

In response to the rapid growth of global videomtraffic and the limitations of traditional wireless transmission systems, we propose a novel dual-stage vector quantization framework, VQ-DeepVSC, tailored to enhance video transmission over wireless channels. In the first stage, we design the adaptive keyframe extractor and interpolator, deployed respectively at the transmitter and receiver, which intelligently select key frames to minimize inter-frame redundancy and mitigate the cliff-effect under challenging channel conditions. In the second stage, we propose the semantic vector quantization encoder and decoder, placed respectively at the transmitter and receiver, which efficiently compress key frames using advanced indexing and spatial normalization modules to reduce redundancy. Additionally, we propose adjustable index selection and recovery modules, enhancing compression efficiency and enabling flexible compression ratio adjustment. Compared to the joint source-channel coding (JSCC) framework, the proposed framework exhibits superior compatibility with current digital communication systems. Experimental results demonstrate that VQ-DeepVSC achieves substantial improvements in both Multi-Scale Structural Similarity (MS-SSIM) and Learned Perceptual Image Patch Similarity (LPIPS) metrics than the H.265 standard, particularly under low channel signal-to-noise ratio (SNR) or multi-path channels, highlighting the significantly enhanced transmission capabilities of our approach.

cs.NI

Hierarchical Salient Patch Identification for Interpretable Fundus Disease Localization

With the widespread application of deep learning technology in medical image analysis, the effective explanation of model predictions and improvement of diagnostic accuracy have become urgent problems that need to be solved. Attribution methods have become key tools to help doctors better understand the diagnostic basis of models, and are used to explain and localize diseases in medical images. However, previous methods suffer from inaccurate and incomplete localization problems for fundus diseases with complex and diverse structures. To solve these problems, we propose a weakly supervised interpretable fundus disease localization method called hierarchical salient patch identification (HSPI) that can achieve interpretable disease localization using only image-level labels and a neural network classifier (NNC). First, we propose salient patch identification (SPI), which divides the image into several patches and optimizes consistency loss to identify which patch in the input image is most important for the network's prediction, in order to locate the disease. Second, we propose a hierarchical identification strategy to force SPI to analyze the importance of different areas to neural network classifier's prediction to comprehensively locate disease areas. Conditional peak focusing is then introduced to ensure that the mask vector can accurately locate the disease area. Finally, we propose patch selection based on multi-sized intersections to filter out incorrectly or additionally identified non-disease regions. We conduct disease localization experiments on fundus image datasets and achieve the best performance on multiple evaluation metrics compared to previous interpretable attribution methods. Additional ablation studies are conducted to verify the effectiveness of each method.

cs.CV

FeaInfNet: Diagnosis in Medical Image with Feature-Driven Inference and Visual Explanations

Interpretable deep learning models have received widespread attention in the field of image recognition. Due to the unique multi-instance learning of medical images and the difficulty in identifying decision-making regions, many interpretability models that have been proposed still have problems of insufficient accuracy and interpretability in medical image disease diagnosis. To solve these problems, we propose feature-driven inference network (FeaInfNet). Our first key innovation involves proposing a feature-based network reasoning structure, which is applied to FeaInfNet. The network of this structure compares the similarity of each sub-region image patch with the disease templates and normal templates that may appear in the region, and finally combines the comparison of each sub-region to make the final diagnosis. It simulates the diagnosis process of doctors to make the model interpretable in the reasoning process, while avoiding the misleading caused by the participation of normal areas in reasoning. Secondly, we propose local feature masks (LFM) to extract feature vectors in order to provide global information for these vectors, thus enhancing the expressive ability of the FeaInfNet. Finally, we propose adaptive dynamic masks (Adaptive-DM) to interpret feature vectors and prototypes into human-understandable image patches to provide accurate visual interpretation. We conducted qualitative and quantitative experiments on multiple publicly available medical datasets, including RSNA, iChallenge-PM, Covid-19, ChinaCXRSet, and MontgomerySet. The results of our experiments validate that our method achieves state-of-the-art performance in terms of classification accuracy and interpretability compared to baseline methods in medical image diagnosis. Additional ablation studies verify the effectiveness of each of our proposed components.

cs.CV

Hierarchical Codebook Design and Analytical Beamforming Solution for IRS Assisted Communication

In intelligent reflecting surface (IRS) assisted communication, beam search is usually time-consuming as the multiple-input multiple-output (MIMO) of IRS is usually very large. Hierarchical codebooks is a widely accepted method for reducing the complexity of searching time. The performance of this method strongly depends on the design scheme of beamforming of different beamwidths. In this paper, a non-constant phase difference (NCPD) beamforming algorithm is proposed. To implement the NCPD algorithm, we first model the phase shift of IRS as a continuous function, and then determine the parameters of the continuous function through the analysis of its array factor. Then, we propose a hierarchical codebook and two beam training schemes, namely the joint searching (JS) scheme and direction-wise searching (DWS) scheme by using the NCPD algorithm which can flexibly change the width, direction and shape of the beam formed by the IRS array. Simulation results show that the NCPD algorithm is more accurate with smaller side lobes, and also more stable on IRS of different sizes compared to other wide beam algorithms. The misalignment rate of the beam formed by the NCPD method is significantly reduced. The time complexity of the NCPD algorithm is constant, thus making it more suitable for solving the beamforming design problem with practically large IRS.

cs.IT

A Cyber-Physical Routing Protocol Exploiting Trajectory Dynamics for Mission-Oriented Flying Ad Hoc Networks

As a special type of mobile ad hoc network (MANET), the flying ad hoc network (FANET) has the potential to enable a variety of emerging applications in both civilian wireless communications (e.g., 5G and 6G) and the defense industry. The routing protocol plays a pivotal role in FANET. However, when designing the routing protocol for FANET, it is conventionally assumed that the aerial nodes move randomly. This is clearly inappropriate for a mission-oriented FANET (MO-FANET), in which the aerial nodes typically move toward a given destination from given departure point(s), possibly along a roughly deterministic flight path while maintaining a well-established formation, in order to carry out certain missions. In this paper, a novel cyber-physical routing protocol exploiting the particular mobility pattern of an MO-FANET is proposed based on cross-disciplinary integration, which makes full use of the mission-determined trajectory dynamics to construct the time sequence of rejoining and separating, as well as the adjacency matrix for each node, as prior information. Compared with the existing representative routing protocols used in FANETs, our protocol achieves a higher packet-delivery ratio (PDR) at the cost of even lower overhead and lower average end-to-end latency, while maintaining a reasonably moderate and stable network jitter, as demonstrated by extensive ns-3-based simulations assuming realistic configurations in an MO-FANET.

cs.NI

Linear Network Coding Based Fast Data Synchronization for Wireless Ad Hoc Networks with Controlled Topology

Fast data synchronization in wireless ad hoc networks is a challenging and critical problem. It is fundamental for efficient information fusion, control and decision in distributed systems. Previously, distributed data synchronization was mainly studied in the latency-tolerant distributed databases, or assuming the general model of wireless ad hoc networks. In this paper, we propose a pair of linear network coding (NC) and all-to-all broadcast based fast data synchronization algorithms for wireless ad hoc networks whose topology is under operator's control. We consider both data block selection and transmitting node selection for exploiting the benefits of NC. Instead of using the store-and-forward protocol as in the conventional uncoded approach, a compute-and-forward protocol is used in our scheme, which improves the transmission efficiency. The performance of the proposed algorithms is studied under different values of network size, network connection degree, and per-hop packet error rate. Simulation results demonstrate that our algorithms significantly reduce the times slots used for data synchronization compared with the baseline that does not use NC.

cs.DC

Neutron Scattering Studies of the Breathing Pyrochlore Antiferromagnet LiGaCr$_{4}$O$_{8}$

We report neutron scattering measurements of the spinel oxide LiGaCr$_{4}$O$_{8}$, in which magnetic ions Cr$^{3+}$ form a breathing pyrochlore lattice. Our experiments reveal the coexistence of a nearly dispersionless resonance mode and dispersive spin wave excitations in the magnetically ordered state, which can be quantitatively described by a quantum spin model of hexagonal loops and linear spin wave theory with the same set of exchange parameters, respectively. Comparison to other Cr spinel oxides reveals a linear relationship between the resonance energy and lattice constant across all these materials, which is in agreement with our hexagonal loop calculations. Our results suggest a unified picture for spin resonances in Cr spinel oxides.

cond-mat.str-el

Learning Smooth Representation for Unsupervised Domain Adaptation

Typical adversarial-training-based unsupervised domain adaptation methods are vulnerable when the source and target datasets are highly-complex or exhibit a large discrepancy between their data distributions. Recently, several Lipschitz-constraint-based methods have been explored. The satisfaction of Lipschitz continuity guarantees a remarkable performance on a target domain. However, they lack a mathematical analysis of why a Lipschitz constraint is beneficial to unsupervised domain adaptation and usually perform poorly on large-scale datasets. In this paper, we take the principle of utilizing a Lipschitz constraint further by discussing how it affects the error bound of unsupervised domain adaptation. A connection between them is built and an illustration of how Lipschitzness reduces the error bound is presented. A \textbf{local smooth discrepancy} is defined to measure Lipschitzness of a target distribution in a pointwise way. When constructing a deep end-to-end model, to ensure the effectiveness and stability of unsupervised domain adaptation, three critical factors are considered in our proposed optimization strategy, i.e., the sample amount of a target domain, dimension and batchsize of samples. Experimental results demonstrate that our model performs well on several standard benchmarks. Our ablation study shows that the sample amount of a target domain, the dimension and batchsize of samples indeed greatly impact Lipschitz-constraint-based methods' ability to handle large-scale datasets. Code is available at https://github.com/CuthbertCai/SRDA.

cs.CV

Polarized neutron scattering studies of magnetic excitations in iron-selenide superconductor (Li$_{0.8}$Fe$_{0.2}$)ODFeSe ($T_c$ = 41 K)

We report polarized neutron scattering measurements of the low energy spin fluctuations of the iron-selenide superconductor Li$_{0.8}$Fe$_{0.2}$ODFeSe below and above its superconducting transition temperature $T_c=41$ K. Our experiments confirmed that the resonance mode near 21 meV is magnetic. Moreover, the spin excitations are essentially isotropic in spin space at 5$\leq E\leq$ 29 meV in the superconducting and normal states. Our results suggest that the resonance mode in iron-based superconductors becomes isotropic when the influence of spin-orbit coupling and magnetic/nematic order is minimized, similar to those observed in cuprate superconductors.

cond-mat.supr-con

Hybrid Interference Mitigation Using Analog Prewhitening

This paper proposes a novel scheme for mitigating strong interferences, which is applicable to various wireless scenarios, including full-duplex wireless communications and uncoordinated heterogenous networks. As strong interferences can saturate the receiver's analog-to-digital converters (ADC), they need to be mitigated both before and after the ADCs, i.e., via hybrid processing. The key idea of the proposed scheme, namely the Hybrid Interference Mitigation using Analog Prewhitening (HIMAP), is to insert an M-input M-output analog phase shifter network (PSN) between the receive antennas and the ADCs to spatially prewhiten the interferences, which requires no signal information but only an estimate of the covariance matrix. After interference mitigation by the PSN prewhitener, the preamble can be synchronized, the signal channel response can be estimated, and thus a minimum mean squared error (MMSE) beamformer can be applied in the digital domain to further mitigate the residual interferences. The simulation results verify that the HIMAP scheme can suppress interferences 80dB stronger than the signal by using off-the-shelf phase shifters (PS) of 6-bit resolution.

eess.SP

Structure of spin excitations in heavily electron-doped Li0.8Fe0.2ODFeSe superconductors

Heavily electron-doped iron-selenide (HEDIS) high-transition-temperature (high-$T_{\rm{c}}$) superconductors, which have no hole Fermi pockets, but have a notably high $T_{\rm{c}}$, have challenged the prevailing $s$$_\pm$ pairing scenario originally proposed for iron pnictides containing both electron and hole pockets. The microscopic mechanism underlying the enhanced superconductivity in HEDIS remains unclear. Here, we used neutron scattering to study the spin excitations of the HEDIS material Li$_{0.8}$Fe$_{0.2}$ODFeSe ($T_{\rm{c}}$ = 41 K). Our data revealed nearly ring-shaped magnetic resonant excitations surrounding ($π$, $π$) at $\sim$ 21 meV. As the energy increased, the spin excitations assumed a diamond shape, and they dispersed outward until the energy reached $\sim$ 60 meV and then inward at higher energies. The observed energy-dependent momentum structure and twisted dispersion of spin excitations near ($π$, $π$) are analogous to those of hole-doped cuprates in several aspects, thus implying that such spin excitations are essential for the remarkably high $T_{\rm{c}}$ in these materials.

cond-mat.supr-con

A Scalable Framework for CSI Feedback in FDD Massive MIMO via DL Path Aligning

Unlike the time-division duplexing (TDD) systems, the downlink (DL) and uplink (UL) channels are not reciprocal anymore in the case of frequency-division duplexing (FDD). However, some long-term parameters, e.g. the time delays and angles of arrival (AoAs) of the channel paths, still enjoy reciprocity. In this paper, by efficiently exploiting the aforementioned limited reciprocity, we address the DL channel state information (CSI) feedback in a practical wideband massive multiple-input multiple-output (MIMO) system operating in the FDD mode. With orthogonal frequency-division multiplexing (OFDM) waveform and assuming frequency-selective fading channels, we propose a scalable framework for the DL pilots design, DL CSI acquisition, and the corresponding CSI feedback in the UL. In particular, the base station (BS) can transmit the FFT-based pilots with the carefully-selected phase shifts. Then the user can rely on the so-called time-domain aggregate channel (TAC) to derive the feedback of reduced imensionality according to either its own knowledge about the statistics of the DL channels or the instruction from the serving BS. We demonstrate that each user can just feed back one scalar number per DL channel path for the BS to recover the DL CSIs. Comprehensive numerical results further corroborate our designs.

cs.IT