Searcharxiv⌕ Search

arXiv subjects

Feng Yang

Publications and source records attributed to Feng Yang.

At least 37 records · Page 2Linked to original sources

Mott-Derived Local Moments and Kondo Hybridization in a d-electron Kagome lattice

Unlike canonical Kondo lattices in f-electron systems, where localized f orbitalsnaturally provide local moments, d-electron Kondo lattices require a distinct mechanism for local-moment formation. However, the study of d-electron Kondo lattices in bulk materials remains far from settled, particularly with regard to the microscopic origin of the local moments. Here, we report a microscopic mechanism for this process in the bilayer kagome metal CsCr6Sb6, where strong correlations drive a Mott splitting of the kagome flat band to supply the requisite local moments. By combining STM/STS and ARPES, we resolve a spectroscopic hierarchy between high-energy correlation effects and low temperature hybridization. Low-temperature STS reveals a robust asymmetric suppression of the density of states near EF that is well captured phenomenologically by a Fano-type lineshape, while ARPES detects a sharp quasiparticlepeak near EF. These low-energy signatures evolveon the same temperature scale and disappear upon warming, consistent with the onset of Kondo hybridization. At the same time, STS resolves symmetric humps at approximately +-50 mV and ARPES identifies a weakly dispersive feature around 50 meV below EF; unlike the near-EF hybridization signatures, these features persist to substantially higher temperatures. This separation of energy and temperature scales supports a two-stage picture in which a kagome flat band first undergoes correlation-driven splitting into lower and upper Hubbard bands, and the occupied lower Hubbard band supplies the local moments that later hybridize with itinerant electrons at lower temperature. Our results therefore move beyond the phenomenology of a kagome Kondo lattice candidate and instead provide a microscopic spectroscopic picture linking Mottness to Kondo hybridization in a frustrated d-electron system.

cond-mat.str-el↗

HICT: High-precision 3D CBCT reconstruction from a single X-ray

Accurate 3D dental imaging is vital for diagnosis and treatment planning, yet CBCT's high radiation dose and cost limit its accessibility. Reconstructing 3D volumes from a single low-dose panoramic X-ray is a promising alternative but remains challenging due to geometric inconsistencies and limited accuracy. We propose HiCT, a two-stage framework that first generates geometrically consistent multi-view projections from a single panoramic image using a video diffusion model, and then reconstructs high-fidelity CBCT from the projections using a ray-based dynamic attention network and an X-ray sampling strategy. To support this, we built XCT, a large-scale dataset combining public CBCT data with 500 paired PX-CBCT cases. Extensive experiments show that HiCT achieves state-of-the-art performance, delivering accurate and geometrically consistent reconstructions for clinical use.

cs.CV↗

Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking

Multi-object tracking (MOT) is primarily dominated by two paradigms: tracking-by-detection (TBD) and tracking-by-query (TBQ). While TBD offers modular efficiency, its fragmented association pipeline often limits robustness in complex scenarios. Conversely, TBQ enhances semantic modeling end-to-end but suffers from high training costs and slow inference due to the tight coupling of detection and association. In this work, we propose the tracking-by-detection-and-query framework, TBDQ-Net, to advance the synergy between TBD and TBQ paradigms. By integrating a frozen detector with a lightweight associator, this architecture ensures intrinsic efficiency. Within this streamlined framework, we introduce tailored designs to address MOT-specific challenges. Concretely, we alleviate task conflicts and occlusions through the dual-stream update of the Basic Information Interaction (BII) module. The Content-Position Alignment (CPA) module further refines both content and positional components, providing well-aligned representations for association decoding. Extensive evaluations on DanceTrack, SportsMOT, and MOT20 benchmarks demonstrate that TBDQ-Net achieves a favorable efficiency-accuracy trade-off in challenging scenarios. Specifically, TBDQ-Net outperforms leading TBD methods by 6.0 IDF1 points on DanceTrack and achieves the best performance among TBQ methods in the crowded MOT20 benchmark. Relative to MOTRv2, TBDQ-Net reduces trainable parameters by approximately 80% while accelerating practical inference by 37.5%. These results highlight TBDQ-Net as an efficient alternative to heavy architectures, showcasing the efficacy of lightweight design. Source code is publicly available at https://github.com/FaithFlow/TBDQ-Net.

cs.CV↗

Expansion of Momentum Space and Full 2$π$ Solid Angle Photoelectron Collection in Laser-Based Angle-Resolved Photoemission Spectroscopy by Applying Sample Bias

Angle-resolved photoemission spectroscopy (ARPES) directly probes the energy and momentum of electrons in quantum materials, but conventional setups capture only a small fraction of the full 2$π$ solid angle. This limitation is acute in laser-based ARPES, where the low photon energy restricts momentum space despite ultrahigh resolution. Here we present systematic studies of bias ARPES, where applying a sample bias expands the accessible momentum range and enables full 2$π$ solid angle collection in two dimension using our 6.994 eV laser source. An analytical conversion relation is established and validated to accurately map the detector angle to the emission angle and the electron momentum in two dimensions. A precise approach is developed to determine the sample work function which is critical in the angle-momentum conversion of the bias ARPES experiments. Energy and angular resolutions are preserved under biases up to 100 V, and minimizing beam size is shown to be crucial. The technique is effective both near normal and off-normal geometries, allowing flexible Brillouin zone access with lower biases. Bias ARPES thus elevates laser ARPES to a new level, extending momentum coverage while retaining high resolution, and is applicable across a broad photon-energy range.

cond-mat.supr-con↗

Intent-Context Synergy Reinforcement Learning for Autonomous UAV Decision-Making in Air Combat

Autonomous UAV infiltration in dynamic contested environments remains a significant challenge due to the partially observable nature of threats and the conflicting objectives of mission efficiency versus survivability. Traditional Reinforcement Learning (RL) approaches often suffer from myopic decision-making and struggle to balance these trade-offs in real-time. To address these limitations, this paper proposes an Intent-Context Synergy Reinforcement Learning (ICS-RL) framework. The framework introduces two core innovations: (1) An LSTM-based Intent Prediction Module that forecasts the future trajectories of hostile units, transforming the decision paradigm from reactive avoidance to proactive planning via state augmentation; (2) A Context-Analysis Synergy Mechanism that decomposes the mission into hierarchical sub-tasks (safe cruise, stealth planning, and hostile breakthrough). We design a heterogeneous ensemble of Dueling DQN agents, each specialized in a specific tactical context. A dynamic switching controller based on Max-Advantage values seamlessly integrates these agents, allowing the UAV to adaptively select the optimal policy without hard-coded rules. Extensive simulations demonstrate that ICS-RL significantly outperforms baselines (Standard DDQN) and traditional methods (PSO, Game Theory). The proposed method achieves a mission success rate of 88\% and reduces the average exposure frequency to 0.24 per episode, validating its superiority in ensuring robust and stealthy penetration in high-dynamic scenarios.

cs.LG↗

Differentiable Rule Induction from Raw Sequence Inputs

Rule learning-based models are widely used in highly interpretable scenarios due to their transparent structures. Inductive logic programming (ILP), a form of machine learning, induces rules from facts while maintaining interpretability. Differentiable ILP models enhance this process by leveraging neural networks to improve robustness and scalability. However, most differentiable ILP methods rely on symbolic datasets, facing challenges when learning directly from raw data. Specifically, they struggle with explicit label leakage: The inability to map continuous inputs to symbolic variables without explicit supervision of input feature labels. In this work, we address this issue by integrating a self-supervised differentiable clustering model with a novel differentiable ILP model, enabling rule learning from raw data without explicit label leakage. The learned rules effectively describe raw data through its features. We demonstrate that our method intuitively and precisely learns generalized rules from time series and image data.

cs.AI↗

Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4

Formal methods (FM) are reliable but costly to apply, often requiring years of expert effort in industrial-scale projects such as seL4, especially for theorem proving. Recent advances in large language models (LLMs) have made automated theorem proving increasingly feasible. However, most prior work focuses on mathematics-oriented benchmarks such as miniF2F, with limited evaluation on real-world verification projects. The few studies that consider industrial-scale verification mostly rely on closed-source models with hundreds of billions of parameters, which cannot be locally deployed and incur substantial usage costs. In this paper, we propose AutoReal, an LLM-driven theorem proving method for real-world industrial-scale systems with support for lightweight local deployment. We evaluate AutoReal on the seL4-Isabelle verification project as a representative and challenging case study. AutoReal incorporates two key improvements: (1) chain-of-thought (CoT)-based proof training, which teaches the LLM the reasoning behind proof steps and enables step-wise explanations alongside proofs, and (2) context augmentation, which leverages proof context from the project to enhance LLM-driven proving. Based on the AutoReal methodology, we fine-tune a base model to obtain AutoReal-Prover, a compact 7B-scale prover for industrial-scale theorem proving. AutoReal-Prover achieves a 51.67% proof success rate on 660 theorems from seL4-designated Important Theories across all 10 seL4 proof categories, substantially outperforming prior attempts on seL4 (27.06%). To evaluate generalization, we further apply AutoReal-Prover to three security-related projects from the Archive of Formal Proofs (AFP), covering all 451 theorems and achieving a proof success rate of 53.88%. Overall, this work advances the application of LLM-driven theorem proving in real-world industrial-scale verification.

cs.CR↗

Electronic Structure and Superconducting Gap of HgBa$_2$Ca$_2$Cu$_3$O$_{8+δ}$ Revealed by Laser-Based Angle-Resolved Photoemission Spectroscopy

The spatially-resolved laser-based high resolution angle resolved photoemission spectroscopy (ARPES) measurements have been performed on the optimally-doped HgBa$_2$Ca$_2$Cu$_3$O$_{8+δ}$ (Hg1223) superconductor with a $T_c$ at 133 K. Two distinct regions are identified on the cleaved surface: the single Fermi surface region where only one Fermi surface is observed, and the double Fermi surface region where two Fermi surface sheets are resolved coming from both the inner (IP) and outer (OP) CuO$_2$ planes. The electronic structure and superconducting gap are measured on both of these two regions. In both cases, the observed electronic states are mainly concentrated near the nodal region. The momentum dependence of superconducting gap deviates from the standard d-wave form. These results indicate that the surface electronic structure of Hg1223 behaves more like that of underdoped cuprates.

cond-mat.supr-con↗

Electronic Origin of Density Wave Orders in a Trilayer Nickelate

The discovery of superconductivity in Ruddlesden-Popper nickelates has established a new frontier in the study of high-temperature superconductors. However, the underlying pairing mechanism and its relationship to the material's electronic and magnetic ground states remain elusive. Since unconventional superconductivity often emerges from a complex interplay of magnetic correlations, elucidating the magnetic ground state of the nickelates at ambient pressure is crucial for understanding the emergence of superconductivity under high pressure. Here, we combine high-resolution angle-resolved photoemission spectroscopy with tight-binding model simulation to investigate the electronic structure of the representative trilayer Ruddlesden-Popper nickelate La$_4$Ni$_3$O$_{10}$. We provide the first experimental evidence of band splitting induced by interlayer coupling and further resolve the momentum-dependent density wave gap structures along all the Fermi surfaces. Our findings identify the mirror-selective Fermi surface nesting as the origin of the interlayer antiferromagnetic spin density wave and demonstrate the dominant role of Ni-3d$_{z^2}$ orbitals in the low-energy physics of La$_4$Ni$_3$O$_{10}$. These results provide a fundamental framework for understanding the magnetic interactions and high-temperature superconductivity mechanism in the Ruddlesden-Popper nickelate family.

cond-mat.supr-con↗

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a result, the model must infer continuously shifting objectives from pixels alone, yielding delayed or overly conservative maneuvers. We argue that effective VLAs for autonomous driving need an online channel in which users can influence driving with specific intentions. To this end, we present EchoVLA, a user-aware VLA that couples camera streams with in situ audio instructions. We augment the nuScenes dataset with temporally aligned, intent-specific speech commands generated by converting ego-motion descriptions into synthetic audios. Further, we compose emotional speech-trajectory pairs into a multimodal Chain-of-Thought (CoT) for fine-tuning a Multimodal Large Model (MLM) based on Qwen2.5-Omni. Specifically, we synthesize the audio-augmented dataset with different emotion types paired with corresponding driving behaviors, leveraging the emotional cues embedded in tone, pitch, and speech tempo to reflect varying user states, such as urgent or hesitant intentions, thus enabling our EchoVLA to interpret not only the semantic content but also the emotional context of audio commands for more nuanced and emotionally adaptive driving behavior. In open-loop benchmarks, our approach reduces the average L2 error by $59.4\%$ and the collision rate by $74.4\%$ compared to the baseline of vision-only perception. More experiments on nuScenes dataset validate that EchoVLA not only steers the trajectory through audio instructions, but also modulates driving behavior in response to the emotions detected in the user's speech.

eess.AS↗

Joint Source-Channel-Generation Coding: From Distortion-oriented Reconstruction to Semantic-consistent Generation

Conventional communication systems, including both separation-based coding and AI-driven joint source-channel coding (JSCC), are largely guided by Shannon's rate-distortion theory. However, relying on generic distortion metrics fails to capture complex human visual perception, often resulting in blurred or unrealistic reconstructions. In this paper, we propose Joint Source-Channel-Generation Coding (JSCGC), a novel paradigm that shifts the focus from deterministic reconstruction to probabilistic generation. JSCGC leverages a generative model at the receiver as a generator rather than a conventional decoder to parameterize the data distribution, enabling direct maximization of mutual information under channel constraints while controlling stochastic sampling to produce outputs residing on the authentic data manifold with high fidelity. We further derive a theoretical lower bound on the maximum semantic inconsistency with given transmitted mutual information, elucidating the fundamental limits of communication in controlling the generative process. Extensive experiments on image transmission demonstrate that JSCGC substantially improves perceptual quality and semantic fidelity, significantly outperforming conventional distortion-oriented JSCC methods.

cs.IT↗

WaTeRFlow: Watermark Temporal Robustness via Flow Consistency

Image watermarking supports authenticity and provenance, yet many schemes are still easy to bypass with various distortions and powerful generative edits. Deep learning-based watermarking has improved robustness to diffusion-based image editing, but a gap remains when a watermarked image is converted to video by image-to-video (I2V), in which per-frame watermark detection weakens. I2V has quickly advanced from short, jittery clips to multi-second, temporally coherent scenes, and it now serves not only content creation but also world-modeling and simulation workflows, making cross-modal watermark recovery crucial. We present WaTeRFlow, a framework tailored for robustness under I2V. It consists of (i) FUSE (Flow-guided Unified Synthesis Engine), which exposes the encoder-decoder to realistic distortions via instruction-driven edits and a fast video diffusion proxy during training, (ii) optical-flow warping with a Temporal Consistency Loss (TCL) that stabilizes per-frame predictions, and (iii) a semantic preservation loss that maintains the conditioning signal. Experiments across representative I2V models show accurate watermark recovery from frames, with higher first-frame and per-frame bit accuracy and resilience when various distortions are applied before or after video generation.

cs.CV↗

Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting

Score Distillation Sampling (SDS) leverages pretrained 2D diffusion models to advance text-to-3D generation but neglects multi-view correlations, being prone to geometric inconsistencies and multi-face artifacts in the generated 3D content. In this work, we propose Coupled Score Distillation (CSD), a framework that couples multi-view joint distribution priors to ensure geometrically consistent 3D generation while enabling the stable and direct optimization of 3D Gaussian Splatting. Specifically, by reformulating the optimization as a multi-view joint optimization problem, we derive an effective optimization rule that effectively couples multi-view priors to guide optimization across different viewpoints while preserving the diversity of generated 3D assets. Additionally, we propose a framework that directly optimizes 3D Gaussian Splatting (3D-GS) with random initialization to generate geometrically consistent 3D content. We further employ a deformable tetrahedral grid, initialized from 3D-GS and refined through CSD, to produce high-quality, refined meshes. Quantitative and qualitative experimental results demonstrate the efficiency and competitive quality of our approach.

cs.CV↗

StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection

Multi-view 3D object detection is a fundamental task in autonomous driving perception, where achieving a balance between detection accuracy and computational efficiency remains crucial. Sparse query-based 3D detectors efficiently aggregate object-relevant features from multi-view images through a set of learnable queries, offering a concise and end-to-end detection paradigm. Building on this foundation, MV2D leverages 2D detection results to provide high-quality object priors for query initialization, enabling higher precision and recall. However, the inherent depth ambiguity in single-frame 2D detections still limits the accuracy of 3D query generation. To address this issue, we propose StereoMV2D, a unified framework that integrates temporal stereo modeling into the 2D detection-guided multi-view 3D detector. By exploiting cross-temporal disparities of the same object across adjacent frames, StereoMV2D enhances depth perception and refines the query priors, while performing all computations efficiently within 2D regions of interest (RoIs). Furthermore, a dynamic confidence gating mechanism adaptively evaluates the reliability of temporal stereo cues through learning statistical patterns derived from the inter-frame matching matrix together with appearance consistency, ensuring robust detection under object appearance and occlusion. Extensive experiments on the nuScenes and Argoverse 2 datasets demonstrate that StereoMV2D achieves superior detection performance without incurring significant computational overhead. Code will be available at https://github.com/Uddd821/StereoMV2D.

cs.CV↗

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially grounded driving actions by leveraging multimodal inputs, including 2D and 3D instance-aware visual representations, ego vehicle states, and language commands. To bridge the modality gap between driving visual representations and language embeddings, we introduce a hierarchical vision language alignment process, projecting both 2D and 3D structured visual tokens into a unified semantic space. Furthermore, we incorporate structured agent environment ego interaction modeling into the autoregressive decoding process, enabling the model to capture fine-grained spatial dependencies and behavior-aware dynamics critical for reliable trajectory planning. Extensive experiments on the nuScenes dataset demonstrate that OpenDriveVLA achieves state-of-the-art results across open-loop trajectory planning and driving-related question answering tasks. Qualitative analyses further illustrate its capability to follow high-level driving commands and generate trajectories under challenging scenarios, highlighting its potential for next-generation end-to-end autonomous driving.

cs.CV↗

Bio-inspired Integrated Networking and Control for Large-Scale Swarm: A Hierarchical Co-design

Unmanned aerial vehicle (UAV) swarms encounter the challenge of high overhead due to both network management and formation control requirements. In this paper, we propose a Bio-inspired Integrated Networking and Control (BINC) scheme, enabling efficient formation management for swarms comprising thousands of UAVs. The scheme forms a two-layer hierarchical structure, where network clusters and formations share the same groups so that cross-cluster control is eliminated. For networking, we design a fused routing message together with control information to reduce overhead, and limit clusters' size to local two-hop topologies for fast command transmission. For controlling, we develop a hybrid bio-inspired control approach, including a pigeon-like leader-follower algorithm within formations under the consideration of cluster topology maintenance, and a starling-like algorithm among formations that helps to improve the ability of obstacle avoidance. We establish a simulation platform for UAV swarms with over 1000 nodes, and experimental results show that the proposed BINC scheme can achieve highly maneuverable swarm formation marching with significant reduction on communication overhead.

cs.NI↗

Multi-Level Damage-Aware Graph Learning for Resilient UAV Swarm Networks

Unmanned aerial vehicle (UAV) swarm networks leverage resilient algorithms to restore connectivity from communication network split issues. However, existing graph learning-based approaches face over-aggregation and non-convergence problems caused by uneven and sparse topology under massive damage. In this paper, we propose a novel Multi-Level Damage-Aware (MLDA) Graph Learning algorithm to generate recovery solutions, explicitly utilizing information about destroyed nodes to guide the recovery process. The algorithm first employs a Multi-Branch Damage Attention (MBDA) module as a pre-processing step, focusing attention on the critical relationships between remaining nodes and destroyed nodes in the global topology. By expanding multi-hop neighbor receptive fields of nodes to those damaged areas, it effectively mitigating the initial sparsity and unevenness before graph learning commences. Second, a Dilated Graph Convolution Network (DGCN) is designed to perform convolution on the MBDA-processed bipartite graphs between remaining and destroyed nodes. The DGCN utilizes a specialized bipartite graph convolution operation to aggregate features and incorporates a residual-connected architecture to extend depth, directly generating the target locations for recovery. We theoretically proved the convergence of the proposed algorithm and the computational complexity is acceptable. Simulation results show that the proposed algorithm can guarantee the connectivity restoration with excellent scalability, while significantly expediting the recovery time and improving the topology uniformity after recovery.

cs.NI↗

Text-Guided Diffusion Model-based Generative Communication for Wireless Image Transmission

Reliable image transmission over wireless channels is particularly challenging at extremely low transmission rates, where conventional compression and channel coding schemes fail to preserve adequate visual quality. To address this issue, we propose a generative communication framework based on diffusion models, which integrates joint source channel coding (JSCC) with semantic-guided reconstruction leveraging a pre-trained generative model. Unlike conventional architectures that aim to recover exact pixel values of the original image, the proposed method focuses on preserving and reconstructing semantically meaningful visual content under severely constrained rates, ensuring perceptual plausibility and faithfulness to the scene intent. Specifically, the transmitter encodes the source image via JSCC and jointly transmits it with a textual prompt over the wireless channel. At the receiver, the corrupted low-rate representation is fused with the prompt and reconstructed through a Stable Diffusion model with ControlNet, enabling high-quality visual recovery. Leveraging both generative priors and semantic guidance, the proposed framework produces perceptually convincing images even under extreme bandwidth limitations. Experimental results demonstrate that the proposed method consistently outperforms conventional coding-based schemes and deep learning baselines, achieving superior perceptual quality and robustness across various channel conditions.

cs.IT↗