Searcharxiv⌕ Search

arXiv subjects

Cong Li

Publications and source records attributed to Cong Li.

At least 37 records · Page 2Linked to original sources

A neuromorphic vision system for open-world visual intelligence

Time-efficient and robust visual intelligence remains a critical challenge in unstructured open-world environments, yet current approaches often rely on computationally intensive neural architectures or task-specific sensors with limited versatility. Inspired by biological vision and information bottleneck theory, we report a neuromorphic vision system that performs task-oriented visual intelligence through an information distillation strategy (named as task traction mechanism) implemented on hardware. The system integrates a polarization-sensitive imager with a resistive random-access memory (RRAM) array to progressively distill task-relevant information via light field selection, region of interest extraction, and target anticipation. The neuromorphic vision system conducts visual tasks within an execution time of 193 μs. Evaluation across eight challenging open-world scenarios shows accuracy improvements of 25.54%, 37.73%, and 36.10% for object tracking, object segmentation, and trajectory prediction, respectively, together with an average 30.6-fold reduction in latency relative to state-of-the-art solutions.

eess.IV↗

Complete Next-to-Next-to-Leading-Order QCD Correction to $J/ψ\to 3γ$ Decay

We address the long-standing problem of negative decay and production rates in perturbative QCD for exclusive processes by proposing amplitude-level NRQCD factorization as a systematic prescription. Building on this, we present the first complete next-to-next-to-leading-order (NNLO) QCD correction to the decay $J/ψ\to 3γ$. The resulting partial width, $Γ(J/ψ\to 3γ) = 0.96^{+4.32}_{-0.13}$ eV, combines this NNLO contribution with the known up to $\mathcal{O}(α_s v^2)$ relativistic correction and shows markedly improved agreement with the high-precision BESIII measurement. In the same way, $Γ(Υ\to 3γ) = 0.0086^{+0.0028}_{-0.0006}$ eV is obtained. The dominant theoretical uncertainty originates from the renormalization scale variation, underscoring the challenge of perturbative convergence at this order and the necessity for future higher-order calculations.

hep-ph↗

Stress-triggered atomic explosion of trapped hydrogen initiates crack nucleation

Hydrogen embrittlement (HE) has persisted for more than a century as one of the most intractable problems in materials science. The prevailing view1 that diffusive H governs embrittlement has fostered the widespread assumption that H trapping at crystal defects mitigates HE. Here we overturn this conventional paradigm. Using plasma/ion irradiation of tungsten, we decouple -- for the first time -- H-induced crack nucleation from subsequent cavity propagation, and reveal nucleation as a two-stage mechanochemical fracture instability enabled by trapped H in the absence of diffusive H. In the first stage, H accumulation to a critical occupancy at dislocation cores acts as a chemical fuse, collapsing the local cohesive strength to a threshold at which infinitesimal external loads can trigger atomic decohesion. This bond rupture instantaneously enables the second stage: confined recombination of atomic hydrogen into molecular form. The abrupt release of chemical energy within an atomically restricted volume generates a transient inflation pressure that drives a dynamic, brittle jump to an internal macroscopic cavity. By separating mechanical decohesion triggering from energetic crack driving, our results provide a deterministic framework for the onset of H-induced crack nucleation under low-stress conditions. Furthermore, we place experimentally the classical H-enhanced decohesion model on an atomistic foundation and elevate it from phenomenology to prediction. Finally, by shifting the focus from experimentally elusive diffusive H to directly measurable trapped H, this work reframes HE as a deterministic, quantifiable instability, establishing a new paradigm for understanding and mitigating H-induced failure in high-strength metals.

cond-mat.mtrl-sci↗

Reinforcement learning in linear embedding space unlocks generalizable control across soft robot configurations

Soft-bodied organisms such as octopuses and elephant trunks exhibit remarkable morphological adaptability, dynamically reconfiguring body shape and stiffness, and flexibly adjusting their control strategies to enable versatile behaviors. Inspired by these biological systems, various soft robots have emerged in recent decades, featuring diverse materials, stiffnesses, and morphologies tailored to specific tasks. Despite substantial advances in the materials and structural designs of soft robots, developing a generalizable control framework capable of rapid adaptation across diverse configurations remains a long-standing challenge. Existing controllers are limited to fixed configurations, demanding laborious configuration-specific remodelling and policy redesign for new configurations. Here, we introduce a generalizable control system that enables rapid adaptation across diverse soft robot configurations via reinforcement learning in a shared linear Koopman embedding space. By encoding robot dynamics into this embedding space, our method decouples control policies from specific morphologies, allowing real-time, model-free policy adaptation across diverse configurations without retraining from scratch. We validate our system across 33 distinct robot configurations. Our system achieves a 75 times reduction in transfer samples across configurations, while sustaining robust performance under high-speed motion, heavy payloads, and multiactuator faults, and achieving real-world skills previously unattainable in soft robotics. This work establishes a unified and adaptable control paradigm for diverse soft robot configurations, bridging mechanical reconfigurability with control flexibility, and may offer broader insights for generalizable control in complex physical systems.

cs.RO↗

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-shelf embedding models, leading to suboptimal performance on massive text embedding benchmarks. In this paper, we identify a potential cause underlying this deficiency. Our motivation stems from an unexpected observation: text embeddings tend to align with frequent but uninformative tokens when projected onto the vocabulary space. We argue that this excessive expression of high-frequency tokens suppresses the model's ability to capture nuanced semantics. To address this, we introduce EmbedFilter, a simple linear transformation designed to refine text embeddings derived from LLMs directly. Specifically, we uncover that the unembedding matrix within LLMs encodes a latent space that is actively writing these frequent tokens into embedding space. By filtering out this subspace, EmbedFilter suppress the influence of high-frequency tokens, thereby enhancing semantic representations. As a compelling byproduct, this enables an inherent dimensionality reduction, lowering index storage and speedup retrieval while fully preserving the refined embedding quality. Our experiments across multiple LLM backbones demonstrate that LLMs equipped with EmbedFilter achieve superior zero-shot downstream performance even with significantly reduced embedding dimensions. We hope our findings provide deeper insights into the mechanisms of LLM-based representations and inspire more principled designs to improve text embeddings training. Our code is available at https://github.com/CentreChen/EmbFilter.

cs.CL↗

Accelerating NBTI Aging Evaluation via Physics-Aware Graph Attention Networks

As semiconductor technology advances to smaller nodes, Negative Bias Temperature Instability (NBTI) under prolonged workloads has emerged as a significant bottleneck constraining reliability-aware Design-Technology Co-Optimization (DTCO). Conventional TCAD simulations incur prohibitive computational overhead when evaluating device aging characteristics, making it difficult to satisfy the demand for efficient iterative design cycles. To address this challenge, this paper proposes an aging evaluation framework based on a physics-aware graph attention network (Physics-Aware RelGAT). By losslessly mapping unstructured device meshes into attributed graphs, this framework constructs a 45-dimensional device encoding scheme that integrates interface trap distributions and macroscopic electro-thermal stresses, achieving a direct mapping from underlying physical quantities to device degradation characteristics. To overcome the challenge of predicting currents that span multiple orders of magnitude, a dual-end normalization strategy and a log-scale loss function optimization are introduced, ensuring the model possesses high-precision fitting capabilities. Experimental results demonstrate that the model achieves a mean error of only 1.27% on an independent test set, achieving an acceleration of approximately 17,000 times compared to traditional TCAD simulations. This framework provides a solution for the assessment of circuit reliability in advanced process nodes that successfully balances physical fidelity with industrial-grade efficiency.

cs.ET↗

Nanoscale Polar Landscapes in Quantum Paraelectric SrTiO3

SrTiO3 is a textbook quantum paraelectric, with ferroelectricity purportedly suppressed by quantum fluctuations of ionic positions down to the lowest temperatures. The precise real space structure of SrTiO3 at low temperature, however, has remained undefined despite decades of study. Here we directly image the low-temperature polar structure of quantum parelectric SrTiO3, using cryogenic scanning transmission electron microscopy down to 20 K. High resolution imaging reveals a spatially fluctuating landscape of nanoscale domains of finite polarization. The short-range polar domains first grow and self-organize into a periodic structure over tens of nanometers. However, the process reverses when entering the quantum paraelectric regime below 40 K and the periodically ordered polar nanodomains fragment into small clusters.

cond-mat.mtrl-sci↗

Entanglement redistribution of hyperon-antihyperon pair via sequential decay

Hyperon-antihyperon pairs produced in high energy electron-positron annihilation constitute a naturally spin-entangled system in the high energy regime. Recently, a probabilistic amplification of entanglement, termed autodistillation, has been found in the daughter baryon-antibaryon pairs from hyperon decay and is constrained by an upper boundary. This work demonstrates that the quantum entanglement in this process may be accompanied by a decrease, constrained by a lower boundary, but will not be completely lost. Thus, the entanglement of these systems undergoes redistribution within the phase space during the sequential decays of hyperons, highlighting an important role of hyperon polarization. By using the explicit spin density matrix of baryon pairs, it is also found that quantumness of the system characterized by quantum discord always have the possibility to increase during decay processes, even when entanglement evaluated by concurrence and negativity does not increase.

hep-ph↗

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)

In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration, existing models still encounter substantial challenges in real-world low-light portrait scenarios. Specifically, they struggle to achieve an optimal balance among noise suppression, detail preservation, and faithful illumination and color reproduction. To bridge this gap, this challenge aims to establish a novel benchmark for real-world low-light portrait restoration. We comprehensively evaluate the proposed algorithms utilizing a hybrid evaluation system that integrates objective quantitative metrics with rigorous subjective assessment protocols. For this competition, we provide a dataset containing 800 groups of real-captured low-light portrait data. Each group consists of a 1K-resolution low-light input image, a 1K ground truth (GT), and a 1K person mask. This challenge has garnered widespread attention from both academia and industry, attracting over 100 participating teams and receiving more than 3,000 valid submissions. This report details the motivation behind the challenge, the dataset construction process, the evaluation metrics, and the various phases of the competition. The released dataset and baseline code for this track are publicly available from the same \href{https://github.com/zsn1434/AI_Flash-BaseLine/tree/main}{GitHub repository}, and the official challenge webpage is hosted on \href{https://www.codabench.org/competitions/12885/}{CodaBench}.

cs.CV↗

A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators

Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based 3D-DRAM has been adopted in LLM accelerators. While this emerging technology provides strong performance gains over existing hardware, current 3D-DRAM accelerators (3D-Accelerators) rely on closed-source evaluation tools, limiting access to publicly available performance analysis methods. Moreover, existing designs are highly customized for specific scenarios, lacking a general and reusable full-stack modeling for 3D-Accelerators across diverse usecases. To bridge this fundamental gap, we present ATLAS, the first silicon-proven Architectural Three-dimesional-DRAM-based LLM Accelerator Simulation framework. Built on commercially deployed multi-layer 3D-DRAM technology, ATLAS introduces unified abstractions for both 3D-Accelerator system architecture and programming primitives to support arbitrary LLM inference scenarios. Validation against real silicon shows that ATLAS achieves $\le$8.57% simulation error and 97.26-99.96\% correlation with measured performance. Through design space exploration with ATLAS, we demonstrate its ability to guide architecture design and distill key takeaways for both 3D-DRAM memory system and 3D-Accelerator microarchitecture across scenarios. ATLAS will be open-sourced upon publication, enabling further research on 3D-Accelerators.

cs.AR↗

Bundle EXTRA for Decentralized Optimization

Decentralized primal-dual methods are widely used for solving decentralized optimization problems, but their updates often rely on the potentially crude first-order Taylor approximations of the objective functions, which can limit convergence speed. To overcome this, we replace the first-order Taylor approximation in the primal update of EXTRA, which can be interpreted as a primal-dual method, with a more accurate multi-cut bundle model, resulting in a fully decentralized bundle EXTRA method. The bundle model incorporates historical information to improve the approximation accuracy, potentially leading to faster convergence. Under mild assumptions, we show that a KKT residual converges to zero. Numerical experiments on decentralized least-squares problems demonstrate that, compared to EXTRA, the bundle EXTRA method converges faster and is more robust to step-size choices.

math.OC↗

Flat band driven competing charge and spin instabilities in the altermagnet CrSb

The confinement of electronic wavefunctions in momentum space can give rise to flat electronic bands, where the quenching of kinetic energy enhances the density of states and amplifies interaction effects. Such conditions are fertile ground for emergent quantum phases, as spin, charge and lattice degrees of freedom become strongly entangled. In these regimes, subtle competitions between intertwined order parameters often dictate the macroscopic ground state, producing complex and sometimes unexpected collective behavior. Here we show that the altermagnet CrSb provides a realization of this scenario, and uncover short-range charge-order fluctuations at the M point of the Brillouin zone, q*=(1/2 0), persisting above the Neel temperature (TN). Remarkably, these fluctuations collapse upon entering the magnetically ordered phase, revealing a direct and robust competition between charge and spin order. At TN, the phonon dispersion at q* develops a pronounced Kohn-like anomaly, signaling strong electron-phonon coupling in the vicinity of the magnetic transition. Below TN, exchange striction dramatically renormalizes the associated soft phonon mode by approximately ~6 meV, the largest spin-phonon coupling ever reported. First-principles calculations attribute this behavior to a strong coupling between nearly dispersionless electronic states and a phonon branch that appears unstable at the harmonic level only when no magnetic order is considered, revealing the large sensitivity of the lattice to magnetic symmetry breaking. The competition between charge and spin order parameters, amplified by flat-band physics, drives the observed phonon anomaly and its abrupt reconstruction at TN. With its chemically simple structure and symmetry-protected altermagnetic state, CrSb emerges as a model platform to explore how flat electronic bands mediate giant spin-phonon coupling and competing broken symmetries.

cond-mat.str-el↗

Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator

Large language models (LLMs) have been widely deployed for online generative services, where numerous LLM instances jointly handle workloads with fluctuating request arrival rates and variable request lengths. To efficiently execute coexisting compute-intensive and memory-intensive operators, near-memory processing (NMP) based computing paradigm has been extensively proposed. However, existing NMP designs adopt coarse-grained KV cache management and inflexible attention execution flow. Such limitations hinder these proposals from efficiently handling \textit{highly dynamic} LLM serving workloads, limiting their ability to accelerate LLM serving. To tackle these problems, we propose Helios, a Hybrid-bonding-based \uline{L}LM \uline{S}erving accelerator. Helios aims to bridge the fundamental gap between the dynamic nature of KV cache management in LLM serving and the distributed, non-uniform memory abstraction among NMP processing engines (PEs). To this end, we design both the intra-PE execution flow and the inter-PE communication primitives for distributed tiled attention execution. We further propose \textit{spatially-aware} KV cache allocation mechanism to balance the attention workload distribution while minimizing the inter-PE data transfer overhead. Compared with existing GPU/NMP designs, Helios achieves 3.25 times (geomean) speedup and 3.36 times (geomean) better energy efficiency, along with up to 72%/76% P50/P99 time-between-tokens degradation.

cs.AR↗

Systematic Characterization of Transmon Qubit Stability with Thermal Cycling

The temporal stability and reproducibility of qubit parameters are critical for the long-term operation and maintenance of superconducting quantum processors. In this work, we present a comprehensive longitudinal characterization of 27 frequency-tunable transmon qubits spanning over one year across four thermal cycles. Our results establish a distinct hierarchy of stability for superconducting hardware. We find that the intrinsic device parameters determining the qubit frequency and the baseline energy relaxation times ($T_1$) exhibit high robustness against thermal stress, characterized by frequency deviations typically confined within 0.5\% and non-degraded coherence baselines. In stark contrast, the environmental variables, specifically the background magnetic flux offsets and the microscopic landscape of two-level system (TLS) defects, undergo a significant stochastic reconfiguration after each cycle. By employing frequency-dependent relaxation spectroscopy and a quantitative metric, the $T_1$ Spectral Topography Fidelity, we demonstrate that thermal cycling acts as a ``hard reset'' for the local defect environment. This process introduces a level of spectral randomization equivalent to thousands of hours of continuous low-temperature evolution. These findings confirm that while the fabrication quality is preserved, the specific noise realization is statistically distinct for each thermal cycle, necessitating automated recalibration strategies for large-scale quantum systems.

quant-ph↗

Low-Loss, High-Coherence Airbridge Interconnects Fabricated by Single-Step Lithography

Airbridges are essential for creating high-performance, low-parasitic interconnects in integrated circuits and quantum devices. Conventional multi-step fabrication methods hinder miniaturization and introduce process-related defects. We report a simplified process for fabricating nanoscale airbridges using only a single electron-beam lithography step. By optimizing a multilayer resist stack with a triple-exposure-dose scheme and a thermal reflow step, we achieve smooth, suspended metallic bridges with sub-200-nm features that exhibit robust mechanical stability. Fabricated within a gradiometric SQUID design for superconducting transmon qubits, these airbridges introduce no measurable additional loss in the relaxation time $T_1$, while enabling a 2.5-fold enhancement of the dephasing time $T_2^*$. This efficient method offers a practical route toward integrating high-performance three-dimensional interconnects in advanced quantum and nano-electronic devices.

quant-ph↗

Two loop QCD corrections to $e^+ e^- \to J/ψ+ η_c$ in asymptotic expansion

Within the framework of NRQCD, the short-distance coefficients (SDCs) for the process $e^+e^-\to J/ψ+η_c$ have been obtained up to NNLO in asymptotic expansions over $r={16m_c^2}/{s}$ up to $r^{15}$. Although these asymptotic expressions are deviated from the full results near the threshold $r= 1$, they provide excellent approximations to the full results for $r<0.8$, with deviations less than $3\%$. Therefore, these asymptotic expressions offer reliable applications for phenomenological predictions across a wide range of center-of-mass energies $\sqrt{s}$. Utilizing these asymptotic expressions, we present phenomenological predictions for the cross sections in both the on-shell mass scheme and the $\overline{\rm MS}$ mass scheme, with the uncertainty arising from the renormalization scale $μ_R$ included. The $μ_R$ uncertainty for predictions from the $\overline{\rm MS}$ mass scheme is slightly larger than that from the on-shell mass scheme, which is partly attributed to the helicity flip in the process $e^+e^-\to J/ψ+η_c$. We observe that both mass schemes yield quite similar predictions, and our theoretical results are consistent with the available experimental data.

hep-ph↗

Exploring Lorentz Invariance Violation from Ultra-high-energy Gamma Rays Observed by LHAASO

Recently the LHAASO Collaboration published the detection of 12 ultra-high-energy gamma-ray sources above 100 TeV, with the highest energy photon reaching 1.4 PeV. The first detection of PeV gamma rays from astrophysical sources may provide a very sensitive probe of the effect of the Lorentz invariance violation (LIV), which results in decay of high-energy gamma rays in the superluminal scenario and hence a sharp cutoff of the energy spectrum. Two highest energy sources are studied in this work. No signature of the existence of LIV is found in their energy spectra, and the lower limits on the LIV energy scale are derived. Our results show that the first-order LIV energy scale should be higher than about 10^5 times the Planck scale M_{pl} and that the second-order LIV scale is >10^{-3}M_{pl}. Both limits improve by at least one order of magnitude the previous results.

astro-ph.HE↗

Constraints on heavy decaying dark matter from 570 days of LHAASO observations

The Kilometer Square Array~(KM2A) of the Large High Altitude Air Shower Observatory (LHAASO) aims at surveying the northern gamma-ray sky at energies above 10 TeV with unprecedented sensitivity. Gamma-ray observations have long been one of the most powerful tools for dark matter searches, as e.g., high-energy gamma-rays could be produced by the decays of heavy dark matter particles. In this letter, we present the first dark matter analysis with LHAASO-KM2A, using the first 340~days of data from 1/2-KM2A and 230~days of data from 3/4-KM2A. Several regions of interest are used to search for a signal and account for the residual cosmic-ray background after gamma/hadron separation. We find no excess of dark matter signals, and thus place some of the strongest gamma-ray constraints on the lifetime of heavy dark matter particles with mass between 10^5 and 10^9~GeV. Our results with LHAASO are robust, and have important implications for dark matter interpretations of the diffuse astrophysical high-energy neutrino emission.

astro-ph.HE↗