Searcharxiv⌕ Search

arXiv subjects

Jun Yang

Publications and source records attributed to Jun Yang.

At least 73 records · Page 4Linked to original sources

AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator

SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip data movement overhead for large DNN models. Existing CIM accelerator studies typically assume that DNN models fit entirely on-chip, leaving efficient dataflow design largely untapped. This paper introduces AccelCIM, a systematic dataflow exploration framework for SRAM CIM accelerator, which addresses two key limitations of prior work. (1) It formulates a systematic dataflow design space spanning CIM macro configurations and macro-array organizations. (2) It introduces rigorous design evaluation using cycle-accurate architectural simulation and post-layout PPA analysis. We conduct an extensive design space exploration and apply AccelCIM to representative LLM applications, providing practical insights for the principled design of CIM accelerators.

cs.AR↗

ChatSVA: Bridging SVA Generation for Hardware Verification via Task-Specific LLMs

Functional verification consumes over 50% of the IC development lifecycle, where SystemVerilog Assertions (SVAs) are indispensable for formal property verification and enhanced simulation-based debugging. However, manual SVA authoring is labor-intensive and error-prone. While Large Language Models (LLMs) show promise, their direct deployment is hindered by low functional accuracy and a severe scarcity of domain-specific data. To address these challenges, we introduce ChatSVA, an end-to-end SVA generation system built upon a multi-agent framework. At its core, the AgentBridge platform enables this multi-agent approach by systematically generating high-purity datasets, overcoming the data scarcity inherent to few-shot scenarios. Evaluated on 24 RTL designs, ChatSVA achieves 98.66% syntax and 96.12% functional pass rates, generating 139.5 SVAs per design with 82.50% function coverage. This represents a 33.3 percentage point improvement in functional correctness and an over 11x enhancement in function coverage compared to the previous state-of-the-art (SOTA). ChatSVA not only sets a new SOTA in automated SVA generation but also establishes a robust framework for solving long-chain reasoning problems in few-shot, domain-specific scenarios. An online service has been publicly released at https://www.nctieda.com/CHATDV.html.

cs.AR↗

Three-Dimensional Ocean Dynamics and Detectability of Tidally Locked Lava Worlds

Tidally locked lava planets are hot, rocky worlds on close-in orbits with a permanent molten dayside. With JWST, their surfaces and atmospheres are beginning to be revealed. This work investigates 3D magma-ocean dynamics, derives scaling laws for the resulting ocean heat transport (OHT), and predicts its detectability. For the first time, the ocean circulation driven by the intense momentum and mass exchanges with the supersonic atmosphere is considered in addition to that by thermal forcing. The wind forcing turns out to overwhelmingly dominate the other two mechanisms, driving ocean currents reaching $\sim$100 m s$^{-1}$ and greatly expanding the latitudinal extent of the Matsuno-Gill response. Despite these extreme flow speeds, scaling analysis and 3D simulations consistently demonstrate that magma-ocean circulation alone does not produce an observable hotspot offset. This inefficiency arises because basin geometry and circulation structure fundamentally constrain zonal heat redistribution, suppressing large-scale longitudinal transport even under vigorous flow.

astro-ph.EP↗

ExVerus: Verus Proof Repair via Counterexample Reasoning

Large Language Models (LLMs) have shown promising results in automating formal verification. However, existing approaches treat proof generation as a static, end-to-end prediction over source code, relying on limited verifier feedback and lacking access to concrete program behaviors. We present EXVERUS, a counterexample-guided framework that enables LLMs to reason about proofs using behavioral feedback via counterexamples. When a proof fails, EXVERUS automatically generates and validates counterexamples, and then guides the LLM to generalize them into inductive invariants to block these failures. Our evaluation shows that EXVERUS significantly improves proof accuracy, robustness, and token efficiency over the state-of-the-art prompting-based Verus proof generator.

cs.PL↗

Strain-released epitaxy of GaN enabled by compliant single-crystalline metal foils

Heteroepitaxy conventionally relies on rigid crystalline substrates, implicitly assuming that lattice and thermal mismatch must be accommodated within the epitaxial layer, leading to residual strain and defects that worsen with increasing substrate size. Here we demonstrate a substrate-mediated strain-partitioning regime in which lattice and thermal mismatch are preferentially partitioned into the substrate rather than stored in the epitaxial layer. We report the epitaxial growth of single-crystalline GaN on mechanically compliant yet crystallographically ordered single-crystalline copper foils. Atomic-resolution microscopy, geometric phase analysis and density functional theory reveal that mismatch-induced stress is primarily screened by elastic deformation of the Cu lattice, accompanied by localized interfacial slip confined to a few atomic layers, leaving the AlN and GaN epilayers nearly strain-free despite large nominal mismatch. Leveraging this strain-released epitaxial platform, we further demonstrate dense GaN micro-light-emitting diode arrays that benefit from efficient vertical electrical conduction and thermal dissipation enabled by the metallic substrate. By establishing compliant single-crystal metal foils as a new substrate class, this work identifies mechanical contrast as an underexplored governing parameter in heteroepitaxial design, with implications extending beyond GaN.

cond-mat.mtrl-sci↗

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self-consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self-evolving multimodal models. The code are available at https://github.com/OPPO-Mente-Lab/LLM-Self-Judge.

cs.CV↗

UAV-DETR: DETR for Anti-Drone Target Detection

Drone detection is pivotal in numerous security and counter-UAV applications. However, existing deep learning-based methods typically struggle to balance robust feature representation with computational efficiency. This challenge is particularly acute when detecting miniature drones against complex backgrounds under severe environmental interference. To address these issues, we introduce UAV-DETR, a novel framework that integrates a small-target-friendly architecture with real-time detection capabilities. Specifically, UAV-DETR features a WTConv-enhanced backbone and a Sliding Window Self-Attention (SWSA-IFI) encoder, capturing the high-frequency structural details of tiny targets while drastically reducing parameter overhead. Furthermore, we propose an Efficient Cross-Scale Feature Recalibration and Fusion Network (ECFRFN) to suppress background noise and aggregate multi-scale semantics. To further enhance accuracy, UAV-DETR incorporates a hybrid Inner-CIoU and NWD loss strategy, mitigating the extreme sensitivity of standard IoU metrics to minor positional deviations in small objects. Extensive experiments demonstrate that UAV-DETR significantly outperforms the baseline RT-DETR on our custom UAV dataset (+6.61% in mAP50:95, with a 39.8% reduction in parameters) and the public DUT-ANTI-UAV benchmark (+1.4% in Precision, +1.0% in F1-Score). These results establish UAV-DETR as a superior trade-off between efficiency and precision in counter-UAV object detection. The code is available at https://github.com/wd-sir/UAVDETR.

cs.CV↗

Quantifying Non-linearity in Topology Optimization with similarity based Visualization

Topology optimization (TO) can be viewed as seeking an optimal solution in the design space of a given TO problem. For weakly non-linear TO problems, e.g., compliance minimization, sensitivity-based methods typically converge well, whereas for strongly non-linear problems, e.g., maximum stress minimization, stabilization strategies such as stabilization terms and projection functions are often required to enhance convergence. Especially in scenarios with massive design variables, it is difficult to intuitively demonstrate the non-linear complexity of different TO problems and to elucidate the mechanisms by which stabilization strategies affect convergence. To address this challenge, we propose a visualization framework and a quantitative non-linearity index for objectives with varying complexity. We employ a multi-start fixed-gradient sampling tailored to similarity-based dimensionality reduction while keeping the computational cost under control. The samples are then parameterized via cosine similarity to obtain a low-dimensional visualization surface of the objective function. Based on this visualization, we construct a dimensionless complexity index with a clear geometric interpretation by measuring the gap between the visualization surface and the discrete approximation of its convex envelope, which enables quantitative comparisons of non-linearity across TO tasks, parameter choices, and stabilization strategies. Extensive comparative experiments show that the proposed approach is both adaptable and discriminative on a variety of representative TO problems, and it provides intuitive and measurable guidance for parameter selection.

math.OC↗

Evidence of Long-Lived Powerful Gyrosynchrotron Radio Emission in the Close Binary FF UMa

RS Canum Venaticorum (RS CVn) close binaries, characterized by tidal locking, rapid rotations, and strong magnetic fields, are ideal laboratories for high-resolution radio observations to probe emission processes, magnetic field configurations, and interaction activity. Despite their importance, only a few RS CVn sources have been explored by polarimetric observations of very long baseline interferometry (VLBI). To expand the effort, we have analyzed the existing Very Long Baseline Array (VLBA) astrometric data for the RS CVn binary FF Ursae Majoris (FF UMa). In the 5GHz VLBA experiments conducted between 2021 and 2024, both total intensity and circularly polarized emission were clearly detected at six of seven epochs. The consistently high brightness temperatures (10^7 K) and the moderate fractional circular polarization (10%-30%) over about three years indicate that the radio emission is mainly produced by gyrosynchrotron radiation from mildly relativistic electrons in the highly-ordered magnetic field. The radio luminosities are also comparable to those of previously studied powerful RS CVn binaries and show a significant anti-correlation with fractional circular polarization. A mean centroid offset of 13.4 +/- 3.1 solar radii between the Stokes I and V emission was found across multiple epochs, indicating a possible additional contribution from the secondary star via a magnetically active corona, a giant magnetic loop, or significant interaction activity with the primary star in the quiescent state.

astro-ph.SR↗

The influence of hypothetical exomoons on planetary thermal phase curves

More than 200 moons exist in our Solar System, yet no exomoon has been confirmed to date. While the innermost two planets of the Solar System lack natural satellites and most studies favour the existence of exomoons around long-period planets, some theoretical studies that take tidal dissipation, orbital decay, and migration processes into account suggest that exomoons may survive around short-period exoplanets. We investigated the impact of exomoons on planetary thermal phase curves and assessed their detectability within a theoretical framework. We simulated the thermal phase curves of exomoon-exoplanet systems, including mutual transits and occultations, and explored their dependence on planetary orbital periods across a wide range of systems. Close-in airless exomoons maintain large day-night temperature contrasts, amplifying the thermal phase-curve signal of the system. When the exomoon transits or is occulted by the exoplanet, the transit depth varies with the planetary phase, and the occultation depth varies with the exomoon's phase. The maximum occultation depth can reach $\sim$ 20 ppm for long-period systems. For short-period planets, the signal can reach up to $\sim$100 ppm, although such configurations may not be dynamically stable over long timescales. If exomoons are not accounted for, the planetary temperature distribution retrieved from observed thermal phase curves may overestimate the planetary day-night temperature contrast and underestimate the planetary horizontal heat transport. In principle, the periodic exomoon-exoplanet mutual occultation signal could be extracted using methods such as box-fitting least squares, providing a framework for future observational studies and instrument planning.

astro-ph.EP↗

Interference-Aware K-Step Reachable Communication in Multi-Agent Reinforcement Learning

Effective communication is pivotal for addressing complex collaborative tasks in multi-agent reinforcement learning (MARL). Yet, limited communication bandwidth and dynamic, intricate environmental topologies present significant challenges in identifying high-value communication partners. Agents must consequently select collaborators under uncertainty, lacking a priori knowledge of which partners can deliver task-critical information. To this end, we propose Interference-Aware K-Step Reachable Communication (IA-KRC), a novel framework that enhances cooperation via two core components: (1) a K-Step reachability protocol that confines message passing to physically accessible neighbors, and (2) an interference-prediction module that optimizes partner choice by minimizing interference while maximizing utility. Compared to existing methods, IA-KRC enables substantially more persistent and efficient cooperation despite environmental interference. Comprehensive evaluations confirm that IA-KRC achieves superior performance compared to state-of-the-art baselines, while demonstrating enhanced robustness and scalability in complex topological and highly dynamic multi-agent scenarios.

cs.AI↗

Shocks in the Symbiotic Recurrent Nova V3890 Sgr: VLBI Radio Imaging and Fermi GeV Gamma-Rays

We present very long baseline interferometric (VLBI) radio imaging and Fermi/LAT GeV $γ$-ray observations of the 2019 eruption of the symbiotic recurrent nova V3890 Sgr.The VLBI imaging spans 8 -- 51 days after eruption, synchronous with the detected $γ$-rays. VLBI imaging shows the eruption starts out asymmetric on day 8 with an eastern component brighter than a western component. By day 32 the blast is rather circularly symmetric, and on day 49, the nova shell is brighter along the north--south axis. This morphological evolution is explained by interaction with circumstellar material (CSM) comprised of a spherical wind plus an over-density in the orbital plane. Comparing radio images to optical line widths gives an expansion parallax distance of 6.8 kpc. In the first 32 days or eruption, VLBI images capture $>$80 per cent of the integrated flux (as measured by the VLA), implying that synchrotron emission dominates. A second peak in the VLA light curve is explained by an image on day 48 that reveals the nova shell surrounded by a diffuse halo, powered by synchrotron emission from particles that have diffused upstream of the shock. The $γ$-rays appear around optical maximum and remain detectable for 23 days; marginally significant $γ$-rays reappear around day 60, concurrent with the second radio peak. Modelling indicates radio and $γ$-ray emission arise in distinct shock regions: $γ$-rays from dense CSM in the orbital plane, radio from the more spherical CSM component. X-ray observations constrain the spherical CSM density, which is higher than in other symbiotic recurrent novae. Assuming equipartition, we estimate the fraction of the post-shock pressure in magnetic fields, $ε_B = 3 \times 10^{-4} - 2 \times 10^{-3}$.

astro-ph.HE↗

ToMPC: Task-oriented Model Predictive Control via ADMM for Safe Robotic Manipulation

This paper proposes a task-oriented model predictive control (ToMPC) framework for safe and efficient robotic manipulation in open workspaces. The framework unifies collision-free motion and robot-environment interaction to address diverse scenarios. Additionally, it introduces task-oriented obstacle avoidance that leverages kinematic redundancy to enhance manipulation efficiency in obstructed environments. This complex optimization problem is solved by the alternating direction method of multipliers (ADMM), which decomposes the problem into two subproblems tackled by differential dynamic programming (DDP) and quadratic programming (QP), respectively. The effectiveness of this approach is validated in simulation and hardware experiments on a Franka Panda robotic manipulator. Results demonstrate that the framework can plan motion and/or force trajectories in real time, maximize the manipulation range while avoiding obstacles, and strictly adhere to safety-related hard constraints.

cs.RO↗

Towards AI Search Paradigm

In this paper, we introduce the AI Search Paradigm, a comprehensive blueprint for next-generation search systems capable of emulating human information processing and decision-making. The paradigm employs a modular architecture of four LLM-powered agents (Master, Planner, Executor and Writer) that dynamically adapt to the full spectrum of information needs, from simple factual queries to complex multi-stage reasoning tasks. These agents collaborate dynamically through coordinated workflows to evaluate query complexity, decompose problems into executable plans, and orchestrate tool usage, task execution, and content synthesis. We systematically present key methodologies for realizing this paradigm, including task planning and tool integration, execution strategies, aligned and robust retrieval-augmented generation, and efficient LLM inference, spanning both algorithmic techniques and infrastructure-level optimizations. By providing an in-depth guide to these foundational components, this work aims to inform the development of trustworthy, adaptive, and scalable AI search systems.

cs.CL↗

Orbital-Selective Spin-Orbit Mott Insulator in Fractional Valence Iridate La$_3$Ir$_3$O$_{11}$

The combination of strong spin-orbit coupling and Coulomb interactions makes the $5d$ iridates a unique platform for realizing novel correlated electronic states. Here, utilizing infrared spectroscopy, we demonstrate that a robust Mott insulating state persists in the $1/3$-hole self-doped system La$_3$Ir$_3$O$_{11}$, evidenced by the collapse of the Drude response and the emergence of sharp excitations across the Mott gap. Our theoretical calculations reveal that the insulating behavior arises from the cooperative interplay of structural distortions, spin-orbit coupling, and Coulomb interactions. Specifically, octahedral distortion and Ir-Ir dimerization split the $t_{2g}$ orbitals, driving the $J_{\mathrm{eff}} = 1/2$ bands toward half-filling while keeping the $J_{\mathrm{eff}} = 3/2$ bands away from it. Consequently, electron correlations induce an orbital-selective Mott transition in the $J_{\mathrm{eff}} = 1/2$ bands, whereas a band-insulating gap develops in the $J_{\mathrm{eff}} = 3/2$ bands, thereby stabilizing the unconventional insulating state in La$_3$Ir$_3$O$_{11}$. These findings provide new insights into the design and understanding of the insulating ground state of spin-orbit-coupled iridates.

cond-mat.str-el↗

Gradient estimates for $p$-Laplacian equation with cubic polynomial nonlinearity on Riemannian manifolds

This paper studies a class of $p$-Laplace equations with cubic polynomial nonlinearity \[ Δ_p v + (v-a_1)(v-a_2)(v-a_3) = 0 \] on complete Riemannian manifolds $M$ with lower Ricci curvature bounds, where $a_1 < a_2 < a_3$ are real constants and $Δ_p v = \operatorname{div}(|\nabla v|^{p-2}\nabla v)$ denotes the $p$-Laplace operator. Depending on whether the solution lies in the intervals $(a_1,a_2), (a_2,a_3)$ or $(a_1,a_3)$, we employ, respectively, a logarithmic transformation or a hyperbolic tangent transformation to convert the original equation to another one for further analysis. Through a detailed analysis of the lower-bound estimate for the linearized operator of the new equation, and by combining Saloff-Coste's Sobolev inequality with a Moser iteration, we establish Cheng-Yau type gradient estimates under an additional assumption on $p$. As applications, the Liouville theorem and a Harnack inequality are further proved.

math.AP↗

RTLocating: Intent-aware RTL Localization for Hardware Design Iteration

Industrial chip development is inherently iterative, favoring localized, intent-driven updates over rewriting RTL from scratch. Yet most LLM-Aided Hardware Design (LAD) work focuses on one-shot synthesis, leaving this workflow underexplored. To bridge this gap, we for the first time formalize $Δ$Spec-to-RTL localization, a multi-positive problem mapping natural language change requests ($Δ$Spec) to the affected Register Transfer Level (RTL) syntactic blocks. We propose RTLocating, an intent-aware RTL localization framework, featuring a dynamic router that adaptively fuses complementary views from a textual semantic encoder, a local structural encoder, and a global interaction and dependency encoder (GLIDE). To enable scalable supervision, we introduce EvoRTL-Bench, the first industrial-scale benchmark for intent-code alignment derived from OpenTitan's Git history, comprising 1,905 validated requests and 13,583 $Δ$Spec-RTL block pairs. On EvoRTL-Bench, RTLocating achieves 0.568 MRR and 15.08% R@1, outperforming the strongest baseline by +22.9% and +67.0%, respectively, establishing a new state-of-the-art for intent-driven localization in evolving hardware designs.

cs.ET↗

Trinity: A Scenario-Aware Recommendation Framework for Large-Scale Cold-Start Users

Early-stage users in a new scenario intensify cold-start challenges, yet prior works often address only parts of the problem through model architecture. Launching a new user experience to replace an established product involves sparse behavioral signals, low-engagement cohorts, and unstable model performance. We argue that effective recommendations require the synergistic integration of feature engineering, model architecture, and stable model updating. We propose Trinity, a framework embodying this principle. Trinity extracts valuable information from existing scenarios while ensuring predictive effectiveness and accuracy in the new scenario. In this paper, we showcase Trinity applied to a billion-user Microsoft product transition. Both offline and online experiments demonstrate that our framework achieves substantial improvements in addressing the combined challenge of new users in new scenarios.

cs.LG↗