SearcharxivSearch

arXiv subjects

Chujie Chen

Publications and source records attributed to Chujie Chen.

6 recordsLinked to original sources

AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism

Large Language Model (LLM) inference services demand exceptionally high availability and low latency, yet multi-GPU Tensor Parallelism (TP) makes them vulnerable to single-GPU failures. We present AnchorTP, a state-preserving elastic TP framework for fast recovery. It (i) enables Elastic Tensor Parallelism (ETP) with unequal-width partitioning over any number of GPUs and compatibility with Mixture-of-Experts (MoE), and (ii) preserves model parameters and KV caches in GPU memory via a daemon decoupled from the inference process. To minimize downtime, we propose a bandwidth-aware planner based on a Continuous Minimal Migration (CMM) algorithm that minimizes reload bytes under a byte-cost dominance assumption, and an execution scheduler that pipelines P2P transfers with reloads. These components jointly restore service quickly with minimal data movement and without changing service interfaces. In typical failure scenarios, AnchorTP reduces Time to First Success (TFS) by up to 11x and Time to Peak (TTP) by up to 59% versus restart-and-reload.

cs.DC

UAVScenes: A Multi-Modal Dataset for UAVs

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support map-level semantic segmentation due to the lack of frame-wise annotations for both camera images and LiDAR point clouds. This limitation prevents them from being used for high-level scene understanding tasks. To address this gap and advance multi-modal UAV perception, we introduce UAVScenes, a large-scale dataset designed to benchmark various tasks across both 2D and 3D modalities. Our benchmark dataset is built upon the well-calibrated multi-modal UAV dataset MARS-LVIG, originally developed only for simultaneous localization and mapping (SLAM). We enhance this dataset by providing manually labeled semantic annotations for both frame-wise images and LiDAR point clouds, along with accurate 6-degree-of-freedom (6-DoF) poses. These additions enable a wide range of UAV perception tasks, including segmentation, depth estimation, 6-DoF localization, place recognition, and novel view synthesis (NVS). Our dataset is available at https://github.com/sijieaaa/UAVScenes

cs.CV

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages - ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval.

cs.LG

ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning

Large Language Models have emerged as powerful tools for automating Register-Transfer Level (RTL) code generation, yet they face critical limitations: existing approaches typically fail to simultaneously optimize functional correctness and hardware efficiency metrics such as Power, Performance, and Area (PPA). Methods relying on supervised fine-tuning commonly produce functionally correct but suboptimal designs due to the lack of inherent mechanisms for learning hardware optimization principles. Conversely, external post-processing techniques aiming to refine PPA performance after generation often suffer from inefficiency and do not improve the LLMs' intrinsic capabilities. To overcome these challenges, we propose ChipSeek, a novel hierarchical reward based reinforcement learning framework designed to encourage LLMs to generate RTL code that is both functionally correct and optimized for PPA metrics. Our approach integrates direct feedback from EDA simulators and synthesis tools into a hierarchical reward mechanism, facilitating a nuanced understanding of hardware design trade-offs. Through Curriculum-Guided Dynamic Policy Optimization (CDPO), ChipSeek enhances the LLM's ability to generate high-quality, optimized RTL code. Evaluations on standard benchmarks demonstrate ChipSeek's superior performance, achieving state-of-the-art functional correctness and PPA performance. Furthermore, it excels in specific optimization tasks, consistently yielding highly efficient designs when individually targeting fine-grained optimization goals such as power, delay, and area. The artifact is open-source in https://github.com/rong-hash/chipseek.

cs.AI

Search for Astrophysical Neutrino Transients with IceCube DeepCore

DeepCore, as a densely instrumented sub-detector of IceCube, extends IceCube's energy reach down to about 10 GeV, enabling the search for astrophysical transient sources, e.g., choked gamma-ray bursts. While many other past and on-going studies focus on triggered time-dependent analyses, we aim to utilize a newly developed event selection and dataset for an untriggered all-sky time-dependent search for transients. In this work, all-flavor neutrinos are used, where neutrino types are determined based on the topology of the events. We extend the previous DeepCore transient half-sky search to an all-sky search and focus only on short timescale sources (with a duration of $10^2 \sim 10^5$ seconds). All-sky sensitivities to transients in an energy range from 10 GeV to 300 GeV will be presented in this poster. We show that DeepCore can be reliably used for all-sky searches for short-lived astrophysical sources.

astro-ph.HE

A Catalog of Astrophysical Neutrino Candidates for IceCube

Multi-messenger astrophysics will enable the discovery of new astrophysical neutrino sources and provide information about the mechanisms that drive these objects. We present a curated online catalog of astrophysical neutrino candidates. Whenever single high energy neutrino events, that are publicly available, get published multiple times from various analyses, the catalog records all these changes and highlights the best information. All studies by IceCube that produce astrophysical candidates will be included in our catalog. All information produced by these searches such as time, type, direction, neutrino energy and signalness will be contained in the catalog. The multi-messenger astrophysical community will be able to select neutrinos with certain characteristics, e.g. within a declination range, visualize data for the selected neutrinos, and finally download data in their preferred form to conduct further studies.

astro-ph.HE