SearcharxivSearch

arXiv subjects

Kenan Li

Publications and source records attributed to Kenan Li.

14 recordsLinked to original sources

SLAMFormer-$\infty$: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing

We introduce the Infinite SLAM Transformer (SLAMFormer-$\infty$), the first geometric transformer capable of supporting both long-range frontend and backend processing without an explicit distance bound. Instead of relying on a first-frame-anchored formulation, SLAMFormer-$\infty$ employs memory conditions to define flexible coordinate systems and scales for input frames, enabling more expressive structural conditioning. Built upon this formulation, the frontend preserves efficient local computation, while the backend jointly optimizes long-range trajectories and scene geometry in a globally consistent manner. Experimental results demonstrate that SLAMFormer-$\infty$ achieves superior or highly competitive performance in both trajectory estimation and scene reconstruction across large-scale datasets. Notably, SLAMFormer-$\infty$ generalizes to extremely long trajectories, successfully operating on sequences exceeding $17\mathrm{km}$.

cs.CV

ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance

Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a continuous maintenance workflow into a series of independent sessions, ignoring the cumulative dependencies that make real-world bug fixing challenging. To bridge this gap, we introduce ChainSWE, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase. We collect chronological chains of 304 issues across 54 Python projects, mined from six SWE-bench-family datasets. Our evaluation across a range of agents and models reveals a consistent performance drop by up to 70% as the chain length increases.

cs.SE

SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

Large language model agents have made strong progress on software engineering, yet current systems suffer from a context coupling problem: the standard code editing interface conflates code inspection, modification planning, and edit execution within a single context window, forcing agents to interleave exploratory viewing with strictly formatted edit generation. Irrelevant context accumulates and edit reliability degrades. We propose SWE-Edit, which decomposes the editing interface into two specialized subagents: a Viewer that extracts task-relevant code on demand, and an Editor that executes modifications from high-level natural language plans -- letting the main agent focus on reasoning while delegating context-intensive operations to clean context windows. On SWE-Bench Verified, this decomposition raises resolve rate by 2.1 pp and cuts inference cost by 17.9%, with consistent gains across multiple reasoning-model families (Kimi-K2, MiniMax-M2.1, GLM-4.7). We further show that effective edit-format selection can be trained into a small model rather than requiring frontier-scale capacity: GRPO training on Qwen3-8B with an adaptive find-replace/whole-file-rewrite policy improves edit success by 12.5 pp and brings an 8B open-source editor to parity with GPT-5-nano on downstream SWE-Bench resolve rate. To enable rapid editor iteration, we release PR-Edit, a lightweight evaluation whose scores correlate strongly with SWE-Bench resolve rate. We release our code at https://github.com/microsoft/SWE-Edit.

cs.SE

ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents

Recent advances in language model (LM) agents have significantly improved automated software engineering (SWE). Prior work has proposed various agentic workflows and training strategies as well as analyzed failure modes of agentic systems on SWE tasks, focusing on several contextual information signals: Reproduction Test, Regression Test, Edit Location, Execution Context, and API Usage. However, the individual contribution of each signal to overall success remains underexplored, particularly their ideal contribution when intermediate information is perfectly obtained. To address this gap, we introduce Oracle-SWE, a unified method to isolate and extract oracle information signals from SWE benchmarks and quantify the impact of each signal on agent performance. To further validate the pattern, we evaluate the performance gain of signals extracted by strong LMs when provided to a base agent, approximating real-world task-resolution settings. These evaluations aim to guide research prioritization for autonomous coding systems.

cs.MA

Complet4R: Geometric Complete 4D Reconstruction

We introduce Complet4R, a novel end-to-end framework for Geometric Complete 4D Reconstruction, which aims to recover temporally coherent and geometrically complete reconstruction for dynamic scenes. Our method formalizes the task of Geometric Complete 4D Reconstruction as a unified framework of reconstruction and completion, by directly accumulating full contexts onto each frame. Unlike previous approaches that rely on pairwise reconstruction or local motion estimation, Complet4R utilizes a decoder-only transformer to operate all context globally directly from sequential video input, reconstructing a complete geometry for every single timestamp, including occluded regions visible in other frames. Our method demonstrates the state-of-the-art performance on our proposed benchmark for Geometric Complete 4D Reconstruction and the 3D Point Tracking task. Code will be released to support future research.

cs.CV

RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel agentic framework that automatically resolves dependencies, compiles source code, and extracts test results across diverse programming languages and operating systems. RepoLaunch achieves a 78% build success rate, outperforming the Python/Linux-only prior system by 18%. To demonstrate its application, we further present a fully automated pipeline for SWE dataset creation driven by RepoLaunch, which only requires human input at the task-design stage. RepoLaunch is open-sourced, and its automated task-generation pipeline has already been adopted by several recent works on agentic benchmarking and training.

cs.SE

ChemBART: A Pre-trained BART Model Assisting Organic Chemistry Analysis

Recent advances in large language models (LLMs) have demonstrated transformative potential across diverse fields. While LLMs have been applied to molecular simplified molecular input line entry system (SMILES) in computer-aided synthesis planning (CASP), existing methodologies typically address single tasks, such as precursor prediction. We introduce ChemBART, a SMILES-based LLM pre-trained on chemical reactions, which enables a unified model for multiple downstream chemical tasks--achieving the paradigm of "one model, one pre-training, multiple tasks." By leveraging outputs from a mask-filling pre-training task on reaction expressions, ChemBART effectively solves a variety of chemical problems, including precursor/reagent generation, temperature-yield regression, molecular property classification, and optimizing the policy and value functions within a reinforcement learning framework, integrated with Monte Carlo tree search for multi-step synthesis route design. Unlike single-molecule pre-trained LLMs constrained to specific applications, ChemBART addresses broader chemical challenges and integrates them for comprehensive synthesis planning. Crucially, ChemBART-designed multi-step synthesis routes and reaction conditions directly inspired wet-lab validation, which confirmed shorter pathways with ~30% yield improvement over literature benchmarks. Our work validates the power of reaction-focused pre-training and showcases the broad utility of ChemBART in advancing the complete synthesis planning cycle.

cs.LG

SLAM-Former: Putting SLAM into One Transformer

We present SLAM-Former, a neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLAM-Former comprises both a frontend and a back-end that operate in tandem. The frontend processes sequential monocular images in real-time for incremental mapping and tracking, while the backend performs global refinement to ensure a geometrically consistent result. This alternating execution allows the frontend and back-end to mutually promote one another, enhancing overall system performance. Comprehensive experimental results demonstrate that SLAM- Former achieves superior or highly competitive performance compared to state-of-the-art dense SLAM methods.

cs.CV

GPx4 is bound to peroxidized membranes by a hydrophobic anchor

Ferroptosis is a form of cell death discovered in recent years, induced by excessive peroxidation of phospholipids. Glutathione peroxidase 4 (GPx4) is an intracellular enzyme that can repair the peroxidized phospholipids on membranes, thus regulating ferroptosis. By combining multiscale molecular dynamics (MD) simulations and experimental assays, we investigate the binding mechanisms of GPx4 on membranes. Using coarse-grained MD simulations, we found that L130 and its adjacent residues on GPx4 can form a stable and unique binding interface with PE/PS-rich and peroxidized membranes. Subsequent all-atom MD simulations verified the stability of the binding interface. The critical residue on the interface, L130, was inserted deeply into the membrane as a hydrophobic anchor and guided the reaction center toward the membrane surface. Enzyme activity assays and in vitro cell experiments showed that mutations of L130 resulted in weaker activities of the enzyme, probably caused by non-functional binding modes of GPx4 on membranes, as revealed by in silico simulations. This study highlights the crucial role of the hydrophobic residue, L130, in the proper anchoring of GPx4 on membranes, the first step of its membrane-repairing function.

q-bio.BM

TrackOcc: Camera-based 4D Panoptic Occupancy Tracking

Comprehensive and consistent dynamic scene understanding from camera input is essential for advanced autonomous systems. Traditional camera-based perception tasks like 3D object tracking and semantic occupancy prediction lack either spatial comprehensiveness or temporal consistency. In this work, we introduce a brand-new task, Camera-based 4D Panoptic Occupancy Tracking, which simultaneously addresses panoptic occupancy segmentation and object tracking from camera-only input. Furthermore, we propose TrackOcc, a cutting-edge approach that processes image inputs in a streaming, end-to-end manner with 4D panoptic queries to address the proposed task. Leveraging the localization-aware loss, TrackOcc enhances the accuracy of 4D panoptic occupancy tracking without bells and whistles. Experimental results demonstrate that our method achieves state-of-the-art performance on the Waymo dataset. The source code will be released at https://github.com/Tsinghua-MARS-Lab/TrackOcc.

cs.CV

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details from RGB input. However, progress in SSC, particularly in large-scale street views, is hindered by the scarcity of high-quality datasets. To address this issue, we introduce SSCBench, a comprehensive benchmark that integrates scenes from widely used automotive datasets (e.g., KITTI-360, nuScenes, and Waymo). SSCBench follows an established setup and format in the community, facilitating the easy exploration of SSC methods in various street views. We benchmark models using monocular, trinocular, and point cloud input to assess the performance gap resulting from sensor coverage and modality. Moreover, we have unified semantic labels across diverse datasets to simplify cross-domain generalization testing. We commit to including more datasets and SSC models to drive further advancements in this field.

cs.CV

A compact single-shot soft X-ray photon spectrometer for free electron laser diagnostics

The photon spectrum from free-electron laser (FEL) light sources offers valuable information in time-resolved experiments and machine optimization in the spectral and temporal domains. We have developed a compact single-shot photon spectrometer to diagnose soft X-ray spectra. The spectrometer consists of an array of off-axis Fresnel zone plates (FZP) that act as transmission-imaging gratings, a Ce-YAG scintillator, and a microscope objective to image the scintillation target onto a two-dimensional imaging detector. This spectrometer operates in an energy range which covers absorption edges associated with several atomic constituents carbon, nitrogen, oxygen, and neon. The spectrometer's performance is demonstrated at a repetition rate of 120 Hz, but our detection scheme can be easily extended to 200 kHz spectral collection by employing a fast complementary metal oxide semiconductor (CMOS) line-scan camera to detect the light from the scintillator. This compact photon spectrometer provides an opportunity for monitoring the spectrum downstream of an endstation in a limited space environment with subelectronvolt energy resolution.

physics.ins-det

On the Uniqueness of Functions that Maximize the Crouzeix Ratio

Let $A$ be an $n$ by $n$ matrix with numerical range $W(A) := \{ q^{*}Aq : q \in \mathbb{C}^n , ~\| q \|_2 = 1 \}$. We are interested in functions $\hat{f}$ that maximize $\| f(A) \|_2$ (the matrix norm induced by the vector 2-norm) over all functions $f$ that are analytic in the interior of $W(A)$ and continuous on the boundary and satisfy $\max_{z \in W(A)} | f(z) | \leq 1$. It is known that there are functions $\hat{f}$ that achieve this maximum and that such functions are of the form $B\circϕ$, where $ϕ$ is any conformal mapping from the interior of $W(A)$ to the unit disk $\mathbb{D}$, extended to be continuous on the boundary of $W(A)$, and $B$ is a Blaschke product of degree at most $n-1$. It is not known if a function $\hat{f}$ that achieves this maximum is unique, up to multiplication by a scalar of modulus one. We show that this is the case when $A$ is a $2\times 2$ nonnormal matrix or a Jordan block, but we give examples of some $3\times 3$ matrices with elliptic numerical range for which two different functions $\hat{f}$, involving the same conformal mapping but Blaschke products of different degrees, achieve the same maximal value of $||f(A)||_2$.

math.CV

Some Extensions of the Crouzeix-Palencia Result

In [{\em The Numerical Range is a $(1 + \sqrt{2})$-Spectral Set}, SIAM J. Matrix Anal. Appl. 38 (2017), pp.~649-655], Crouzeix and Palencia show that the numerical range of a square matrix or linear operator $A$ is a $(1 + \sqrt{2})$-spectral set for $A$; that is, for any function $f$ analytic in the interior of the numerical range $W(A)$ and continuous on its boundary, the inequality $\| f(A) \| \leq (1 + \sqrt{2} ) \| f \|_{W(A)}$ holds, where the norm on the left is the operator 2-norm and $\| f \|_{W(A)}$ on the right denotes the supremum of $| f(z) |$ over $z \in W(A)$. In this paper, we show how the arguments in their paper can be extended to show that other regions in the complex plane that do {\em not} necessarily contain $W(A)$ are $K$-spectral sets for a value of $K$ that may be close to $1 + \sqrt{2}$. We also find some special cases in which the constant $(1 + \sqrt{2})$ for $W(A)$ can be replaced by $2$, which is the value conjectured by Crouzeix.

math.NA