SearcharxivSearch

arXiv subjects

Yixin Li

Publications and source records attributed to Yixin Li.

At least 19 recordsLinked to original sources

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the resolution of complex real-world SE tasks. However, the trial-and-error nature of these agents generates lengthy interaction trajectories, creating severe bottlenecks in terms of context window limits and cost. While context compression offers a potential remedy, prior approaches suffer from static pruning strategies and granularity mismatches, often failing to preserve the semantic dependencies and syntactic details crucial for SE tasks. To strictly preserve critical task evidence while reducing context length, we introduce AttnCompress, a dynamic attention-guided trajectory compression framework. Unlike existing approaches, AttnCompress bridges the gap between semantic integrity and dynamic adaptability through three key mechanisms: (1) structure-aware segmentation via perplexity (PPL) spikes to preserve the syntactic structure of code and logs; (2) relevance estimation using proxy attention weights to quantify the precise relevance of historical blocks to the agent's current reasoning; and (3) a dynamic rolling window to re-evaluate and recall historical context as the task evolves. Extensive evaluation on SWE-Bench-Verified and Multi-SWE-Bench demonstrates that AttnCompress achieves a pass rate of 53.17%, outperforming prior state-of-the-art baselines while reducing token consumption by 21.6% and total costs by 33.6%. The framework proves to be model-agnostic and generalizes effectively across diverse programming languages.

cs.SE

Enhanced Yield Rate of \textsuperscript{229m}Th via Cascade Decay in Storage Rings and Electron Beam Ion Traps

The low-energy nuclear isomeric state of \textsuperscript{229m}Th provides a unique bridge between nuclear and atomic physics, enabling applications such as nuclear clocks and precision metrology. However, efficient and controllable production of \textsuperscript{229m}Th remains a major experimental challenge. We propose an efficient scheme to produce the $^{229\mathrm{m}}$Th in storage rings (SRs) and electron beam ion traps (EBITs), using a cascade decay pathway. Highly charged ions are excited to higher nuclear states via nuclear excitation by inelastic electron scattering (NEIES) and nuclear excitation by electron capture (NEEC), followed by radiative or internal conversion cascades that populate the isomer. Our calculations demonstrate that, under typical SRs and EBITs conditions, optimized indirect excitation pathways significantly enhance \textsuperscript{229m}Th production rate. In particular, NEIES can provide an enhancement of up to four orders of magnitude through cascade de-excitation at high energies, while NEEC can contribute an additional enhancement of up to several tens of times. Such a significant increase in the \textsuperscript{229m}Th yield rate would facilitate its application in various nuclear photonics fields, especially in the development of atomic nuclear clocks.

nucl-th

SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data

The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQP to benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. % In our quality evaluations, non-private models achieved near-perfect machine-learning efficacy \(\ge0.97\). Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold. Code available at https://github.com/CAN-SYNH/SynQP

cs.LG

Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development

The development of LLM-based autonomous agents for end-to-end software development represents a significant paradigm shift in software engineering. However, the scientific evaluation of these systems is hampered by significant challenges, including overly simplistic benchmarks and the difficulty of conducting fair comparisons between different agent architectures due to confounding implementation variables. To address these limitations, we first construct a challenging and dynamically curated E2EDevBench to simulate realistic development scenarios. Second, we propose a hybrid evaluation framework that combines test-case-based functional assessment with fine-grained, LLM-based requirement verification. Using this framework, we conduct a controlled empirical study on three representative agent architectures implemented upon a unified foundation to isolate the impact of workflow design. Our findings reveal that state-of-the-art agents can fulfill approximately 50\% of requirements on \bench{}, but their success is critically dependent on the architectural strategy for task decomposition and collaboration. Furthermore, our analysis indicates that the primary bottleneck is the omission of requirements and inadequate self-verification. This work provides the community with a more realistic benchmark, a comprehensive evaluation framework, and crucial insights into the current capabilities and core challenges of software development agents, guiding future research toward enhancing requirement comprehension and planning.

cs.SE

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation

Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Language Models (LLMs). Current benchmarks frequently focus on fine-grained code units, lack complete project context, and use inadequate evaluation metrics. To address these limitations, we introduce SWRBench , a new benchmark comprising 1000 manually verified Pull Requests (PRs) from GitHub, offering PR-centric review with full project context. SWRBench employs an objective LLM-based evaluation method that aligns strongly with human judgment (~90 agreement) by verifying if issues from a structured ground truth are covered in generated reviews. Our systematic evaluation of mainstream ACR tools and LLMs on SWRBench reveals that current systems underperform, and ACR tools are more adept at detecting functional errors. Subsequently, we propose and validate a simple multi-review aggregation strategy that significantly boosts ACR performance, increasing F1 scores by up to 43.67%. Our contributions include the SWRBench benchmark, its objective evaluation method, a comprehensive study of current ACR capabilities, and an effective enhancement approach, offering valuable insights for advancing ACR research.

cs.SE

Physically-Based Inverse Rendering Framework for PET Image Reconstruction

Differentiable rendering has been widely adopted in computer graphics as a powerful approach to inverse problems, enabling efficient gradient-based optimization by differentiating the image formation process with respect to millions of scene parameters. Inspired by this paradigm, we propose a physically-based inverse rendering (IR) framework, the first ever platform for PET image reconstruction using Dr.Jit, for PET image reconstruction. Our method integrates Monte Carlo sampling with an analytical projector in the forward rendering process to accurately model photon transport and physical process in the PET system. The emission image is iteratively optimized using voxel-wise gradients obtained via automatic differentiation, eliminating the need for manually derived update equations. The proposed framework was evaluated using both phantom studies and clinical brain PET data acquired from a Siemens Biograph mCT scanner. Implementing the Maximum Likelihood Expectation Maximization (MLEM) algorithm across both the CASToR toolkit and our IR framework, the IR reconstruction achieved a higher signal-to-noise ratio (SNR) and improved image quality compared to CASToR reconstructions. In clinical evaluation compared with the Siemens Biograph mCT platform, the IR reconstruction yielded higher hippocampal standardized uptake value ratios (SUVR) and gray-to-white matter ratios (GWR), indicating enhanced tissue contrast and the potential for more accurate tau localization and Braak staging in Alzheimer's disease assessment. The proposed IR framework offers a physically interpretable and extensible platform for high-fidelity PET image reconstruction, demonstrating strong performance in both phantom and real-world scenarios.

physics.med-ph

An improved peridynamic framework to eliminate unphysical stress and fictitious yield at geometry surface for geomaterials elastoplastic deformation and fracture analysis

This paper presents an improved non-ordinary state-based peridynamics (NOSB PD) framework for modelling the elastoplastic behaviour and damage of geomaterials, such as soil, rock, and concrete, under quasi static conditions. Conventional NOSB PD for elastoplastic materials faces two primary challenges: the surface effect due to the low accuracy of the approximate deformation gradient (FPD) near boundaries and fictitious yielding during explicit time integration. These issues can lead to numerical errors, such as inaccurate crack predictions and potential simulation failure. The proposed framework thoroughly analyses the reason for the surface effect by demonstrating that FPD exhibits only first order accuracy within a horizon radius (delta) from the surface but introduces residual stresses within a larger range of 2 delta with significant surface effect. To mitigate this, a divergence formulation of the non-local differential operator (NDO) is applied to enforce a reasonable stress gradient within a 2 delta subregion, alongside a traction boundary condition consistent with the divergence of stress. Additionally, a loading balance correction algorithm is introduced to enhance the conventional explicit time integration process. The model is validated by integrating it with a modified hyperbolic-hardening Drucker-Prager model, where the stress integration is performed using the closest point projection method (CMMP). Numerical results, compared with finite element simulations and experimental data, demonstrate the effective elimination of the surface effect and false yielding, providing a robust simulation of elastoplastic deformation and progressive failure in geomaterials.

cond-mat.mtrl-sci

Combining Automation and Expertise: A Semi-automated Approach to Correcting Eye Tracking Data in Reading Tasks

In reading tasks drift can move fixations from one word to another or even another line, invalidating the eye tracking recording. Manual correction is time-consuming and subjective, while automated correction is fast yet limited in accuracy. In this paper we present Fix8 (Fixate), an open-source GUI tool that offers a novel semi-automated correction approach for eye tracking data in reading tasks. The proposed approach allows the user to collaborate with an algorithm to produce accurate corrections faster without sacrificing accuracy. Through a usability study (N=14) we assess the time benefits of the proposed technique, and measure the correction accuracy in comparison to manual correction. In addition, we assess subjective workload through NASA Task Load Index, and user opinions through Likert-scale questions. Our results show that on average the proposed technique was 44% faster than manual correction without any sacrifice in accuracy. In addition, users reported a preference for the proposed technique, lower workload, and higher perceived performance compared to manual correction. Fix8 is a valuable tool that offers useful features for generating synthetic eye tracking data, visualization, filters, data converters, and eye movement analysis in addition to the main contribution in data correction.

cs.HC

Study the quantum resolution sizes and atomic bonding states of two-dimensional tin monoxide

Understanding the interatomic bonding and electronic properties of two-dimensional (2D) materials is crucial for preparing high-performance 2D semiconductor materials. We have calculated the band structure, electronic properties, and bonding characteristics of SnO in 2D materials by using density functional theory (DFT) and combining bond energy and bond charge models. Atomic bonding analysis enables us to deeply and meticulously analyze the interatomic bonding and charge transfer in the layered structure of SnO. This study greatly enhances our understanding of the local bonding state on the surface of 2D structural materials. In addition, we use the renormalization method to operate energy to determine the wave function at different quantum resolutions. This is of great significance for describing the size and phase transition of nanomaterials.

cond-mat.mtrl-sci

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains. Our code is available at https://github.com/qq456cvb/3DCorrEnhance.

cs.CV

CXL-Interference: Analysis and Characterization in Modern Computer Systems

Compute Express Link (CXL) is a promising technology that addresses memory and storage challenges. Despite its advantages, CXL faces performance threats from external interference when co-existing with current memory and storage systems. This interference is under-explored in existing research. To address this, we develop CXL-Interplay, systematically characterizing and analyzing interference from memory and storage systems. To the best of our knowledge, we are the first to characterize CXL interference on real CXL hardware. We also provide reverse-reasoning analysis with performance counters and kernel functions. In the end, we propose and evaluate mitigating solutions.

cs.AR

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between tasks and reduce video clarity because the interaction space is limited to the low-level semantic RGB space. To address this issue, we propose an innovative unified framework, SwapTalk, which accomplishes both face swapping and lip synchronization tasks in the same latent space. Referring to recent work on face generation, we choose the VQ-embedding space due to its excellent editability and fidelity performance. To enhance the framework's generalization capabilities for unseen identities, we incorporate identity loss during the training of the face swapping module. Additionally, we introduce expert discriminator supervision within the latent space during the training of the lip synchronization module to elevate synchronization quality. In the evaluation phase, previous studies primarily focused on the self-reconstruction of lip movements in synchronous audio-visual videos. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experimental results on the HDTF demonstrate that our method significantly surpasses existing techniques in video quality, lip synchronization accuracy, face swapping fidelity, and identity consistency. Our demo is available at http://swaptalk.cc.

cs.CV

Distributionally Robust Evaluation for Real-Time Flexibility of Electric Vehicles Considering Uncertain Departure Behavior and State-of-Charge

Accurately evaluating the real-time flexibility of electric vehicles (EVs) is necessary for EV aggregators to offer ancillary services. However, regulation-caused uncertain state-of-charge and random departure behavior complicate the evaluation and badly impact the evaluation accuracy. To resolve this issue, this letter proposes a distributionally robust real-time flexibility evaluation model that formulates the uncertain departure behavior and state-of-charge of EVs in an online updating pattern. Thanks to dualization, this model can be efficiently solved via off-the-shelf solvers. Case studies validate the superiority of the proposed method and its scalability regarding EV numbers.

eess.SY

Spin fluctuations and charge properties of core shell C$_{80}$+M$_{13}$ (V, Mn, Cr, Ni, Co)

Transition metal clusters have a broad spectrum of potential applications in electronic and magnetic devices owing to their unique properties. Protective shells such as fullerene C$_{80}$ can be introduced to improve their stability. In this study, we optimized five core shell structures, C$_{80}$+M$_{13}$ (V, Mn, Cr, Ni, Co), and calculated their electromagnetic properties using density functional theory.We determined that there is electron transfer between C$_{80}$ and the transition metal clusters near the Fermi surface, and that the d orbitals contribute most to the magnetism of the structure. C$_{80}$+Ni$_{13}$ was antiferromagnetic. The magnetic properties of the clusters were significantly altered, revealing antiferromagnetism. The results establish a theoretical starting point for tuning the electronic and magnetic properties of 13-atom clusters embedded in fullerene cages.

cond-mat.mtrl-sci

Topological Bonding and Electronic properties of Cd$_{43}$Te$_{28}$ semiconductor material with microporous structure

CdTe is II-VI semiconductor material with excellent characteristics and has demonstrated promising potential for application in the photovoltaic field. The electronic properties of Cd43Te28 with microporous structures have been investigated based on density functional theory. The newly established binding-energy and bond-charge model have been used to convert the value of Hamiltonian into bonding values. We provide a method for describing topological chemical bonds by atomic coordinates and wave phases. We also discuss the dynamic process of the wave function with time and the magic cube matrix. This study provides an innovative method and technology for the accurate analysis of the topological bonding and electronic properties of microporous semiconductor materials.

cond-mat.mtrl-sci

Characterizing Spectral Properties of Bridge

The Bridge graph is a special type of graph which are constructed by connecting identical connected graphs with path graphs. We discuss different types of bridge graphs $B_{n\times l}^{m\times k}$ in this paper. In particular, we discuss the following: complete-type bridge graphs, star-type bridge graphs, and full binary tree bridge graphs. We also bound the second eigenvalues of the graph Laplacian of these graphs using methods from Spectral Graph Theory. In general, we prove that for general bridge graphs, $B_{n\times l}^2$, the second eigenvalue of the graph Laplacian should be between $0$ and $2$, inclusive. In the end, we talk about future work on infinite bridge graphs. We created definitions and found the related theorems to support our future work about infinite bridge graphs.

math.CO

Blockchain for Data Sharing at the Network Edge: Trade-Off Between Capability and Security

Blokchain is a promising technology to enable distributed and reliable data sharing at the network edge. The high security in blockchain is undoubtedly a critical factor for the network to handle important data item. On the other hand, according to the dilemma in blockchain, an overemphasis on distributed security will lead to poor transaction-processing capability, which limits the application of blockchain in data sharing scenarios with high-throughput and low-latency requirements. To enable demand-oriented distributed services, this paper investigates the relationship between capability and security in blockchain from the perspective of block propagation and forking problem. First, a Markov chain is introduced to analyze the gossiping-based block propagation among edge servers, which aims to derive block propagation delay and forking probability. Then, we study the impact of forking on blockchain capability and security metrics, in terms of transaction throughput, confirmation delay, fault tolerance, and the probability of malicious modification. The analytical results show that with the adjustment of block generation time or block size, transaction throughput improves at the sacrifice of fault tolerance, and vice versa. Meanwhile, the decline in security can be offset by adjusting confirmation threshold, at the cost of increasing confirmation delay. The analysis of capability-security trade-off can provide a theoretical guideline to manage blockchain performance based on the requirements of data sharing scenarios.

cs.DC

Predict the water level of the Lake Mead for the next 30 years based on ARIMA

In this study, a mathematical model is developed for the drought problem of Lake Mead. First, a polynomial fitting of the elevation of Lake Mead to the area of the lake is done by the least-squares method, and the volume of Lake Mead is approximated by the numerical integration of the product of the height and the area solved by the trapezoidal rule. The accuracy of the fitting reached more than 96%at all four different locations. Second, the minimum and maximum water levels were transformed into volume numbers by the above method, and the historical data of Lake Mead were classified into three classes of water resources by sequential clustering. According to these data, the optimal cut point of the most recent drought period was 2008 and has continued until now. Finally, two prediction models were constructed using ARIMA(2,2,2) and ARIMA(3,2,2) to study the water level data from 2008 to 2020 and 2005 to 2020, respectively, to predict the water level data of Lake Mead from 2022 to 2050, and to compare and analyze them.

physics.ao-ph