SearcharxivSearch

arXiv subjects

Tingting Yu

Publications and source records attributed to Tingting Yu.

At least 19 recordsLinked to original sources

HyperFL: Query-Adaptive Representation Learning for Software Fault Localization

Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugging and program repair. Recent retrieval-based approaches formulate fault localization as a dense retrieval task by learning a shared embedding space between issue reports and source code. However, these methods encode all issue reports using a fixed query representation, despite the substantial diversity of real-world issue reports in length, structure, and debugging information. To address this limitation, we propose HyperFL, a query-adaptive representation learning framework for software fault localization. HyperFL employs a lightweight hypernetwork to generate query-specific LoRA parameters for the query encoder, enabling dynamic query adaptation while keeping the code encoder fixed and reusable. Experiments on a real-world issue localization benchmark demonstrate that HyperFL consistently improves retrieval performance across multiple embedding backbones, achieving up to 13.3% relative improvement in function-level MRR@10 and 16.7% relative improvement in Hit@1 over the state-of-the-art method SweRank. Further analysis shows that HyperFL learns distinct adaptation patterns for different issue characteristics, highlighting the effectiveness of query-adaptive representations for software issue localization.

cs.SE

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert. To address this gap, we propose MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input-specific low-rank updates. This design decouples expert capacity from adapter storage while enabling instance-conditioned adaptation. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE-based PEFT baselines across three backbones. MoEGen also performs strongly in joint medical and legal-domain adaptation.

cs.CL

ConFL: Explainable Concurrent Fault Localization via Hierarchy-Guided LLM Reasoning

Localizing concurrent bugs from bug reports alone is challenging due to incomplete information, misleading program-entity mentions, and complex cross-thread interactions, causing existing LLM-based approaches to suffer from unstable reasoning and limited explainability. We propose ConFL, an explainable concurrent fault localization framework that augments LLM reasoning with structured concurrency knowledge. ConFL constructs a Concurrent Knowledge Base (CKB) from source code and performs LLM-guided hierarchical retrieval to progressively narrow the search space from components to interaction-level concurrency contexts. An interaction-level DSL explicitly encodes cross-thread interactions over shared resources, enabling focused reasoning without traversing deep call chains. Experiments on real-world concurrent bugs from eight large-scale Java projects show that ConFL significantly outperforms state-of-the-art IR-based and LLM-based baselines, achieving an MRR of 0.503 and a MAP of 0.486, while remaining robust to noisy bug reports, unseen bugs, and different LLM backbones.

cs.SE

ADVENT: LLM-Driven Automatic Predicate Invention for ILP

Predicate invention (PI), the creation of new predicates to extend the hypothesis space, remains a critical bottleneck in Inductive Logic Programming (ILP). Existing methods rely on domain expertise and produce semantically opaque predicates, hindering adaptation to unfamiliar domains and cross-task reuse. We present ADVENT, an LLM-driven PI mechanism for ILP. ADVENT pairs LLM abductive generation with Prolog deductive verification, forming an iterative loop in which concrete execution results guide the LLM to refine candidate predicates. The mechanism leverages Large Language Models to identify implicit patterns in structured relational data and invent auxiliary predicates with meaningful names and definitions. Invented predicates and learned rules accumulate in a knowledge pool for cross-task reuse. Experiments on nine poker-hand concepts across seven LLMs show that LLM-driven PI achieves 58% success rate where ILP alone fails entirely, formal verification raises this to 80%, and the knowledge pool yields gains up to +31 percentage points, while producing human-interpretable rules. These results suggest that ADVENT offers a promising direction for automating predicate invention and enabling cross-task knowledge reuse in ILP.

cs.LO

Structural relaxation and stagnation of grain boundary during migration

Instability is a major bottleneck in nanomaterials due to grain boundary (GB) activities under thermal or mechanical stimuli. The relaxation of GB will stabilize the properties of materials by structure modification of GBs to lower energy states. However, lack of understanding of mechanisms limits the application of GB relaxation. In this study, we certify that GB can realize relaxation through defect emission including vacancy, twinning and dislocations, etc, which lower the average atomic energy of GB. In particular, we found stagnation and shear-coupling migration accompany the relaxation process, where a lower average atomic energy and lower average atomic volume of grain boundaries can be the reasons for stagnation. In this study, we found when simulated by ramped energy-conserving orientated driving force, under a specific driving force, the grain boundary suddenly stops migrating during the migration process, and as the driving force increases to a certain value, the grain boundary continues to migrate. Even at fixed driving forces, the grain boundary migration process can stall. This phenomenon has been found in many grain boundaries, and it has been found that the reason for the stagnation is the change of the average atomic energy of grain boundaries and the average atomic volume of grain boundaries. The discovery of this stagnation phenomenon is helpful to better understand the migration characteristics of grain boundaries and lays a foundation for improving the strength of materials by designing grain boundary microstructures.

cond-mat.mtrl-sci

Latent Confidence Alignment for LLM Self-Assessment

Confidence calibration in large language models (LLMs) is commonly evaluated by comparing predicted confidence with observed accuracy. However, such approaches do not model item difficulty, making it difficult to interpret discrepancies and to determine whether model confidence reflects genuine self-assessment or is merely a byproduct of the response generation process. To address this, we adopt a Rasch model-based latent ability framework and a metacognitive perspective, and propose Latent Confidence Alignment Error (LCAE) to measure the consistency between model self-assessment and the latent error probability implied by model ability and item difficulty. We further incorporate item difficulty as an external signal with a reasoning mechanism. Experiments on a medical-domain dataset with 20 models show that the proposed approach improves self-assessment quality without affecting model ability, and reveals an association between reliability and inference cost.

cs.CY

Passive repetition-rate stabilization for a mode-locked fiber laser by electro-optic modulation

We report a passive stabilization of the repetition rate for a mode-locked fiber laser by using an electro-optic modulator in a phase-biased nonlinear amplifying loop mirror. The underlying mechanism, in contrast to active feedback operations, lies in the cross-phase modulation between electrical and optical pulses within an electro-optic crystal. The resulting spectral shift can automatically compensate the cavity-length drift via the group velocity dispersion. Consequently, the artificial actuator enables to obtain a capture range up to 2.3 mm, much longer than that achieved by index changes of the modulator. A robust and tight locking for the repetition rate is then realized with a standard deviation as low as 9 $\mu$Hz with a 1-s sample time over 11 hours, corresponding to a fractional instability of 4.3$\times$10$^{-13}$. Furthermore, a dynamic optical sampling by repetition-rate tuning has been manifested with a fast refresh rate at 100 kHz and a broad scanning range over 305 ps. The demonstrated passive servo action may provide a simple yet effective way to stabilize the repetition rate with high precision, large bandwidth and wide tunability.

physics.optics

High-precision passive stabilization of repetition rate for a mode-locked fiber laser based on optical pulse injection

We have proposed and implemented a novel scheme to obtain high-precision repetition rate stabilization for a polarization-maintaining mode-locked fiber laser. The essential technique lies in the periodic injection of electronically modulated optical pulses into a nonlinear amplifying loop mirror within the laser resonator. Thanks to the nonlinear cross-phase modulation effect, the injected pulses referenced to an external clock serves as a stable and precise timing trigger for an effective intensity modulator. Consequently, synchronous mode-locking can be initiated to output ultrafast pulses with a passively stabilized repetition rate. The capture range of the locking system reaches to a record of 1 mm, which enables a long-term stable operation over 15 hours without the need of temperature stabilization and vibration isolation. Meanwhile, the achieved standard deviation is as low as 100 $\mu$Hz with a 1-s sample time, corresponding to a fluctuation instability of 5.0$\times10^{-12}$. Additionally, the repetition rate stabilization performance based on the passive synchronization has been systematically investigated by varying the average power, central wavelength and pulse duration of the optical injection.

physics.optics

High-resolution mid-infrared single-photon upconversion ranging

Single-photon laser ranging has widespread applications in remote sensing and target recognition. However, highly-sensitive light detection and ranging (LiDAR) has long been restricted in visible or near-infrared bands. An appealing quest is to extend the operation wavelength into the mid-infrared (MIR) region, which calls for an infrared photon counting system at high detection sensitivity and precise temporal resolution. Here, we devise and demonstrate a MIR upconversion LiDAR based on nonlinear asynchronous optical sampling. Specifically, the infrared probe is interrogated in a nonlinear crystal by a train of pump pulses at a slightly different repetition rate, which favors for a temporal optical scanning at a picosecond timing resolution and a kilohertz refreshing rate over $\sim$50 ns. Moreover, the cross-correlation upconversion trace is temporally stretched by a factor of 2$\times$10$^4$, which can thus be recorded by a low-bandwidth silicon detector. In combination with time-correlated photon-counting technique, the achieved effective resolution is about two orders of magnitude better than the timing jitter of the detector itself, which facilitates a ranging precision of 4 $\mu$m under a low detected flux of 8$\times$10$^{-5}$ photons per pulse. The presented MIR time-of-flight range finder is featured with single-photon sensitivity and high positioning resolution, which would be particularly useful in infrared sensing and imaging in photon-starved scenarios.

physics.optics

Widely tunable mid-infrared fiber-feedback optical parametric oscillator

Synchronously pumped optical parametric oscillators (OPOs) provide uniquely versatile platforms to generate ultrafast mid-infrared pulses within a spectral range beyond the access of conventional mode-locked lasers. However, conventional OPO sources based on bulk crystals have been plagued by complex optical alignment and large physical footprint. Here, we devise and implement two OPO variants based on a polarization-maintaining fiber-feedback cavity, which allow to robustly deliver sub-picosecond MIR pulses without the need of active stabilization. The first one integrates an erbium-doped fiber into the OPO cavity as the additional gain medium, which significantly reduces the pump threshold and allows stable optical pulse formation within a spectral range of 1553-1586 nm. The second one adopts a chirped poling nonlinear crystal in a passive-fiber cavity to further extend the operation spectral coverage, which facilitates broad tuning ranges of 1350-1768 nm and 2450-4450 nm for the signal and idler bands, respectively. Therefore, the presented mid-infrared OPO source is featured with high compactness, robust operation, and wide tunability, which would be attractive for subsequent applications such as infrared photonics, biomedical examination, and molecular spectroscopy.

physics.optics

ViBR: Automated Bug Replay from Video-based Reports using Vision-Language Models

Bug reports play a critical role in software maintenance by helping users convey encountered issues to developers. Recently, GUI screen capture videos have gained popularity as a bug reporting artifact due to their ease of use and ability to retain rich contextual information. However, automatically reproducing bugs from such recordings remains a significant challenge. Existing methods often rely on fragile image-processing heuristics, explicit touch indicators, or pre-constructed UI transition graphs, which require non-trivial instrumentation and app-specific setup. This paper presents ViBR, a lightweight and fully automated approach that reproduces bugs directly from GUI recordings. Specifically, ViBR combines CLIP-based embedding similarity for action boundary segmentation with Vision-Language Models (VLMs) for region-aware GUI state comparison and guided bug replay. Experimental results show that ViBR successfully reproduces 72% of bug recordings, significantly outperforming state-of-the-art baselines and ablation variants.

cs.SE

QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models

While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to capture the unstructured, ambiguous, and long-tail complexities inherent in genuine user inquiries. To bridge this gap, we introduce QuarkMedBench, an ecologically valid benchmark tailored for real-world medical LLM assessment. We compiled a massive dataset spanning Clinical Care, Wellness Health, and Professional Inquiry, comprising 20,821 single-turn queries and 3,853 multi-turn sessions. To objectively evaluate open-ended answers, we propose an automated scoring framework that integrates multi-model consensus with evidence-based retrieval to dynamically generate 220,617 fine-grained scoring rubrics (~9.8 per query). During evaluation, hierarchical weighting and safety constraints structurally quantify medical accuracy, key-point coverage, and risk interception, effectively mitigating the high costs and subjectivity of human grading. Experimental results demonstrate that the generated rubrics achieve a 91.8% concordance rate with clinical expert blind audits, establishing highly dependable medical reliability. Crucially, baseline evaluations on this benchmark reveal significant performance disparities among state-of-the-art models when navigating real-world clinical nuances, highlighting the limitations of conventional exam-based metrics. Ultimately, QuarkMedBench establishes a rigorous, reproducible yardstick for measuring LLM performance on complex health issues, while its framework inherently supports dynamic knowledge updates to prevent benchmark obsolescence.

cs.CL

Identifying Concurrency Bug Reports via Linguistic Patterns

With the growing ubiquity of multi-core architectures, concurrent systems have become essential but increasingly prone to complex issues such as data races and deadlocks. While modern issue-tracking systems facilitate the reporting of such problems, labeling concurrency-related bug reports remains a labor-intensive and error-prone task. This paper presents a linguistic-pattern-based framework for automatically identifying concurrency bug reports. We derive 58 distinct linguistic patterns from 730 manually labeled concurrency bug reports, organized across four levels: word-level (keywords), phrase-level (n-grams), sentence-level (semantic), and bug report-level (contextual). To assess their effectiveness, we evaluate four complementary approaches-matching, learning, prompt-based, and fine-tuning-spanning traditional machine learning, large language models (LLMs), and pre-trained language models (PLMs). Our comprehensive evaluation on 12 large-scale open-source projects (10,920 issue reports from GitHub and Jira) demonstrates that fine-tuning PLMs with linguistic-pattern-enriched inputs achieves the best performance, reaching a precision of 91% on GitHub and 93% on Jira, and maintaining strong precision on post cut-off data (91%). The contributions of this work include: (1) a comprehensive taxonomy of linguistic patterns for concurrency bugs, (2) a novel fine-tuning strategy that integrates domain-specific linguistic knowledge into PLMs, and (3) a curated, labeled dataset to support reproducible research. Together, these advances provide a foundation for improving the automation, precision, and interpretability of concurrency bug classification.

cs.SE

SysPro: Reproducing System-level Concurrency Bugs from Bug Reports

Reproducing system-level concurrency bugs requires both input data and the precise interleaving order of system calls. This process is challenging because such bugs are non-deterministic, and bug reports often lack the detailed information needed. Additionally, the unstructured nature of reports written in natural language makes it difficult to extract necessary details. Existing tools are inadequate to reproduce these bugs due to their inability to manage the specific interleaving at the system call level. To address these challenges, we propose SysPro, a novel approach that automatically extracts relevant system call names from bug reports and identifies their locations in the source code. It generates input data by utilizing information retrieval, regular expression matching, and the category-partition method. This extracted input and interleaving data are then used to reproduce bugs through dynamic source code instrumentation. Our empirical study on real-world benchmarks demonstrates that SysPro is both effective and efficient at localizing and reproducing system-level concurrency bugs from bug reports.

cs.SE

HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks

Instruction-based text editing is increasingly critical for real-world applications such as code editors (e.g., Cursor), but Large Language Models (LLMs) continue to struggle with this task. Unlike free-form generation, editing requires faithfully implementing user instructions while preserving unchanged content, as even minor unintended modifications can break functionality. Existing approaches treat editing as generic text generation, leading to two key failures: they struggle to faithfully align edits with diverse user intents, and they often over-edit unchanged regions. We propose HyperEdit to address both issues. First, we introduce hypernetwork-based dynamic adaptation that generates request-specific parameters, enabling the model to tailor its editing strategy to each instruction. Second, we develop difference-aware regularization that focuses supervision on modified spans, preventing over-editing while ensuring precise, minimal changes. HyperEdit achieves a 9%--30% relative improvement in BLEU on modified regions over state-of-the-art baselines, despite utilizing only 3B parameters.

cs.CL

A Fully Automatic Framework for Intracranial Pressure Grading: Integrating Keyframe Identification, ONSD Measurement and Clinical Data

Intracranial pressure (ICP) elevation poses severe threats to cerebral function, thus necessitating monitoring for timely intervention. While lumbar puncture is the gold standard for ICP measurement, its invasiveness and associated risks drive the need for non-invasive alternatives. Optic nerve sheath diameter (ONSD) has emerged as a promising biomarker, as elevated ICP directly correlates with increased ONSD. However, current clinical practices for ONSD measurement suffer from inconsistency in manual operation, subjectivity in optimal view selection, and variability in thresholding, limiting their reliability. To address these challenges, we introduce a fully automatic two-stage framework for ICP grading, integrating keyframe identification, ONSD measurement and clinical data. Specifically, the fundus ultrasound video processing stage performs frame-level anatomical segmentation, rule-based keyframe identification guided by an international consensus statement, and precise ONSD measurement. The intracranial pressure grading stage then fuses ONSD metrics with clinical features to enable the prediction of ICP grades, thereby demonstrating an innovative blend of interpretable ultrasound analysis and multi-source data integration for objective clinical evaluation. Experimental results demonstrate that our method achieves a validation accuracy of $0.845 \pm 0.071$ (with standard deviation from five-fold cross-validation) and an independent test accuracy of 0.786, significantly outperforming conventional threshold-based method ($0.637 \pm 0.111$ validation accuracy, $0.429$ test accuracy). Through effectively reducing operator variability and integrating multi-source information, our framework establishes a reliable non-invasive approach for clinical ICP evaluation, holding promise for improving patient management in acute neurological conditions.

cs.CV

TreeDiff: AST-Guided Code Generation with Diffusion LLMs

Code generation is increasingly critical for real-world applications. Still, diffusion-based large language models continue to struggle with this demand. Unlike free-form text, code requires syntactic precision; even minor structural inconsistencies can render a program non-executable. Existing diffusion-based large language models rely on random token masking for corruption, leading to two key failures: they lack awareness of syntactic boundaries during the iterative denoising process, and they fail to capture the long-range hierarchical dependencies essential for program correctness. We propose TreeDiff to address both issues. Specifically, we propose a syntax-aware diffusion framework that incorporates structural priors from Abstract Syntax Tree (AST) into the corruption process. Instead of masking individual tokens at random, we selectively mask tokens belonging to key AST nodes. By aligning the corruption process with the underlying structure of code, our method encourages the model to internalize the compositional nature of programming languages, enabling it to reconstruct programs that respect grammatical boundaries and capture long-range dependencies. Our method achieves a 13.3% relative improvement over the random masking training method, demonstrating its effectiveness in code generation task by leveraging underlying structures.

cs.CL

An Empirical Study on Leveraging Images in Automated Bug Report Reproduction

Automated bug reproduction is a challenging task, with existing tools typically relying on textual steps-to-reproduce, videos, or crash logs in bug reports as input. However, images provided in bug reports have been overlooked. To address this gap, this paper presents an empirical study investigating the necessity of including images as part of the input in automated bug reproduction. We examined the characteristics and patterns of images in bug reports, focusing on (1) the distribution and types of images (e.g., UI screenshots), (2) documentation patterns associated with images (e.g., accompanying text, annotations), and (3) the functional roles they served, particularly their contribution to reproducing bugs. Furthermore, we analyzed the impact of images on the performance of existing tools, identifying the reasons behind their influence and the ways in which they can be leveraged to improve bug reproduction. Our findings reveal several key insights that demonstrate the importance of images in supporting automated bug reproduction. Specifically, we identified six distinct functional roles that images serve in bug reports, each exhibiting unique patterns and specific contributions to the bug reproduction process. This study offers new insights into tool advancement and suggests promising directions for future research.

cs.SE