SearcharxivSearch

arXiv subjects

Guoxiang Guo

Publications and source records attributed to Guoxiang Guo.

3 recordsLinked to original sources

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues

Large Language Models (LLMs) are increasingly used in software engineering to generate and refine code. In practice, developers often continue from an initial code generation request with follow-up refinement instructions, such as requests to improve style, restructure implementation, or change the execution strategy while preserving the intended behaviour. However, existing benchmarks generally omit this multi-turn code refinement dialogue setting and therefore cannot evaluate whether LLMs maintain functional correctness, i.e., whether the refined code still passes the test suite for the original task. To address this limitation, we introduce CodeChat-Eval, an evaluation framework that constructs evaluation sessions from multi-turn code refinement dialogues using a dynamic instruction selection algorithm. Our empirical study on open-weight and proprietary LLMs observes a statistically significant decrease ranging from 19.2% (GPT-5 Nano) to 69.2% (Llama 3.1 8B) in functional correctness over multi-turn refinement. The largest correctness drops are associated with logic-level refinements and additive change requests. These findings indicate that LLMs struggle to maintain functional correctness during multi-turn code refinement dialogues, and highlight the need for benchmarks that evaluate functionality-preserving refinement beyond single-turn generation.

cs.SE

Next Generation Active Learning: Mixture of LLMs in the Loop

With the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering the annotation quality, labels generated by LLMs often fall short of real-world applicability. To address this, we propose a novel active learning framework, Mixture of LLMs in the Loop Active Learning, replacing human annotators with labels generated through a Mixture-of-LLMs-based annotation model, aimed at enhancing LLM-based annotation robustness by aggregating the strengths of multiple LLMs. To further mitigate the impact of the noisy labels, we introduce annotation discrepancy and negative learning to identify the unreliable annotations and enhance learning effectiveness. Extensive experiments demonstrate that our framework achieves performance comparable to human annotation and consistently outperforms single-LLM baselines and other LLM-ensemble-based approaches. Moreover, our framework is built on lightweight LLMs, enabling it to operate fully on local machines in real-world applications.

cs.LG

Stable gain-switched thulium fiber laser with 140 nm tuning range

We demonstrate a gain-switched thulium fiber laser that can be continuously tuned over 140 nm, while maintaining stable nanosecond single-pulse operation. To the best of our knowledge, this system represents the broadest tuning range for a gain-switched fiber laser. The system simplicity and wideband wavelength tunability combined with the ability to control the temporal characteristics of the gain-switched pulses mean this is a versatile source highly suited to a wide range of applications in the eye-safe region of the infrared, including spectroscopy, sensing and material processing, as well as being a practical seed source for pumping nonlinear processes.

physics.optics