SearcharxivSearch

arXiv subjects

Chung-Yu Wang

Publications and source records attributed to Chung-Yu Wang.

7 recordsLinked to original sources

Deep-Bench: Deep Learning Benchmark Dataset for Code Generation

Deep learning (DL) has revolutionized areas such as computer vision, natural language processing, and more. However, developing DL systems is challenging due to the complexity of DL workflows. Large Language Models (LLMs), such as GPT, Claude, Llama, Mistral, etc., have emerged as promising tools to assist in DL code generation, offering potential solutions to these challenges. Despite this, existing benchmarks such as DS-1000 are limited, as they primarily focus on small DL code snippets related to pre/post-processing tasks and lack a comprehensive coverage of the full DL pipeline, including different DL phases and input data types. To address this, we introduce DeepBench, a novel benchmark dataset designed for function-level DL code generation. DeepBench categorizes DL problems based on three key aspects: phases such as pre-processing, model construction, and training; tasks, including classification, regression, and recommendation; and input data types such as tabular, image, and text. GPT-4o -- the state-of-the-art LLM -- achieved 31% accuracy on DeepBench, significantly lower than its 60% on DS-1000. We observed similar difficulty for other LLMs (e.g., 28% vs. 54% for Claude, 21% vs. 41% for LLaMA, and 15% vs. 20% for Mistral). This result underscores DeepBench's greater complexity. We also construct a taxonomy of issues and bugs found in LLM-generated DL code, which highlights the distinct challenges that LLMs face when generating DL code compared to general code. Furthermore, our analysis also reveals substantial performance variations across categories, with differences of up to 7% among phases and 37% among tasks. These disparities suggest that DeepBench offers valuable insights into the LLMs' performance and areas for potential improvement in the DL domain.

cs.SE

Ultrafast spintronics with geometric effects in non-adiabatic wave-packet dynamics

Motivated by the intriguing possibilities of steering ultrafast non-adiabatic processes through the geometric properties of bands in quantum materials by laser pulses, we extend a wave-packet transport theory, previously well-established in the adiabatic regime that intuitively captured geometric properties of bands, to the transient and non-adiabatic regime. This extension facilitates us to investigate macroscopic ways of manifesting microscopic band-geometric effects that highlight the special capability of non-adiabatic drivings not available to adiabatic drivings. These include imprinting band-geometric properties to the current rate after switching off the laser pulses and the induction of intrinsic macroscopic spin polarisation with an orientation not accessible by adiabatic processes. In particular, the microscopic geometrically-rooted intrinsic spin coherence is shown to underlie the spin-mediated parts of the macroscopic photocurrents. Through explicit calculations of an example with Rashba spin-orbit coupling, the spin-mediated part is shown to be discernible from the non-spin-mediated part in terms of the anisotropy of the photocurrents. Working principles behind the above theoretical results allegedly applicable beyond the Rashba example are distilled to inspect experimental data collected for SnSe, exhibiting considerable anisotropic effects. Consistency between theory and experiment is observed, paving the way of further exploration into the above intended direction.

cond-mat.mes-hall

Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity

Large Language Models (LLMs) have demonstrated impressive performance in software engineering tasks. However, improving their accuracy in generating correct and reliable code remains challenging. Numerous prompt engineering techniques (PETs) have been developed to address this, but no single approach is universally optimal. Selecting the right PET for each query is difficult for two primary reasons: (1) interactive prompting techniques may not consistently deliver the expected benefits, especially for simpler queries, and (2) current automated prompt engineering methods lack adaptability and fail to fully utilize multi-stage responses. To overcome these challenges, we propose PET-Select, a PET-agnostic selection model that uses code complexity as a proxy to classify queries and select the most appropriate PET. By incorporating contrastive learning, PET-Select effectively distinguishes between simple and complex problems, allowing it to choose PETs that are best suited for each query's complexity level. Our evaluations on the MBPP and HumanEval benchmarks using GPT-3.5 Turbo and GPT-4o show up to a 1.9% improvement in pass@1 accuracy, along with a 74.8% reduction in token usage. Additionally, we provide both quantitative and qualitative results to demonstrate how PET-Select effectively selects the most appropriate techniques for each code generation query, further showcasing its efficiency in optimizing PET selection.

cs.SE

Task-oriented Prompt Enhancement via Script Generation

Large Language Models (LLMs) have demonstrated remarkable abilities across various tasks, leveraging advanced reasoning. Yet, they struggle with task-oriented prompts due to a lack of specific prior knowledge of the task answers. The current state-of-the-art approach, PAL, utilizes code generation to address this issue. However, PAL depends on manually crafted prompt templates and examples while still producing inaccurate results. In this work, we present TITAN-a novel strategy designed to enhance LLMs' performance on task-oriented prompts. TITAN achieves this by generating scripts using a universal approach and zero-shot learning. Unlike existing methods, TITAN eliminates the need for detailed task-specific instructions and extensive manual efforts. TITAN enhances LLMs' performance on various tasks by utilizing their analytical and code-generation capabilities in a streamlined process. TITAN employs two key techniques: (1) step-back prompting to extract the task's input specifications and (2) chain-of-thought prompting to identify required procedural steps. This information is used to improve the LLMs' code-generation process. TITAN further refines the generated script through post-processing and the script is executed to retrieve the final answer. Our comprehensive evaluation demonstrates TITAN's effectiveness in a diverse set of tasks. On average, TITAN outperforms the state-of-the-art zero-shot approach by 7.6% and 3.9% when paired with GPT-3.5 and GPT-4. Overall, without human annotation, TITAN achieves state-of-the-art performance in 8 out of 11 cases while only marginally losing to few-shot approaches (which needed human intervention) on three occasions by small margins. This work represents a significant advancement in addressing task-oriented prompts, offering a novel solution for effectively utilizing LLMs in everyday life tasks.

cs.SE

Can ChatGPT Support Developers? An Empirical Evaluation of Large Language Models for Code Generation

Large language models (LLMs) have demonstrated notable proficiency in code generation, with numerous prior studies showing their promising capabilities in various development scenarios. However, these studies mainly provide evaluations in research settings, which leaves a significant gap in understanding how effectively LLMs can support developers in real-world. To address this, we conducted an empirical analysis of conversations in DevGPT, a dataset collected from developers' conversations with ChatGPT (captured with the Share Link feature on platforms such as GitHub). Our empirical findings indicate that the current practice of using LLM-generated code is typically limited to either demonstrating high-level concepts or providing examples in documentation, rather than to be used as production-ready code. These findings indicate that there is much future work needed to improve LLMs in code generation before they can be integral parts of modern software development.

cs.SE

Coupled Kohn-Sham equations for electrons and phonons

This work establishes the algebraic structure of the Kohn-Sham equations to be solved in a density formulation of electron and phonon dynamics, including the superconducting order parameter. A Bogoliubov transform is required to diagonalize both the fermionic and bosonic Kohn-Sham Hamiltonians since they both represent a non-interacting quantum field theory. The Bogoliubov transform for phonons is non-Hermitian in the general case, and the corresponding time-evolution is non-unitary. Several sufficient conditions for ensuring that the bosonic eigenvalues are real are provided and a practical method for solving the system is described. Finally, we produce a set of approximate mean-field potentials which are functionals of the electronic and phononic density matrices and depend on the electron-phonon vertex.

cond-mat.supr-con

Nonlinear Optical Properties of Transition Metal Dichalcogenide MX$_2$ (M = Mo, W; X = S, Se) Monolayers and Trilayers from First-principles Calculations

Due to the absence of interlayer coupling and inversion symmetry, transition metal dichalcogenide (MX$_2$) semiconductor monolayers exhibit novel properties that are distinctly different from their bulk crystals such as direct optical band gaps, large band spin splittings, spin-valley coupling, piezoelectric and nonlinear optical responses, and thus have promising applications in, e.g., opto-electronic and spintronic devices. Here we have performed a systematic first-principles study of the second-order nonlinear optical properties of MX$_2$ (M = Mo, W; X = S, Se) monolayers and trilayers within the density functional theory with the generalized gradient approximation plus scissors correction. We find that all the four MX$_2$ monolayers possess large second-order optical susceptibility $χ^{(2)}$ in the optical frequency range and significant linear electro-optical coefficients in low frequency limit, thus indicating their potential applications in non-linear optical devices and electric optical switches. The $χ^{(2)}$ spectra of the MX$_2$ trilayers are overall similar to the corresponding MX$_2$ monolayers, {\it albeit} with the magnitude reduced by roughly a factor of 3. The prominent features in the $χ^{(2)}$ spectra of the MX$_2$ multilayers are analyzed in terms of the underlying band structures and optical dielectric function, and also compared with available experiments.

cond-mat.mes-hall