SearcharxivSearch

arXiv subjects

Lu Han

Publications and source records attributed to Lu Han.

At least 37 records · Page 2Linked to original sources

Molly: Making Large Language Model Agents Solve Python Problem More Logically

Applying large language models (LLMs) as teaching assists has attracted much attention as an integral part of intelligent education, particularly in computing courses. To reduce the gap between the LLMs and the computer programming education expert, fine-tuning and retrieval augmented generation (RAG) are the two mainstream methods in existing researches. However, fine-tuning for specific tasks is resource-intensive and may diminish the model`s generalization capabilities. RAG can perform well on reducing the illusion of LLMs, but the generation of irrelevant factual content during reasoning can cause significant confusion for learners. To address these problems, we introduce the Molly agent, focusing on solving the proposed problem encountered by learners when learning Python programming language. Our agent automatically parse the learners' questioning intent through a scenario-based interaction, enabling precise retrieval of relevant documents from the constructed knowledge base. At generation stage, the agent reflect on the generated responses to ensure that they not only align with factual content but also effectively answer the user's queries. Extensive experimentation on a constructed Chinese Python QA dataset shows the effectiveness of the Molly agent, indicating an enhancement in its performance for providing useful responses to Python questions.

cs.CL

MIETT: Multi-Instance Encrypted Traffic Transformer for Encrypted Traffic Classification

Network traffic includes data transmitted across a network, such as web browsing and file transfers, and is organized into packets (small units of data) and flows (sequences of packets exchanged between two endpoints). Classifying encrypted traffic is essential for detecting security threats and optimizing network management. Recent advancements have highlighted the superiority of foundation models in this task, particularly for their ability to leverage large amounts of unlabeled data and demonstrate strong generalization to unseen data. However, existing methods that focus on token-level relationships fail to capture broader flow patterns, as tokens, defined as sequences of hexadecimal digits, typically carry limited semantic information in encrypted traffic. These flow patterns, which are crucial for traffic classification, arise from the interactions between packets within a flow, not just their internal structure. To address this limitation, we propose a Multi-Instance Encrypted Traffic Transformer (MIETT), which adopts a multi-instance approach where each packet is treated as a distinct instance within a larger bag representing the entire flow. This enables the model to capture both token-level and packet-level relationships more effectively through Two-Level Attention (TLA) layers, improving the model's ability to learn complex packet dynamics and flow patterns. We further enhance the model's understanding of temporal and flow-specific dynamics by introducing two novel pre-training tasks: Packet Relative Position Prediction (PRPP) and Flow Contrastive Learning (FCL). After fine-tuning, MIETT achieves state-of-the-art (SOTA) results across five datasets, demonstrating its effectiveness in classifying encrypted traffic and understanding complex network behaviors. Code is available at \url{https://github.com/Secilia-Cxy/MIETT}.

cs.CR

SOFTS: Efficient Multivariate Time Series Forecasting with Series-Core Fusion

Multivariate time series forecasting plays a crucial role in various fields such as finance, traffic management, energy, and healthcare. Recent studies have highlighted the advantages of channel independence to resist distribution drift but neglect channel correlations, limiting further enhancements. Several methods utilize mechanisms like attention or mixer to address this by capturing channel correlations, but they either introduce excessive complexity or rely too heavily on the correlation to achieve satisfactory results under distribution drifts, particularly with a large number of channels. Addressing this gap, this paper presents an efficient MLP-based model, the Series-cOre Fused Time Series forecaster (SOFTS), which incorporates a novel STar Aggregate-Redistribute (STAR) module. Unlike traditional approaches that manage channel interactions through distributed structures, \textit{e.g.}, attention, STAR employs a centralized strategy to improve efficiency and reduce reliance on the quality of each channel. It aggregates all series to form a global core representation, which is then dispatched and fused with individual series representations to facilitate channel interactions effectively.SOFTS achieves superior performance over existing state-of-the-art methods with only linear complexity. The broad applicability of the STAR module across different forecasting models is also demonstrated empirically. For further research and development, we have made our code publicly available at https://github.com/Secilia-Cxy/SOFTS.

cs.LG

Sharingan: Extract User Action Sequence from Desktop Recordings

Video recordings of user activities, particularly desktop recordings, offer a rich source of data for understanding user behaviors and automating processes. However, despite advancements in Vision-Language Models (VLMs) and their increasing use in video analysis, extracting user actions from desktop recordings remains an underexplored area. This paper addresses this gap by proposing two novel VLM-based methods for user action extraction: the Direct Frame-Based Approach (DF), which inputs sampled frames directly into VLMs, and the Differential Frame-Based Approach (DiffF), which incorporates explicit frame differences detected via computer vision techniques. We evaluate these methods using a basic self-curated dataset and an advanced benchmark adapted from prior work. Our results show that the DF approach achieves an accuracy of 70% to 80% in identifying user actions, with the extracted action sequences being re-playable though Robotic Process Automation. We find that while VLMs show potential, incorporating explicit UI changes can degrade performance, making the DF approach more reliable. This work represents the first application of VLMs for extracting user action sequences from desktop recordings, contributing new methods, benchmarks, and insights for future research.

cs.CV

Bulk Crystal Growth and Single-Crystal-to-Single-Crystal Phase Transitions in the Averievite CsClCu5V2O10

Quasi-two-dimensional averievites with triangle-kagome-triangle trilayers are of interest due to their rich structural and magnetic transitions and strong spin frustration that are expected to host quantum spin liquid ground state with suitable substitution or doping. Herein, we report growth of bulk single crystals of averievite CsClCu5V2O10 with dimensions of several millimeters on edge in order to (1) address the open question whether the room temperature crystal structure is P-3m1, P-3, P21/c or else, (2) to elucidate the nature of phase transitions, and (3) to study direction-dependent physical properties. Single-crystal-to-single-crystal structural transitions at ~305 K and ~127 K were observed in the averievite CsClCu5V2O10 single crystals. The nature of the transition at ~305 K, which was reported as P-3m1-P21/c transition, was found to be a structural transition from high temperature P-3m1 to low temperature P-3 by combining variable temperature synchrotron X-ray single crystal and high-resolution powder diffraction. In-plane and out-of-plane magnetic susceptibility and heat capacity measurements confirm a first-order transition at 305 K, a structural transition at 127 K and an antiferromagnetic transition at 24 K. These averievites are thus ideal model systems for a deeper understanding of structural transitions and magnetism.

cond-mat.str-el

Cascade of phase transitions and large magnetic anisotropy in a triangle-kagome-triangle trilayer antiferromagnet

Spins in strongly frustrated systems are of intense interest due to the emergence of intriguing quantum states including superconductivity and quantum spin liquid. Herein we report the discovery of cascade of phase transitions and large magnetic anisotropy in the averievite CsClCu5P2O10 single crystals. Under zero field, CsClCu5P2O10 undergoes a first-order structural transition at around 225 K from high temperature centrosymmetric P-3m1 to low temperature noncentrosymmetric P321, followed by an AFM transition at 13.6 K, another structural transition centering at ~3 K, and another AFM transition at ~2.18 K. Based upon magnetic susceptibility and magnetization data with magnetic fields perpendicular to the ab plane, a phase diagram, consisting of a paramagnetic state, two AFM states and four field-induced states including two magnetization plateaus, has been constructed. Our findings demonstrate that the quasi-2D CsClCu5P2O10 exhibits rich structural and metamagnetic transitions and the averievite family is a fertile platform for exploring novel quantum states.

cond-mat.str-el

Low-temperature aqueous solution growth of the acousto-optic TeO2 single crystals

$α$-TeO2 is widely used in acousto-optic devices due to its excellent physical properties. Conventionally, $α$-TeO2 single crystals were grown using melt methods. Here, we report for the first time the growth of $α$-TeO2 single crystals using the aqueous solution method below 100 °C. Solubility curve of $α$-TeO2 was measured, and then single crystals with dimensions of 3.5x3.5x2.5 mm3 were successfully grown using seed crystals that were synthesized from spontaneous nucleation. The as-grown single crystals belong to the P41212 space group, evidenced by single crystal X-ray diffraction and Rietveld refinement on powder diffraction. Rocking curve measurements show that the as-grown crystals exhibit high crystallinity with a full-width at half maxima (FWHM) of 57.2''. Ultraviolet-Visible absorption spectroscopy indicates the absorption edge is 350 nm and the band gap is estimated to be 3.58 eV. The density and Vickers hardness of as-grown single crystals are measured to be 6.042 g/cm3 and 404 kg/mm2, repectively. Our findings provide an easy-to-access and energy-saving method for growing single crystals of inorganic compounds.

cond-mat.mtrl-sci

QACP: An Annotated Question Answering Dataset for Assisting Chinese Python Programming Learners

In online learning platforms, particularly in rapidly growing computer programming courses, addressing the thousands of students' learning queries requires considerable human cost. The creation of intelligent assistant large language models (LLMs) tailored for programming education necessitates distinct data support. However, in real application scenarios, the data resources for training such LLMs are relatively scarce. Therefore, to address the data scarcity in intelligent educational systems for programming, this paper proposes a new Chinese question-and-answer dataset for Python learners. To ensure the authenticity and reliability of the sources of the questions, we collected questions from actual student questions and categorized them according to various dimensions such as the type of questions and the type of learners. This annotation principle is designed to enhance the effectiveness and quality of online programming education, providing a solid data foundation for developing the programming teaching assists (TA). Furthermore, we conducted comprehensive evaluations of various LLMs proficient in processing and generating Chinese content, highlighting the potential limitations of general LLMs as intelligent teaching assistants in computer programming courses.

cs.CL

Inelastic electron scattering at large angles: the phonon polariton contribution

We explore the inelastic electron scattering in SrTiO3, PbTiO3, and SiC in their phonon energy range, challenging the assumption that phonon polaritons are excluded at large angles in high-resolution transmission electron energy-loss spectroscopy. We demonstrate that through multiple scattering, the electron beam can excite both phonons and phonon polaritons, and the relative proportion of each varies depending on the structure factor and scattering angle. Integrating dielectric theory, density functional theory, and multi-slice simulations, we provide a comprehensive framework for understanding these interactions in materials with polar optical phonons.

cond-mat.mtrl-sci

Twice Class Bias Correction for Imbalanced Semi-Supervised Learning

Differing from traditional semi-supervised learning, class-imbalanced semi-supervised learning presents two distinct challenges: (1) The imbalanced distribution of training samples leads to model bias towards certain classes, and (2) the distribution of unlabeled samples is unknown and potentially distinct from that of labeled samples, which further contributes to class bias in the pseudo-labels during training. To address these dual challenges, we introduce a novel approach called \textbf{T}wice \textbf{C}lass \textbf{B}ias \textbf{C}orrection (\textbf{TCBC}). We begin by utilizing an estimate of the class distribution from the participating training samples to correct the model, enabling it to learn the posterior probabilities of samples under a class-balanced prior. This correction serves to alleviate the inherent class bias of the model. Building upon this foundation, we further estimate the class bias of the current model parameters during the training process. We apply a secondary correction to the model's pseudo-labels for unlabeled samples, aiming to make the assignment of pseudo-labels across different classes of unlabeled samples as equitable as possible. Through extensive experimentation on CIFAR10/100-LT, STL10-LT, and the sizable long-tailed dataset SUN397, we provide conclusive evidence that our proposed TCBC method reliably enhances the performance of class-imbalanced semi-supervised learning.

cs.LG

Learning Robust Precipitation Forecaster by Temporal Frame Interpolation

Recent advances in deep learning have significantly elevated weather prediction models. However, these models often falter in real-world scenarios due to their sensitivity to spatial-temporal shifts. This issue is particularly acute in weather forecasting, where models are prone to overfit to local and temporal variations, especially when tasked with fine-grained predictions. In this paper, we address these challenges by developing a robust precipitation forecasting model that demonstrates resilience against such spatial-temporal discrepancies. We introduce Temporal Frame Interpolation (TFI), a novel technique that enhances the training dataset by generating synthetic samples through interpolating adjacent frames from satellite imagery and ground radar data, thus improving the model's robustness against frame noise. Moreover, we incorporate a unique Multi-Level Dice (ML-Dice) loss function, leveraging the ordinal nature of rainfall intensities to improve the model's performance. Our approach has led to significant improvements in forecasting precision, culminating in our model securing \textit{1st place} in the transfer learning leaderboard of the \textit{Weather4cast'23} competition. This achievement not only underscores the effectiveness of our methodologies but also establishes a new standard for deep learning applications in weather forecasting. Our code and weights have been public on \url{https://github.com/Secilia-Cxy/UNetTFI}.

cs.LG

Single Diamond Structured Titania Scaffold

The single diamond (SD) network, discovered in beetle and weevil skeletons, is the 'holy grail' of photonic materials with the widest complete bandgap known to date. However, the thermodynamic instability of SD has made its self-assembly long been a formidable challenge. By imitating the simultaneous co-folding process of nonequilibrium skeleton formation in natural organisms, we devised an unprecedented bottom-up approach to fabricate SD networks via the synergistic self-assembly of diblock copolymer and inorganic precursors and successfully obtained tetrahedral connected polycrystalline anatase SD frameworks. A photonic bandstructure calculation showed that the resulting SD structure has a wide and complete photonic bandgap. This work provides an ingenious design solution to the complex synthetic puzzle and offers new opportunities for biorelevant materials, next-generation optical devices, etc.

physics.app-ph

The Capacity and Robustness Trade-off: Revisiting the Channel Independent Strategy for Multivariate Time Series Forecasting

Multivariate time series data comprises various channels of variables. The multivariate forecasting models need to capture the relationship between the channels to accurately predict future values. However, recently, there has been an emergence of methods that employ the Channel Independent (CI) strategy. These methods view multivariate time series data as separate univariate time series and disregard the correlation between channels. Surprisingly, our empirical results have shown that models trained with the CI strategy outperform those trained with the Channel Dependent (CD) strategy, usually by a significant margin. Nevertheless, the reasons behind this phenomenon have not yet been thoroughly explored in the literature. This paper provides comprehensive empirical and theoretical analyses of the characteristics of multivariate time series datasets and the CI/CD strategy. Our results conclude that the CD approach has higher capacity but often lacks robustness to accurately predict distributionally drifted time series. In contrast, the CI approach trades capacity for robust prediction. Practical measures inspired by these analyses are proposed to address the capacity and robustness dilemma, including a modified CD method called Predict Residuals with Regularization (PRReg) that can surpass the CI strategy. We hope our findings can raise awareness among researchers about the characteristics of multivariate time series and inspire the construction of better forecasting models.

cs.LG

Augmentation Component Analysis: Modeling Similarity via the Augmentation Overlaps

Self-supervised learning aims to learn a embedding space where semantically similar samples are close. Contrastive learning methods pull views of samples together and push different samples away, which utilizes semantic invariance of augmentation but ignores the relationship between samples. To better exploit the power of augmentation, we observe that semantically similar samples are more likely to have similar augmented views. Therefore, we can take the augmented views as a special description of a sample. In this paper, we model such a description as the augmentation distribution and we call it augmentation feature. The similarity in augmentation feature reflects how much the views of two samples overlap and is related to their semantical similarity. Without computational burdens to explicitly estimate values of the augmentation feature, we propose Augmentation Component Analysis (ACA) with a contrastive-like loss to learn principal components and an on-the-fly projection loss to embed data. ACA equals an efficient dimension reduction by PCA and extracts low-dimensional embeddings, theoretically preserving the similarity of augmentation distribution between samples. Empirical results show our method can achieve competitive results against various traditional contrastive learning methods on different benchmarks.

cs.LG

On Pseudo-Labeling for Class-Mismatch Semi-Supervised Learning

When there are unlabeled Out-Of-Distribution (OOD) data from other classes, Semi-Supervised Learning (SSL) methods suffer from severe performance degradation and even get worse than merely training on labeled data. In this paper, we empirically analyze Pseudo-Labeling (PL) in class-mismatched SSL. PL is a simple and representative SSL method that transforms SSL problems into supervised learning by creating pseudo-labels for unlabeled data according to the model's prediction. We aim to answer two main questions: (1) How do OOD data influence PL? (2) What is the proper usage of OOD data with PL? First, we show that the major problem of PL is imbalanced pseudo-labels on OOD data. Second, we find that OOD data can help classify In-Distribution (ID) data given their OOD ground truth labels. Based on the findings, we propose to improve PL in class-mismatched SSL with two components -- Re-balanced Pseudo-Labeling (RPL) and Semantic Exploration Clustering (SEC). RPL re-balances pseudo-labels of high-confidence data, which simultaneously filters out OOD data and addresses the imbalance problem. SEC uses balanced clustering on low-confidence data to create pseudo-labels on extra classes, simulating the process of training with ground truth. Experiments show that our method achieves steady improvement over supervised baseline and state-of-the-art performance under all class mismatch ratios on different benchmarks.

cs.LG

Revisiting Unsupervised Meta-Learning via the Characteristics of Few-Shot Tasks

Meta-learning has become a practical approach towards few-shot image classification, where "a strategy to learn a classifier" is meta-learned on labeled base classes and can be applied to tasks with novel classes. We remove the requirement of base class labels and learn generalizable embeddings via Unsupervised Meta-Learning (UML). Specifically, episodes of tasks are constructed with data augmentations from unlabeled base classes during meta-training, and we apply embedding-based classifiers to novel tasks with labeled few-shot examples during meta-test. We observe two elements play important roles in UML, i.e., the way to sample tasks and measure similarities between instances. Thus we obtain a strong baseline with two simple modifications -- a sufficient sampling strategy constructing multiple tasks per episode efficiently together with a semi-normalized similarity. We then take advantage of the characteristics of tasks from two directions to get further improvements. First, synthesized confusing instances are incorporated to help extract more discriminative embeddings. Second, we utilize an additional task-specific embedding transformation as an auxiliary component during meta-training to promote the generalization ability of the pre-adapted embeddings. Experiments on few-shot learning benchmarks verify that our approaches outperform previous UML methods and achieve comparable or even better performance than its supervised variants.

cs.CV

Abnormal Signal Recognition with Time-Frequency Spectrogram: A Deep Learning Approach

With the increasingly complex and changeable electromagnetic environment, wireless communication systems are facing jamming and abnormal signal injection, which significantly affects the normal operation of a communication system. In particular, the abnormal signals may emulate the normal signals, which makes it very challenging for abnormal signal recognition. In this paper, we propose a new abnormal signal recognition scheme, which combines time-frequency analysis with deep learning to effectively identify synthetic abnormal communication signals. Firstly, we emulate synthetic abnormal communication signals including seven jamming patterns. Then, we model an abnormal communication signals recognition system based on the communication protocol between the transmitter and the receiver. To improve the performance, we convert the original signal into the time-frequency spectrogram to develop an image classification algorithm. Simulation results demonstrate that the proposed method can effectively recognize the abnormal signals under various parameter configurations, even under low signal-to-noise ratio (SNR) and low jamming-to-signal ratio (JSR) conditions.

eess.SP

Electron transfer under the Floquet modulation in donor-bridge-acceptor systems

Electron transfer (ET) processes are of broad interest in modern chemistry. With the advancements of experimental techniques, one may modulate the ET via such as the light-matter interactions. In this work, we study the ET under a Floquet modulation occurring in the donor-bridge-acceptor systems, with the rate kernels projected out from the exact disspaton equation of motion formalism. This together with the Floquet theorem enables us to investigate the interplay between the intrinsic non-Markovianity and the driving periodicity. The observed rate kernel exhibits a Herzberg-Teller-like mechanism induced by the bridge fluctuation subject to effective modulation.

physics.chem-ph