SearcharxivSearch

arXiv subjects

Jing Chang

Publications and source records attributed to Jing Chang.

10 recordsLinked to original sources

ShapleyPipe: Hierarchical Shapley Search for Data Preparation Pipeline Construction

Automated data preparation pipeline construction is critical for machine learning success, yet existing methods suffer from two fundamental limitations: they treat pipeline construction as black-box optimization without quantifying individual operator contributions, and they struggle with the combinatorial explosion of the search space ($N^M$ configurations for N operators and pipeline length M). We introduce ShapleyPipe, a principled framework that leverages game-theoretic Shapley values to systematically quantify each operator's marginal contribution while maintaining full interpretability. Our key innovation is a hierarchical decomposition that separates category-level structure search from operator-level refinement, reducing the search complexity from exponential to polynomial. To make Shapley computation tractable, we develop: (1) a Multi-Armed Bandit mechanism for intelligent category evaluation with provable convergence guarantees, and (2) Permutation Shapley values to correctly capture position-dependent operator interactions. Extensive evaluation on 18 diverse datasets demonstrates that ShapleyPipe achieves 98.1\% of high-budget baseline performance while using 24\% fewer evaluations, and outperforms the state-of-the-art reinforcement learning method by 3.6\%. Beyond performance gains, ShapleyPipe provides interpretable operator valuations ($\rho$=0.933 correlation with empirical performance) that enable data-driven pipeline analysis and systematic operator library refinement.

cs.DB

SoftPipe: A Soft-Guided Reinforcement Learning Framework for Automated Data Preparation

Data preparation is a foundational yet notoriously challenging component of the machine learning lifecycle, characterized by a vast combinatorial search space. While reinforcement learning (RL) offers a promising direction, state-of-the-art methods suffer from a critical limitation: to manage the search space, they rely on rigid ``hard constraints'' that prematurely prune the search space and often preclude optimal solutions. To address this, we introduce SoftPipe, a novel RL framework that replaces these constraints with a flexible ``soft guidance'' paradigm. SoftPipe formulates action selection as a Bayesian inference problem. A high-level strategic prior, generated by a Large Language Model (LLM), probabilistically guides exploration. This prior is combined with empirical estimators from two sources through a collaborative process: a fine-grained quality score from a supervised Learning-to-Rank (LTR) model and a long-term value estimate from the agent's Q-function. Through extensive experiments on 18 diverse datasets, we demonstrate that SoftPipe achieves up to a 13.9\% improvement in pipeline quality and 2.8$\times$ faster convergence compared to existing methods.

cs.DB

LLaPipe: LLM-Guided Reinforcement Learning for Automated Data Preparation Pipeline Construction

Automated data preparation is crucial for democratizing machine learning, yet existing reinforcement learning (RL) based approaches suffer from inefficient exploration in the vast space of possible preprocessing pipelines. We present LLaPipe, a novel framework that addresses this exploration bottleneck by integrating Large Language Models (LLMs) as intelligent policy advisors. Unlike traditional methods that rely solely on statistical features and blind trial-and-error, LLaPipe leverages the semantic understanding capabilities of LLMs to provide contextually relevant exploration guidance. Our framework introduces three key innovations: (1) an LLM Policy Advisor that analyzes dataset semantics and pipeline history to suggest promising preprocessing operations, (2) an Experience Distillation mechanism that mines successful patterns from past pipelines and transfers this knowledge to guide future exploration, and (3) an Adaptive Advisor Triggering strategy (Advisor\textsuperscript{+}) that dynamically determines when LLM intervention is most beneficial, balancing exploration effectiveness with computational cost. Through extensive experiments on 18 diverse datasets spanning multiple domains, we demonstrate that LLaPipe achieves up to 22.4\% improvement in pipeline quality and 2.3$\times$ faster convergence compared to state-of-the-art RL-based methods, while maintaining computational efficiency through selective LLM usage (averaging only 19.0\% of total exploration steps).

cs.DB

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and differences present in multilingual and cross-lingual contexts. To address this gap, we introduce MaXIFE: a comprehensive evaluation benchmark designed to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. MaXIFE integrates both Rule-Based Evaluation and Model-Based Evaluation, ensuring a balance of efficiency and accuracy. We applied MaXIFE to evaluate several leading commercial LLMs, establishing baseline results for future comparisons. By providing a standardized tool for multilingual instruction-following evaluation, MaXIFE aims to advance research and development in natural language processing.

cs.CL

Yi: Open Foundation Models by 01.AI

We introduce the Yi model family, a series of language and multimodal models that demonstrate strong multi-dimensional capabilities. The Yi model family is based on 6B and 34B pretrained language models, then we extend them to chat models, 200K long context models, depth-upscaled models, and vision-language models. Our base models achieve strong performance on a wide range of benchmarks like MMLU, and our finetuned chat models deliver strong human preference rate on major evaluation platforms like AlpacaEval and Chatbot Arena. Building upon our scalable super-computing infrastructure and the classical transformer architecture, we attribute the performance of Yi models primarily to its data quality resulting from our data-engineering efforts. For pretraining, we construct 3.1 trillion tokens of English and Chinese corpora using a cascaded data deduplication and quality filtering pipeline. For finetuning, we polish a small scale (less than 10K) instruction dataset over multiple iterations such that every single instance has been verified directly by our machine learning engineers. For vision-language, we combine the chat language model with a vision transformer encoder and train the model to align visual representations to the semantic space of the language model. We further extend the context length to 200K through lightweight continual pretraining and demonstrate strong needle-in-a-haystack retrieval performance. We show that extending the depth of the pretrained checkpoint through continual pretraining further improves performance. We believe that given our current results, continuing to scale up model parameters using thoroughly optimized data will lead to even stronger frontier models.

cs.CL

Designing superhard magnetic material in clathrate \b{eta}-C3N2 through atom embeddedness

Designing new compounds with the coexistence of diverse physical properties is of great significance for broad applications in multifunctional electronic devices. In this work, based on density functional theory, we predict the coexistence of mechanical superhardness and the controllable magnetism in the clathrate material \b{eta}-C3N2 through the implant of the external atom into the intrinsic cage structure. Taking hydrogen-doping (H@\b{eta}-C3N2) and fluorine-doping (F@\b{eta}-C3N2) as examples, our calculations indicate these two doped configurations are stable and discovered that they belong to antiferromagnetic semiconductor and ferromagnetic semi-metal, respectively. These intriguing magnetic phase transitions originate from their distinctive band structure around the Fermi level and can be well understood by the 3D Hubbard model with half-filling occupation and the Stoner model. Moreover, the high Vickers hardness of 49.0 GPa for H@\b{eta}-C3N2 and 48.2 GPa for F@\b{eta}-C3N2 are obtained, suggesting they are clathrate superhard materials as its host. Therefore, the incorporation of H and F in \b{eta}-C3N2 gives rise to a new type of superhard antiferromagnetic semiconductor and superhard ferromagnetic semimetal, respectively, which could have potential applications in harsh conditions. Our work provides an effective strategy to design a new class of highly desirable multifunctional materials with excellent mechanical properties and magnetic properties, which may arouse spintronic applications in superhard materials in the future.

cond-mat.mtrl-sci

Evaluating the Effect of Crutch-using on Trunk Muscle Loads

As a traditional tool of external assistance, crutches play an important role in society. They have a wide range of applications to help either the elderly and disabled to walk or to treat certain illnesses or for post-operative rehabilitation. But there are many different types of crutches, including shoulder crutches and elbow crutches. How to choose has become an issue that deserves to be debated. Because while crutches help people walk, they also have an impact on the body. Inappropriate choice of crutches or long-term misuse can lead to problems such as scoliosis. Previous studies were mainly experimental measurements or the construction of dynamic models to calculate the load on joints with crutches. These studies focus only on the level of the joints, ignoring the role that muscles play in this process. Although some also take into account the degree of muscle activation, there is still a lack of quantitative analysis. The traditional dynamic model can be used to calculate the load on each joint. However, due to the activation of the muscle, this situation only causes part of the load transmitted to the joint, and the work of the chair will compensate the other part of the load. Analysis at the muscle level allows a better understanding of the impact of crutches on the body. By comparing the levels of activation of the trunk muscles, it was found that the use of crutches for walking, especially a single crutch, can cause a large difference in the activation of the back muscles on the left and right sides, and this difference will cause muscle degeneration for a long time, leading to scoliosis. In this article taking scoliosis as an example, by analyzing the muscles around the spine, we can better understand the pathology and can better prevent diseases. The objective of this article is to analyze normal walking compared to walking with one or two crutches using OpenSim software to obtain the degree of activation of different muscles in order to analyze the impact of crutches on the body.

cs.RO

Using 3D Scan to Determine Human Body Segment Mass in OpenSim Model

Biomechanical motion simulation and dynamic analysis of human joint moments will provide insights into Musculoskeletal Disorders. As one of the mainstream simulation tools, OpenSim uses proportional scaling to specify model segment masses to the simulated subject, which may bring about errors. This study aims at estimating the errors caused by the specifying method used in OpenSim as well as the influence of these errors on dynamic analysis. A 3D scan is used to construct subject's 3D geometric model, according to which segment masses are determined. The determined segment masses data is taken as the yardstick to assess the errors of OpenSim scaled model. Then influence of these errors on the dynamic calculation is evaluated in the simulation of a motion in which the subject walks in an ordinary gait. Result shows that the mass error in one segment can be as large as 5.31\% of overall body weight. The mean influence on calculated joint moment varies from 0.68\% to 12.68\% in 18 joints. In conclusion, a careful specification of segment masses will increase the accuracy of the dynamic simulation. As far as estimating human segment masses, the use of segment volume and density data can be an economical choice apart from referring to population mass distribution data.

physics.med-ph

Muscle Fatigue Analysis Using OpenSim

In this research, attempts are made to conduct concrete muscle fatigue analysis of arbitrary motions on OpenSim, a digital human modeling platform. A plug-in is written on the base of a muscle fatigue model, which makes it possible to calculate the decline of force-output capability of each muscle along time. The plug-in is tested on a three-dimensional, 29 degree-of-freedom human model. Motion data is obtained by motion capturing during an arbitrary running at a speed of 3.96 m/s. Ten muscles are selected for concrete analysis. As a result, the force-output capability of these muscles reduced to 60%-70% after 10 minutes' running, on a general basis. Erector spinae, which loses 39.2% of its maximal capability, is found to be more fatigue-exposed than the others. The influence of subject attributes (fatigability) is evaluated and discussed.

q-bio.TO

Estimating the EMG response exclusively to fatigue during sustained static maximum voluntary contraction

The increase of surface electromyography (sEMG) root-mean-square (RMS) is very frequently used to determine fatigue. However, as RMS is also influenced by muscle force,its effective usage as indicator of fatigue is mainly limited to isometric, constant force tasks.This research develops a simple me-thodto preclude the effect of muscle force, hereby estimates the EMG amplitude response exclusively to fatigue with RMS. Experiment was carried out on the biceps brachiis of 15 subjects (7males, 8 females) during sustained static maximum voluntary contractions (sMVC).Result shows that the sEMG RMS response to fatigue increasesto 21.27% while muscle force decreasing to 50%MVC, which implies that more and more extra effort is needed as muscle fatigue intensifies. It would be promising to use the RMS response exclusively to fatigue as an indicator of muscle fatigue.

q-bio.TO