SearcharxivSearch

arXiv subjects

Yuting Huang

Publications and source records attributed to Yuting Huang.

18 recordsLinked to original sources

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present $\Phi$-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, $\Phi$-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.

cs.CL

LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse labware and instruments and execute long-horizon, state-dependent experimental procedures. Yet existing benchmarks do not jointly capture dexterous hand use, real-world laboratory interactions, and multi-stage experimental procedures, limiting systematic training and evaluation. To bridge this gap, we introduce LabDex, a large-scale real-world dataset and benchmark for dexterous manipulation in chemistry laboratories, organized around a hierarchical task taxonomy spanning atomic skills, compositional tasks, and long-horizon experiments. First, LabDex is cross-platform and, for the first time, unifies real-world and simulation platforms under a common framework, providing standardized task definitions, demonstrations, and evaluation protocols. Second, LabDex is large-scale and systematically organizes chemistry laboratory operations into three interconnected levels: Atomic Skills, which characterize fundamental dexterous manipulation capabilities; Compositional Skills; and Long-Horizon Laboratory Workflows. This hierarchical design not only supports the evaluation of end-task performance, but also enables the analysis of how fundamental dexterous skills compose and influence more complex laboratory operations. We conduct cross-level evaluations of representative robot learning methods in both real-world and simulation environments. The experimental results validate the effectiveness of the LabDex task design and demonstration data, and show that the benchmark supports the training and systematic evaluation of existing robotic policies across laboratory dexterous manipulation tasks at different levels, providing a foundation for further research and development of autonomous laboratory robots.

cs.RO

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs

Large language models (LLMs) have achieved remarkable success in general-domain tasks, yet their direct application to the legal domain remains challenging due to hallucinated legal citations, incomplete knowledge coverage, and weak structured reasoning. To address these issues, we propose PoliLegalLM, a domain-specific large language model tailored for political and legal applications. Our approach adopts a unified training framework that integrates continued pretraining, progressive supervised fine-tuning, and preference-based reinforcement learning to jointly enhance legal knowledge grounding, task alignment, and reasoning capability. We construct a large-scale, high-quality legal corpus and design a structured post-training pipeline, enabling the model to effectively learn domain-specific knowledge and adapt to diverse legal tasks. We evaluate PoliLegalLM on three representative benchmarks, including LawBench, LexEval, and a real-world dataset, PoliLegal. Experimental results demonstrate that PoliLegalLM achieves strong and consistent performance, outperforming competitive models of similar scale and remaining highly competitive with significantly larger models, while achieving the best results on real-world legal scenarios. These results highlight the effectiveness of our training paradigm and the practical value of domain-specific LLMs for real-world legal applications.

cs.CL

Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity patterns change accordingly. Static pruning lacks the adaptability required for environment dynamics, whereas fixed-interval dynamic layer pruning suffers from coarse granularity and high retraining overheads. To bridge this gap, we propose EcoVLA, a training-free, plug-and-play adaptive pruning framework that supports orthogonal combination with existing VLA acceleration methods. EcoVLA comprises two components: Environment-aware Adaptive Pruning (EAP) and Interleaved Inference Orchestration ($I^2O$). EAP is a lightweight adaptive channel pruning method that incorporates the temporal consistency of the physical environment to update sparsity patterns. $I^2O$ leverages the FLOPs bubbles inherent in VLA inference to schedule the pruning method in parallel, ensuring negligible impact on latency. Evaluated on diverse VLA models and benchmarks, EcoVLA delivers state-of-the-art performance, achieving up to 1.60$\times$ speedup with only a 0.4% drop in success rate, and further reaches 2.18$\times$ speedup with only a 0.5% degradation when combined with token pruning. We further validate the effectiveness of EcoVLA on real-world robots.

cs.AI

Building AI Agents to Improve Job Referral Requests to Strangers

This paper develops AI agents that help job seekers write effective requests for job referrals in a professional online community. The basic workflow consists of an improver agent that rewrites the referral request and an evaluator agent that measures the quality of revisions using a model trained to predict the probability of receiving referrals from other users. Revisions suggested by the LLM (large language model) increase predicted success rates for weaker requests while reducing them for stronger requests. Enhancing the LLM with Retrieval-Augmented Generation (RAG) prevents edits that worsen stronger requests while it amplifies improvements for weaker requests. Overall, using LLM revisions with RAG increases the predicted success rate for weaker requests by 14\% without degrading performance on stronger requests. Although improvements in model-predicted success do not guarantee more referrals in the real world, they provide low-cost signals for promising features before running higher-stakes experiments on real users.

cs.AI

An Improved Inverse Method for Estimating Disease Transmission Rates in Low-Prevalence Epidemics

The accurate estimation of time-varying transmission rates is fundamental for understanding infectious disease dynamics and implementing effective public health interventions. To this end, we propose an improved inverse method for estimating time-varying transmission rates in low-prevalence settings, where conventional data preprocessing approaches often fail due to sparse case observations. To overcome this difficulty, we introduce an exponential B-spline interpolation approach that integrates both continuous and discrete inverse methods. This method ensures that transmission rate estimates remain non-negative and smooth, even when the observed data exhibit low cases. We apply this approach to several infectious disease models using real-world data from China, including a scarlet fever model, a multi-strain influenza model, and an age-structured influenza model. The results show that our method provides accurate transmission rate estimates, particularly in low-prevalence infectious diseases and multi-group epidemic models, demonstrating its robustness and applicability across various epidemiological contexts. The improved inverse method offers a new perspective for epidemiological modeling and provides reliable technical support for related theoretical exploration and public health decision-making.

q-bio.PE

Mechanical performance of hybrid polymer-lipid vesicles with leaflet asymmetry engineered using microfluidics

Lipid vesicles consist of aqueous cores surrounded by a bilayer of phospholipids. Hybrid polymer-lipid vesicles incorporate both polymers and lipids, offering promising properties for developing pharmaceuticals, biosensors, and artificial cells. The hybrid vesicles can be symmetric, in which their two leaflets contain identical compositions, or asymmetric, in which the leaflets possess dissimilar compositions and can lead to dramatically modified properties. However, methods to produce both symmetric and asymmetric hybrid vesicles result in heterogenous compositions and sizes, making it challenging to quantify the effect of asymmetry and limiting applications. Here, we use a microfluidic approach to produce hybrid vesicles containing symmetric or asymmetric leaflets with precisely engineered compositions. We find the vesicles with asymmetric leaflets are significantly stiffer and tougher than those with symmetric leaflets; moreover, the lateral diffusivity of lipids is greatly decreased. The structure for improved toughness consists of an inner leaflet that is a stretchable lipid leaflet and an outer leaflet that is a fully continuous polymer leaflet. This technique of precisely engineering asymmetric structures may be applied to hybrid vesicles composed of block copolymers and phospholipids dissolvable in chloroform and hexane, further expanding their applications.

cond-mat.soft

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

Modeling temporal multimodal data poses significant challenges in classification tasks, particularly in capturing long-range temporal dependencies and intricate cross-modal interactions. Audiovisual data, as a representative example, is inherently characterized by strict temporal order and diverse modalities. Effectively leveraging the temporal structure is essential for understanding both intra-modal dynamics and inter-modal correlations. However, most existing approaches treat each modality independently and rely on shallow fusion strategies, which overlook temporal dependencies and hinder the model's ability to represent complex structural relationships. To address the limitation, we propose the hybrid hypergraph network (HHN), a novel framework that models temporal multimodal data via a segmentation-first, graph-later strategy. HHN splits sequences into timestamped segments as nodes in a heterogeneous graph. Intra-modal structures are captured via hyperedges guided by a maximum entropy difference criterion, enhancing node heterogeneity and structural discrimination, followed by hypergraph convolution to extract high-order dependencies. Inter-modal links are established through temporal alignment and graph attention for semantic fusion. HHN achieves state-of-the-art (SOTA) results on four multimodal datasets, demonstrating its effectiveness in complex classification tasks.

cs.LG

VLMPlanner: Integrating Visual Language Models with Motion Planning

Integrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controllability, and improved generalization in rare and long-tail scenarios. However, existing methods often rely on abstracted perception or map-based inputs, missing crucial visual context, such as fine-grained road cues, accident aftermath, or unexpected obstacles, which are essential for robust decision-making in complex driving environments. To bridge this gap, we propose VLMPlanner, a hybrid framework that combines a learning-based real-time planner with a vision-language model (VLM) capable of reasoning over raw images. The VLM processes multi-view images to capture rich, detailed visual information and leverages its common-sense reasoning capabilities to guide the real-time planner in generating robust and safe trajectories. Furthermore, we develop the Context-Adaptive Inference Gate (CAI-Gate) mechanism that enables the VLM to mimic human driving behavior by dynamically adjusting its inference frequency based on scene complexity, thereby achieving an optimal balance between planning performance and computational efficiency. We evaluate our approach on the large-scale, challenging nuPlan benchmark, with comprehensive experimental results demonstrating superior planning performance in scenarios with intricate road conditions and dynamic elements. Code will be available.

cs.AI

Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach

Spatio-temporal prediction plays a crucial role in intelligent transportation, weather forecasting, and urban planning. While integrating multi-modal data has shown potential for enhancing prediction accuracy, key challenges persist: (i) inadequate fusion of multi-modal information, (ii) confounding factors that obscure causal relations, and (iii) high computational complexity of prediction models. To address these challenges, we propose E^2-CSTP, an Effective and Efficient Causal multi-modal Spatio-Temporal Prediction framework. E^2-CSTP leverages cross-modal attention and gating mechanisms to effectively integrate multi-modal data. Building on this, we design a dual-branch causal inference approach: the primary branch focuses on spatio-temporal prediction, while the auxiliary branch mitigates bias by modeling additional modalities and applying causal interventions to uncover true causal dependencies. To improve model efficiency, we integrate GCN with the Mamba architecture for accelerated spatio-temporal encoding. Extensive experiments on 4 real-world datasets show that E^2-CSTP significantly outperforms 9 state-of-the-art methods, achieving up to 9.66% improvements in accuracy as well as 17.37%-56.11% reductions in computational overhead.

cs.LG

AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios

Recent advances in LegalAI have primarily focused on individual case judgment analysis, often overlooking the critical appellate process within the judicial system. Appeals serve as a core mechanism for error correction and ensuring fair trials, making them highly significant both in practice and in research. To address this gap, we present the AppealCase dataset, consisting of 10,000 pairs of real-world, matched first-instance and second-instance documents across 91 categories of civil cases. The dataset also includes detailed annotations along five dimensions central to appellate review: judgment reversals, reversal reasons, cited legal provisions, claim-level decisions, and whether there is new information in the second instance. Based on these annotations, we propose five novel LegalAI tasks and conduct a comprehensive evaluation across 20 mainstream models. Experimental results reveal that all current models achieve less than 50% F1 scores on the judgment reversal prediction task, highlighting the complexity and challenge of the appeal scenario. We hope that the AppealCase dataset will spur further research in LegalAI for appellate case analysis and contribute to improving consistency in judicial decision-making.

cs.CL

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents

Large Language Models (LLMs) exhibit substantial promise in enhancing task-planning capabilities within embodied agents due to their advanced reasoning and comprehension. However, the systemic safety of these agents remains an underexplored frontier. In this study, we present Safe-BeAl, an integrated framework for the measurement (SafePlan-Bench) and alignment (Safe-Align) of LLM-based embodied agents' behaviors. SafePlan-Bench establishes a comprehensive benchmark for evaluating task-planning safety, encompassing 2,027 daily tasks and corresponding environments distributed across 8 distinct hazard categories (e.g., Fire Hazard). Our empirical analysis reveals that even in the absence of adversarial inputs or malicious intent, LLM-based agents can exhibit unsafe behaviors. To mitigate these hazards, we propose Safe-Align, a method designed to integrate physical-world safety knowledge into LLM-based embodied agents while maintaining task-specific performance. Experiments across a variety of settings demonstrate that Safe-BeAl provides comprehensive safety validation, improving safety by 8.55 - 15.22%, compared to embodied agents based on GPT-4, while ensuring successful task completion.

cs.AI

Effective and Efficient Cross-City Traffic Knowledge Transfer: A Privacy-Preserving Perspective

Traffic prediction aims to forecast future traffic conditions using historical traffic data, serving a crucial role in urban computing and transportation management. While transfer learning and federated learning have been employed to address the scarcity of traffic data by transferring traffic knowledge from data-rich to data-scarce cities without traffic data exchange, existing approaches in Federated Traffic Knowledge Transfer (FTT) still face several critical challenges such as potential privacy leakage, cross-city data distribution discrepancies, and low data quality, hindering their practical application in real-world scenarios. To this end, we present FedTT, a novel privacy-aware and efficient federated learning framework for cross-city traffic knowledge transfer. Specifically, our proposed framework includes three key innovations: (i) a traffic view imputation method for missing traffic data completion to enhance data quality, (ii) a traffic domain adapter for uniform traffic data transformation to address data distribution discrepancies, and (iii) a traffic secret aggregation protocol for secure traffic data aggregation to safeguard data privacy. Extensive experiments on 4 real-world datasets demonstrate that the proposed FedTT framework outperforms the 14 state-of-the-art baselines.

cs.LG

Spatio-temporal characterization of nonlinear forcing and response in turbulent channel flow

The quadratic convection term in the incompressible Navier-Stokes equations is considered as a nonlinear forcing to the linear resolvent operator, and it is studied in the Fourier domain through the analysis of interactions between triadically compatible wavenumber-frequency triplets. A framework to quantify the triadic contributions to the forcing and response by each pair of triplets is developed and applied to data from direct numerical simulations of a turbulent channel at $Re_{\tau} \approx 550$. The linear resolvent operator is incorporated to provide the missing link from energy transfer between modes to the effect on the spectral turbulent kinetic energy. The coefficients highlight the importance of interactions involving large-scale structures, providing a natural connection to the modeling assumptions in quasi-linear (QL) and generalized quasi-linear (GQL) analyses. Specifically, it is revealed that the QL and GQL reductions efficiently capture important triadic interactions in the flow, especially when including of a small number of wavenumbers into the GQL large-scale base flow. Additionally, spatio-temporal analyses of the triadic contributions to a single mode representative of the near-wall cycle demonstrate the spatio-temporal nature of the triadic interactions and the effect of the resolvent operator, which selectively amplifies certain forcing profiles. The tools presented are expected to be useful for improving modeling of the nonlinearity, especially in QL, GQL, and resolvent analyses, and understanding the amplitude modulation mechanism relating large-scale fluctuations to the modulation of near-wall structures.

physics.flu-dyn

Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction

As Large Language Models (LLMs) are widely applied in various domains, the safety of LLMs is increasingly attracting attention to avoid their powerful capabilities being misused. Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or suffix tokens to achieve a specific representation manually or automatically. However, they suffer from low efficiency and explicit jailbreak patterns, far from the real deployment of mass attacks to LLMs. In this paper, we point out that simply rewriting the original instruction can achieve a jailbreak, and we find that this rewriting approach is learnable and transferable. We propose the Rewrite to Jailbreak (R2J) approach, a transferable black-box jailbreak method to attack LLMs by iteratively exploring the weakness of the LLMs and automatically improving the attacking strategy. The jailbreak is more efficient and hard to identify since no additional features are introduced. Extensive experiments and analysis demonstrate the effectiveness of R2J, and we find that the jailbreak is also transferable to multiple datasets and various types of models with only a few queries. We hope our work motivates further investigation of LLM safety. The code can be found at https://github.com/ythuang02/R2J/.

cs.CL

Intelligent Legal Assistant: An Interactive Clarification System for Legal Question Answering

The rise of large language models has opened new avenues for users seeking legal advice. However, users often lack professional legal knowledge, which can lead to questions that omit critical information. This deficiency makes it challenging for traditional legal question-answering systems to accurately identify users' actual needs, often resulting in imprecise or generalized advice. In this work, we develop a legal question-answering system called Intelligent Legal Assistant, which interacts with users to precisely capture their needs. When a user poses a question, the system requests that the user select their geographical location to pinpoint the applicable laws. It then generates clarifying questions and options based on the key information missing from the user's initial question. This allows the user to select and provide the necessary details. Once all necessary information is provided, the system produces an in-depth legal analysis encompassing three aspects: overall conclusion, jurisprudential analysis, and resolution suggestions.

cs.CL

Spatio-temporal characterization of non-linear forcing in turbulence

The quadratic convection term in the incompressible Navier-Stokes equations is considered as a non-linear forcing to the linear operator, and it is studied in the Fourier domain through the analysis of interactions between triadically compatible wavenumber-frequency triplets. Interaction coefficients are proposed to quantify the contribution to the forcing by each pair of triplets and are computed using data from direct numerical simulations of a turbulent channel at $Re_{\tau} \approx 550$. The coefficients show the importance of interactions involving streamwise large scales. The regions of non-linear interactions permitted under the quasi-linear (QL) and generalized quasi-linear (GQL) assumptions are shown to be significant contributors to the forcing and increasing the number of GQL-large scales is shown to monotonically increase the total forcing captured, providing a possible reason for the success of QL and GQL simulations. The tools presented are expected to be useful for improving modeling of the nonlinearity, especially in QL, GQL, and resolvent analyses, and understanding the amplitude modulation mechanism relating large-scale fluctuations to the modulation of near-wall structures.

physics.flu-dyn

Direct identification of Mott Hubbard band pattern beyond charge density wave superlattice in monolayer 1T-NbSe2

Understanding Mott insulators and charge density waves (CDW) is critical for both fundamental physics and future device applications. However, the relationship between these two phenomena remains unclear, particularly in systems close to two-dimensional (2D) limit. In this study, we utilize scanning tunneling microscopy/spectroscopy to investigate monolayer 1T-NbSe2 to elucidate the energy of the Mott upper Hubbard band (UHB), and reveal that the spin-polarized UHB is spatially distributed away from the dz2 orbital at the center of the CDW unit. Moreover, the UHB shows a root3 x root3 R30{\deg} periodicity in addition to the typically observed CDW pattern. Furthermore, a pattern similar to the CDW order is visible deep in the Mott gap, exhibiting CDW without contribution of the Mott Hubbard band. Based on these findings in monolayer 1T-NbSe2, we provide novel insights into the relation between the correlated and collective electronic structures in monolayer 2D systems.

cond-mat.str-el