SearcharxivSearch

arXiv subjects

Ting Zhu

Publications and source records attributed to Ting Zhu.

At least 19 recordsLinked to original sources

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples. We propose two variants: CritICL-dynamic, which adaptively predicts input-specific failure modes and retrieves critiques, and CritICL-static, which uses a global failure mode profile to provide stable guidance. Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost. Code available at: https://github.com/umwyf/CRITICL

cs.CL

Phase-field modeling of elastically driven abnormal grain growth

Grain-refined metals typically exhibit high strength, yet their engineering applications are often constrained by grain coarsening under thermo-mechanical loading. Recent experiments have revealed abnormal grain growth (AGG) in ultrafine-grained Ni thin films subjected to cyclic loading at room temperature. Unlike conventional AGG, which generally requires significant plastic deformation or high temperatures, this phenomenon occurs within the regime of macroscopic elastic deformation. This AGG is characterized by the preferential growth of grains with an in-plane <100> orientation aligned with the loading direction. Here, we investigate the underlying physical mechanisms by combining phase-field simulations with micromechanical analysis. The results indicate that elastic energy reduction provides a thermodynamically plausible driving force for this orientation-selective grain growth. Phase-field simulations reveal the evolution kinetics of AGG and confirm that local grain geometry and stress states play critical roles in determining the grain growth pathway. By applying this framework to systems with varying elastic anisotropy, we establish a general approach for investigating elastically driven AGG in polycrystalline materials.

cond-mat.mtrl-sci

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored, requiring comprehensive e-commercial dataset and elaborative utilization of multi-modal condition. In this paper, we propose a brand-new e-commerce dataset, named Fashion130k, with various occasions, models, and garment types. For the consistent generation of garment, we design a framework with Unified Multi-modal Condition (UMC) to align and integrate the text and visual prompts into generation model. Specifically, we explore an embedding refiner to extract the unified embeddings of multi-modal prompts, within which a Fusion Transformer is proposed to align the multi-modal embeddings by adjusting the modality gap between text and image. Based on unified embeddings, the attention in generation model is redesigned to emphasis the correlations between prompts and noise image, inducing that the noise image can select the pivotal tokens of prompts for consistent outfit generation. Our dataset and proposed framework offer a general and nuanced exploration of multi-modal prompts for generation models. Extensive experiments on real-world applications and benchmark demonstrate the effectiveness of UMC in visual consistency, achieving promising result than that of SoTA methods.

cs.CV

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the image. Previous methods often involve complex steps of specifying the text content and attributes, such as font size, color, and layout, without considering the stylistic consistency with the reference image. To address this, we propose UM-Text, a unified multimodal model for context understanding and visual text editing by natural language instructions. Specifically, we introduce a Visual Language Model (VLM) to process the instruction and reference image, so that the text content and layout can be elaborately designed according to the context information. To generate an accurate and harmonious visual text image, we further propose the UM-Encoder to combine the embeddings of various condition information, where the combination is automatically configured by VLM according to the input instruction. During training, we propose a regional consistency loss to offer more effective supervision for glyph generation on both latent and RGB space, and design a tailored three-stage training strategy to further enhance model performance. In addition, we contribute the UM-DATA-200K, a large-scale visual text image dataset on diverse scenes for model training. Extensive qualitative and quantitative results on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance.

cs.CV

Chain Reactions in Space: Analyzing the Impact of Satellite Collisions and Debris Accumulation

The exponential increase in artificial satellites, growing from 852 in 2004 to over 9,000 in 2023, has intensified the risk of the Kessler Syndrome: a cascading chain reaction of orbital collisions. This paper analyzes the dynamics of space debris accumulation to identify the primary orbital features contributing to this systemic risk. We compiled and analyzed Two-Line Element (TLE) datasets from Space-Track.org and historical collision data using a Python-based data mining approach. Specifically, we derived satellite velocities using the Vis-Viva equation and evaluated the correlation of five key features, launch piece count, orbital period, apogee, perigee, and Radar Cross Section (RCS) size, with debris density. Our evaluation reveals that apogee and orbital period exhibit the strongest correlation with the risk of the Kessler Syndrome, indicating that satellites in higher orbits pose a disproportionately greater threat to long-term sustainability due to navigational constraints. Contrary to common assumptions, our data suggests that velocity and object size (RCS) show negligible direct correlation with collision incidence in the current dataset. Based on these findings, we propose mitigation strategies focusing on integrating AI-driven autonomous navigation systems and deploying advanced radiation-resistant shielding materials to enhance the resilience of high-orbit assets.

astro-ph.EP

SLASh: Simulation of LISLs Aboard LEO Satellite Shells

Recent advances in satellite technology have introduced a new frontier of wireless networking by establishing Low Earth Orbit (LEO) Satellite networks that work to connect difficult to reach areas and improve global connectivity. These novel advancements lack robust open-source simulation models that can highlight potential bottlenecks or potential wasted resources, wasting terrestrial users and the companies that provide these networks time and money. To that end, we propose SLASh, a highly-customizable satellite network simulation which allows users to design a simulated network with specific characteristics, and constructs them analog to real-world conditions. Additionally, SLASh can generate abstract telemetry that can be simulated moving throughout the network, allowing users to compare network capabilities across a variety of frameworks.

cs.NI

Satellite Cybersecurity Across Orbital Altitudes: Analyzing Ground-Based Threats to LEO, MEO, and GEO

The rapid proliferation of satellite constellations, particularly in Low Earth Orbit (LEO), has fundamentally altered the global space infrastructure, shifting the risk landscape from purely kinetic collisions to complex cyber-physical threats. While traditional safety frameworks focus on debris mitigation, ground-based adversaries increasingly exploit radio-frequency links, supply chain vulnerabilities, and software update pathways to degrade space assets. This paper presents a comparative analysis of satellite cybersecurity across LEO, Medium Earth Orbit (MEO), and Geostationary Earth Orbit (GEO) regimes. By synthesizing data from 60 publicly documented security incidents with key vulnerability proxies--including Telemetry, Tracking, and Command (TT&C) anomalies, encryption weaknesses, and environmental stressors--we characterize how orbital altitude dictates attack feasibility and impact. Our evaluation reveals distinct threat profiles: GEO systems are predominantly targeted via high-frequency uplink exposure, whereas LEO constellations face unique risks stemming from limited power budgets, hardware constraints, and susceptibility to thermal and radiation-induced faults. We further bridge the gap between security and sustainability, arguing that unmitigated cyber vulnerabilities accelerate hardware obsolescence and debris accumulation, undermining efforts toward carbon-neutral space operations. The results demonstrate that weak encryption and command path irregularities are the most consistent predictors of adversarial success across all orbits.

cs.CR

Investigating How MacBook Accessories Evolve across Generations, and Their Potential Environmental, Economical Impacts

The technological transition of MacBook charging solutions from MagSafe to USB-C, followed by a return to MagSafe 3, encapsulates the dynamic interplay between technological advancement, environmental considerations, and economic factors. This study delves into the broad implications of these charging technology shifts, particularly focusing on the environmental repercussions associated with electronic waste and the economic impacts felt by both manufacturers and consumers. By investigating the lifecycle of these technologies - from development and market introduction through to their eventual obsolescence - this paper underscores the importance of devising strategies that not only foster technological innovation but also prioritize environmental sustainability and economic feasibility. This comprehensive analysis illuminates the crucial factors influencing the evolution of charging technologies and their wider societal and environmental implications, advocating for a balanced approach that ensures technological progress does not compromise ecological health or economic stability.

cs.CY

EdgeFlex-Transformer: Transformer Inference for Edge Devices

Deploying large-scale transformer models on edge devices presents significant challenges due to strict constraints on memory, compute, and latency. In this work, we propose a lightweight yet effective multi-stage optimization pipeline designed to compress and accelerate Vision Transformers (ViTs) for deployment in resource-constrained environments. Our methodology combines activation profiling, memory-aware pruning, selective mixed-precision execution, and activation-aware quantization (AWQ) to reduce the model's memory footprint without requiring costly retraining or task-specific fine-tuning. Starting from a ViT-Huge backbone with 632 million parameters, we first identify low-importance channels using activation statistics collected via forward hooks, followed by structured pruning to shrink the MLP layers under a target memory budget. We further apply FP16 conversion to selected components and leverage AWQ to quantize the remaining model weights and activations to INT8 with minimal accuracy degradation. Our experiments on CIFAR-10 demonstrate that the fully optimized model achieves a 76% reduction in peak memory usage and over 6x lower latency, while retaining or even improving accuracy compared to the original FP32 baseline. This framework offers a practical path toward efficient transformer inference on edge platforms, and opens future avenues for integrating dynamic sparsity and Mixture-of-Experts (MoE) architectures to further scale performance across diverse tasks.

cs.LG

On-device Large Multi-modal Agent for Human Activity Recognition

Human Activity Recognition (HAR) has been an active area of research, with applications ranging from healthcare to smart environments. The recent advancements in Large Language Models (LLMs) have opened new possibilities to leverage their capabilities in HAR, enabling not just activity classification but also interpretability and human-like interaction. In this paper, we present a Large Multi-Modal Agent designed for HAR, which integrates the power of LLMs to enhance both performance and user engagement. The proposed framework not only delivers activity classification but also bridges the gap between technical outputs and user-friendly insights through its reasoning and question-answering capabilities. We conduct extensive evaluations using widely adopted HAR datasets, including HHAR, Shoaib, Motionsense to assess the performance of our framework. The results demonstrate that our model achieves high classification accuracy comparable to state-of-the-art methods while significantly improving interpretability through its reasoning and Q&A capabilities.

cs.LG

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly annotation, underscoring the need for more flexible feeling-aligned methods. To support multimodal music generation, we construct ArtiCaps, a pseudo feeling-aligned image-music-text dataset created by semantically matching descriptions from ArtEmis and MusicCaps. We further propose Art2Music, a lightweight cross-modal framework that synthesizes music from artistic images and user comments. In the first stage, images and text are encoded with OpenCLIP and fused using a gated residual module; the fused representation is decoded by a bidirectional LSTM into Mel-spectrograms with a frequency-weighted L1 loss to enhance high-frequency fidelity. In the second stage, a fine-tuned HiFi-GAN vocoder reconstructs high-quality audio waveforms. Experiments on ArtiCaps show clear improvements in Mel-Cepstral Distortion, Frechet Audio Distance, Log-Spectral Distance, and cosine similarity. A small LLM-based rating study further verifies consistent cross-modal feeling alignment and offers interpretable explanations of matches and mismatches across modalities. These results demonstrate improved perceptual naturalness, spectral fidelity, and semantic consistency. Art2Music also maintains robust performance with only 50k training samples, providing a scalable solution for feeling-aligned creative audio generation in interactive art, personalized soundscapes, and digital art exhibitions.

cs.SD

Defining structural gradient hardening through Type II back stress for heterostructured materials

The recently proposed term "heterostructured (HS) materials" serves as an umbrella classification encompassing a wide range of materials that hold great promise for enhanced mechanical properties. Most HS materials exhibit back-stress strengthening, as is typical for all plastically non-homogeneous materials. To better embody the distinctiveness of materials crafted via innovative heterostructuring, here we introduce the concept of "structural gradient hardening" (SGH), which captures an essential feature of HS materials and complements traditional strengthening mechanisms. SGH refers to the extra strengthening that arises from a characteristic gradient structure introduced by heterostructuring, beyond what is predicted by the rule of mixtures. This distinction is useful, as the overall back stress can in fact be partitioned into Type I and Type II components, with the latter specifically quantifying the extra hardening originating from the structural and strain gradients established by heterostructuring, as articulated in this Viewpoint article.

cond-mat.mtrl-sci

SmartFlow: A CFD-solver-agnostic deep reinforcement learning framework for computational fluid dynamics on HPC platforms

Deep reinforcement learning (DRL) is emerging as a powerful tool for fluid-dynamics research, encompassing active flow control, autonomous navigation, turbulence modeling and discovery of novel numerical schemes. We introduce SmartFlow, a CFD-solver-agnostic framework for both single- and multi-agent DRL algorithms that can easily integrate with MPI-parallel CPU and GPU-accelerated solvers. Built on Relexi and SmartSOD2D, SmartFlow uses the SmartSim infrastructure library and our newly developed SmartRedis-MPI library to enable asynchronous, low-latency, in-memory communication between CFD solvers and Python-based DRL algorithms. SmartFlow leverages PyTorch's Stable-Baselines3 for training, which provides a modular, Gym-like environment API. We demonstrate its versatility via three case studies: single-agent synthetic-jet control for drag reduction in a cylinder flow simulated by the high-order FLEXI solver, multi-agent cylinder wake control using the GPU-accelerated spectral-element code SOD2D, and multi-agent wall-model learning for large-eddy simulation with the finite-difference solver CaLES. SmartFlow's CFD-solver-agnostic design and seamless HPC integration is promising to accelerate RL-driven fluid-mechanics studies.

physics.flu-dyn

Analysis of Security in OS-Level Virtualization

Virtualization is a technique that allows multiple instances typically running different guest operating systems on top of single physical hardware. A hypervisor, a layer of software running on top of the host operating system, typically runs and manages these different guest operating systems. Rather than to run different services on different servers for reliability and security reasons, companies started to employ virtualization over their servers to run these services within a single server. This approach proves beneficial to the companies as it provides much better reliability, stronger isolation, improved security and resource utilization compared to running services on multiple servers. Although hypervisor based virtualization offers better resource utilization and stronger isolation, it also suffers from high overhead as the host operating system has to maintain different guest operating systems. To tackle this issue, another form of virtualization known as Operating System-level virtualization has emerged. This virtualization provides light-weight, minimal and efficient virtualization, as the different instances are run on top of the same host operating system, sharing the resources of the host operating system. But due to instances sharing the same host operating system affects the isolation of the instances. In this paper, we will first establish the basic concepts of virtualization and point out the differences between the hyper-visor based virtualization and operating system-level virtualization. Next, we will discuss the container creation life-cycle which helps in forming a container threat model for the container systems, which allows to map different potential attack vectors within these systems. Finally, we will discuss a case study, which further looks at isolation provided by the containers.

cs.CR

Optimizing Global Quantum Communication via Satellite Constellations

In this paper, we investigate the optimization of global quantum communication through satellite constellations. We address the challenge of quantum key distribution (QKD) across vast distances and the limitations posed by terrestrial fiber-optic networks. Our research focuses on the configuration of satellite constellations to improve QKD between ground stations and the application of innovative orbital mechanics to reduce latency in quantum information transfer. We introduce a novel approach using quantum relay satellites in Molniya orbits, enhancing communication efficiency and coverage. The use of these high eccentricity orbits allows us to extend the operational presence of satellites over targeted hemispheres, thus maximizing the quantum network's reach. Our findings provide a strategic framework for deploying quantum satellites and relay systems to achieve a robust and efficient global quantum communication network.

quant-ph

Achieving Carbon Neutrality for I/O Devices

Achieving carbon neutrality has become a critical goal in mitigating the environmental impacts of human activities, particularly in the face of global climate challenges. Input/Output (I/O) devices, such as keyboards, mice, displays, and printers, contribute significantly to greenhouse gas emissions through their manufacturing, operation, and disposal processes. In this paper, we explores sustainable strategies for achieving carbon neutrality in I/O devices, emphasizing the importance of environmentally conscious design and development. Through a comprehensive review of existing literature and best approaches, we introduces a framework to outline approaches for reducing the carbon footprint of I/O devices. The result underscore the necessity of integrating sustainability into the lifecycle of I/O devices to support global carbon neutrality goals and promote long-term environmental sustainability.

cs.CY

Environmental and Economic Impact of I/O Device Obsolescence

This paper analyzes the proportion of Input/output devices made obsolete by changes in technology generations. This obsolescence may be by new software/hardware generations rendering otherwise functional devices unusable. Concluding with brief analysis on the economic and environmental impacts of the e-waste produced.

cs.CY

Energy Efficient LoRaWAN in LEO Satellites

LPWAN service's inexpensive cost and long range capabilities make it a promising addition and countless satellite companies have started taking advantage of this technology to connect IoT users across the globe. However, LEO satellites have the unique challenge of using rechargeable batteries and green solar energy to power their components. LPWAN technology is not optimized to maximize battery lifespan of network nodes. By incorporating a MAC protocol that maximizes node the battery lifespan across the network, we can reduce battery waste and usage of scarce Earth resources to develop satellite batteries.

cs.ET