SearcharxivSearch

arXiv subjects

Yikai Li

Publications and source records attributed to Yikai Li.

14 recordsLinked to original sources

FOX: Visual Exploration of Data Fact Outliers

Exploratory Data Analysis (EDA) systems extract and present data facts to summarize meaningful patterns such as trends and correlations for efficient dataset exploration. However, existing approaches rarely consider outlier detection at the level of data facts,and heterogeneous facts from different analytical scopes are often aggregated in a single view, making it difficult to define meaningful metrics and effectively analyze data fact outliers. To fill this gap, we present FOX, a novel visual analytics system for interactive data Fact Outlier eXploration. FOX organizes data facts into groups with consistent analytical scopes and computes a unified outlier score that combines distribution-based and pattern-based components. Its interface comprises an Upload Panel for data preparation and two coordinated exploration panels: the Overview Panel employs a matrix-based visualization to enable an intuitive overview of all data facts, and the Main Panel provides four linked views for cluster-level and fact-level analysis. We evaluated the usability and effectiveness of the system through two usage scenarios on public datasets and in-depth interviews with 12 participants. The results show that FOX enables meaningful detection, analysis, and explanation of data fact outliers.

cs.HC

PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving

Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. Serving them is hard because the workload inverts what GPUs provide: terabytes of memory against only tens of TFLOPS, and because every published system stages its adapters from CPU DRAM over PCIe, where each access pays a kernel stop and a host-run copy and capacity ends at the motherboard's DIMM slots. Meanwhile, memory-semantic fabrics such as CXL and NVLink are converging on pooled memory that an accelerator addresses with its own loads and stores, and near-data processing (NDP) can place compute beside the pooled data. How to serve multi-LoRA workloads on such hardware remains unexplored. This paper introduces PLoRA, an NDP-enhanced pooled-memory system for cost-efficient multi-LoRA serving. PLoRA keeps adapters and KV cache in the pool and returns only reduced results over the link, through a read-compute interface the GPU drives with its own loads and stores. Above this architecture, a GPU memory management system picks among four LoRA and two attention execution strategies for each adapter and caches the most performance-critical bytes in GPU memory, guided by a link-parameterized cost model. On one H100 serving 1000 adapters, PLoRA attains the lowest decode latency on every model and workload we measure, averaging 6.6x below a real-machine S-LoRA at under 3.4% added device area. The link itself stops mattering: throughput saturates at 32 GB/s on short contexts, a quarter of CXL 3.1, and the verdict survives scale: per-GPU demand falls from 7B to a modeled 1.2T deployment once adapter traffic shards with the tensor parallelism. The design runs unchanged from CXL-class to NVLink-class fabrics, and surplus bandwidth buys pooled capacity rather than speed.

cs.AR

Spontaneous translation of charged droplets during evaporation on dry surfaces

Evaporating sessile droplets are usually treated as capillary objects, but droplets generated by routine handling can carry tens to hundreds of picocoulombs of electric charge. Here we combine Faraday-cup charge measurements with optical imaging to determine how such charge evolves as water droplets evaporate on dry polymer substrates. A zero-time protocol shows that a reproducible initial charge is preserved on poly(methylpentene) (PMP), whereas PDMS, SOCAL-coated surfaces, and polystyrene either exchange, dissipate, or inject charge on contact. On PMP, ensemble-resolved measurements reveal two regimes: the charge remains nearly constant during early evaporation and then decreases abruptly once the droplet reaches a small-volume state. This charge collapse coincides with spontaneous lateral translation rather than jetting or breakup. A Rayleigh-normalized analysis, including a spherical-cap stress correction and measured contact-angle retention scale, shows that motion occurs only after evaporation drives the droplet into a high electro-pinning state. High-speed imaging and kinematic analysis support a picture in which the subsequent motion is governed by repeated contact-line depinning and re-pinning: the total distance traveled is strongly affected by dry-surface pinning, whereas the peak translational velocity serves as a more robust indicator of the discharge strength. These results identify a dry-substrate mode of evaporation-driven electrostatic relaxation, distinct from Coulomb fission on lubricated surfaces, in which substrate electrostatic passivity enables charge retention, droplet geometry selects the instability onset, and whole-droplet translation provides the charge-release pathway.

cond-mat.soft

JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation

Modern Text-to-Image (T2I) generation increasingly relies on token-centric architectures that are trained with self-supervision, yet effectively fusing text with visual tokens remains a challenge. We propose \textbf{JEPA-T}, a unified multimodal framework that encodes images and captions into discrete visual and textual tokens, processed by a joint-embedding predictive Transformer. To enhance fusion, we incorporate cross-attention after the feature predictor for conditional denoising while maintaining a task-agnostic backbone. Additionally, raw texts embeddings are injected prior to the flow matching loss to improve alignment during training. During inference, the same network performs both class-conditional and free-text image generation by iteratively denoising visual tokens conditioned on text. Evaluations on ImageNet-1K demonstrate that JEPA-T achieves strong data efficiency, open-vocabulary generalization, and consistently outperforms non-fusion and late-fusion baselines. Our approach shows that late architectural fusion combined with objective-level alignment offers an effective balance between conditioning strength and backbone generality in token-based T2I.The code is now available: https://github.com/justin-herry/JEPA-T.git

cs.CV

RTBAgent: A LLM-based Agent System for Real-Time Bidding

Real-Time Bidding (RTB) enables advertisers to place competitive bids on impression opportunities instantaneously, striving for cost-effectiveness in a highly competitive landscape. Although RTB has widely benefited from the utilization of technologies such as deep learning and reinforcement learning, the reliability of related methods often encounters challenges due to the discrepancies between online and offline environments and the rapid fluctuations of online bidding. To handle these challenges, RTBAgent is proposed as the first RTB agent system based on large language models (LLMs), which synchronizes real competitive advertising bidding environments and obtains bidding prices through an integrated decision-making process. Specifically, obtaining reasoning ability through LLMs, RTBAgent is further tailored to be more professional for RTB via involved auxiliary modules, i.e., click-through rate estimation model, expert strategy knowledge, and daily reflection. In addition, we propose a two-step decision-making process and multi-memory retrieval mechanism, which enables RTBAgent to review historical decisions and transaction records and subsequently make decisions more adaptive to market changes in real-time bidding. Empirical testing with real advertising datasets demonstrates that RTBAgent significantly enhances profitability. The RTBAgent code will be publicly accessible at: https://github.com/CaiLeng/RTBAgent.

cs.AI

AIGC-Assisted Digital Watermark Services in Low-Earth Orbit Satellite-Terrestrial Edge Networks

Low Earth Orbit (LEO) satellite communication is a crucial component of future 6G communication networks, contributing to the development of an integrated satellite-terrestrial network. In the forthcoming satellite-to-ground network, the idle computational resources of LEO satellites can serve as edge servers, delivering intelligent task computation services to ground users. Existing research on satellite-to-ground computation primarily focuses on designing efficient task scheduling algorithms to provide straightforward computation services to ground users. This study aims to integrate satellite edge networks with Artificial Intelligence-Generated Content (AIGC) technology to offer personalized AIGC services to ground users, such as customized digital watermarking services. Firstly, we propose a satellite-to-ground edge network architecture, enabling bidirectional communication between visible LEO satellites and ground users. Each LEO satellite is equipped with intelligent algorithms supporting various AIGC-assisted digital watermarking technologies with different precision levels. Secondly, considering metrics like satellite visibility, satellite-to-ground communication stability, digital watermark quality, satellite-to-ground communication time, digital watermarking time, and ground user energy consumption, we construct an AIGC-assisted digital watermarking model based on the satellite-to-ground edge network. Finally, we introduce a reinforcement learning-based task scheduling algorithm to obtain an optimal strategy. Experimental results demonstrate that our approach effectively meets the watermark generation needs of ground users, achieving a well-balanced trade-off between generation time and user energy consumption. We anticipate that this work will provide an effective solution for the intelligent services in satellite-to-ground edge networks.

cs.NI

Security-Sensitive Task Offloading in Integrated Satellite-Terrestrial Networks

With the rapid development of sixth-generation (6G) communication technology, global communication networks are moving towards the goal of comprehensive and seamless coverage. In particular, low earth orbit (LEO) satellites have become a critical component of satellite communication networks. The emergence of LEO satellites has brought about new computational resources known as the \textit{LEO satellite edge}, enabling ground users (GU) to offload computing tasks to the resource-rich LEO satellite edge. However, existing LEO satellite computational offloading solutions primarily focus on optimizing system performance, neglecting the potential issue of malicious satellite attacks during task offloading. In this paper, we propose the deployment of LEO satellite edge in an integrated satellite-terrestrial networks (ISTN) structure to support \textit{security-sensitive computing task offloading}. We model the task allocation and offloading order problem as a joint optimization problem to minimize task offloading delay, energy consumption, and the number of attacks while satisfying reliability constraints. To achieve this objective, we model the task offloading process as a Markov decision process (MDP) and propose a security-sensitive task offloading strategy optimization algorithm based on proximal policy optimization (PPO). Experimental results demonstrate that our algorithm significantly outperforms other benchmark methods in terms of performance.

eess.SP

Deep Reinforcement Learning for Privacy-Preserving Task Offloading in Integrated Satellite-Terrestrial Networks

Satellite communication networks have attracted widespread attention for seamless network coverage and collaborative computing. In satellite-terrestrial networks, ground users can offload computing tasks to visible satellites that with strong computational capabilities. Existing solutions on satellite-assisted task computing generally focused on system performance optimization such as task completion time and energy consumption. However, due to the high-speed mobility pattern and unreliable communication channels, existing methods still suffer from serious privacy leakages. In this paper, we present an integrated satellite-terrestrial network to enable satellite-assisted task offloading under dynamic mobility nature. We also propose a privacy-preserving task offloading scheme to bridge the gap between offloading performance and privacy leakage. In particular, we balance two offloading privacy, called the usage pattern privacy and the location privacy, with different offloading targets (e.g., completion time, energy consumption, and communication reliability). Finally, we formulate it into a joint optimization problem, and introduce a deep reinforcement learning-based privacy-preserving algorithm for an optimal offloading policy. Experimental results show that our proposed algorithm outperforms other benchmark algorithms in terms of completion time, energy consumption, privacy-preserving level, and communication reliability. We hope this work could provide improved solutions for privacy-persevering task offloading in satellite-assisted edge computing.

cs.CR

Dual-Function Radar-Communication System Aided by Intelligent Reflecting Surfaces

We propose a novel design of a dual-function radar communication (DFRC) system aided by an Intelligent Reflecting Surface (IRS). We consider a scenario with one target and multiple communication receivers, where there is no line-of-sight between the radar and the target. The radar precoding matrix and the IRS weights are optimally designed to maximize the weighted sum of the signal-to-noise ratio (SNR) at the radar receiver and the SNR at the communication receivers subject to power constraints and constant modulus constraints on the IRS weights. The problem is decoupled into two sub-problems, namely, waveform design and IRS weight design, and is solved via alternating optimization. The former subproblem is solved via linear programming, and the latter via manifold optimization with a quartic polynomial objective. The key contribution of this paper lies in solving the IRS weight design sub-problem that is based on the optimization of a quartic objective function in the IRS weights, and is subject to unit modulus-constraint on the IRS weights. Simulation results are provided to show the convergence behavior of the proposed algorithm under different system configurations, and the effectiveness of using IRS to improve radar and communication performance.

eess.SP

Theoretical study on the interfacial instability of a spherical droplet subject to vertical vibration

Interfacial instability would be aroused on a spherical liquid droplet when it is subject to external vertical vibration. In this paper, a linear analysis was conducted on this instability problem. The polar-angle dependent acceleration in the spherical coordinate is strongly coupled with the temporal and spatial component of the surface deformation displacement, which gives a recursion equation that implicitly expresses the dispersion relation between the growth rate and spherical mode numbers. The unstable regions (or unstable tongues) for the inviscid fluids considering latitudinal mode (longitudinal mode number m = 0) were derived and presented in the parameter plane. Compared with the solution of the spherical Faraday instability under radial vibration acceleration, the regions of harmonic unstable tongues for the mono-directional vibration case is much narrowed and the subharmonic unstable tongues almost become straight lines. The analysis shows that the latitudinal waves emerging on the spherical droplet surface ought to oscillate harmonically instead of subharmonically, which is opposite to the results for the case under radial vibration acceleration. A corresponding experiment of a liquid droplet lying on a vertically vibrating plate was conducted and the observations substantiate our theoretical predictions.

physics.flu-dyn

Multi-Plane Program Induction with 3D Box Priors

We consider two important aspects in understanding and editing images: modeling regular, program-like texture or patterns in 2D planes, and 3D posing of these planes in the scene. Unlike prior work on image-based program synthesis, which assumes the image contains a single visible 2D plane, we present Box Program Induction (BPI), which infers a program-like scene representation that simultaneously models repeated structure on multiple 2D planes, the 3D position and orientation of the planes, and camera parameters, all from a single image. Our model assumes a box prior, i.e., that the image captures either an inner view or an outer view of a box in 3D. It uses neural networks to infer visual cues such as vanishing points, wireframe lines to guide a search-based algorithm to find the program that best explains the image. Such a holistic, structured scene representation enables 3D-aware interactive image editing operations such as inpainting missing pixels, changing camera parameters, and extrapolate the image contents.

cs.CV

Perspective Plane Program Induction from a Single Image

We study the inverse graphics problem of inferring a holistic representation for natural images. Given an input image, our goal is to induce a neuro-symbolic, program-like representation that jointly models camera poses, object locations, and global scene structures. Such high-level, holistic scene representations further facilitate low-level image manipulation tasks such as inpainting. We formulate this problem as jointly finding the camera pose and scene structure that best describe the input image. The benefits of such joint inference are two-fold: scene regularity serves as a new cue for perspective correction, and in turn, correct perspective correction leads to a simplified scene structure, similar to how the correct shape leads to the most regular texture in shape from texture. Our proposed framework, Perspective Plane Program Induction (P3I), combines search-based and gradient-based algorithms to efficiently solve the problem. P3I outperforms a set of baselines on a collection of Internet images, across tasks including camera pose estimation, global structure inference, and down-stream image manipulation tasks.

cs.CV

Program-Guided Image Manipulators

Humans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structures involves reasoning over repetition and symmetry of the objects in the image. In this paper, we present the Program-Guided Image Manipulator (PG-IM), inducing neuro-symbolic program-like representations to represent and manipulate images. Given an image, PG-IM detects repeated patterns, induces symbolic programs, and manipulates the image using a neural network that is guided by the program. PG-IM learns from a single image, exploiting its internal statistics. Despite trained only on image inpainting, PG-IM is directly capable of extrapolation and regularity editing in a unified framework. Extensive experiments show that PG-IM achieves superior performance on all the tasks.

cs.CV

Auto-Retoucher(ART) - A framework for Background Replacement and Image Editing

Replacing the background and simultaneously adjusting foreground objects is a challenging task in image editing. Current techniques for generating such images relies heavily on user interactions with image editing softwares, which is a tedious job for professional retouchers. To reduce their workload, some exciting progress has been made on generating images with a given background. However, these models can neither adjust the position and scale of the foreground objects, nor guarantee the semantic consistency between foreground and background. To overcome these limitations, we propose a framework -- ART(Auto-Retoucher), to generate images with sufficient semantic and spatial consistency. Images are first processed by semantic matting and scene parsing modules, then a multi-task verifier model will give two confidence scores for the current background and position setting. We demonstrate that our jointly optimized verifier model successfully improves the visual consistency, and our ART framework performs well on images with the human body as foregrounds.

cs.CV