SearcharxivSearch

arXiv subjects

Chao Lin

Publications and source records attributed to Chao Lin.

17 recordsLinked to original sources

Exploring the Performance Frontier of Compact Unified Image Generation Models

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.

cs.CV

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

cs.CV

CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and introduces multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings, serving as an effective proxy for public human evaluations.

cs.CV

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Recent advances in image generation and editing have opened new opportunities for virtual try-on. However, existing methods still struggle to meet complex real-world demands. We present Tstars-Tryon 1.0, a commercial-scale virtual try-on system that is robust, realistic, versatile, and highly efficient. First, our system maintains a high success rate across challenging cases like extreme poses, severe illumination variations, motion blur, and other in-the-wild conditions. Second, it delivers highly photorealistic results with fine-grained details, faithfully preserving garment texture, material properties, and structural characteristics, while largely avoiding common AI-generated artifacts. Third, beyond apparel try-on, our model supports flexible multi-image composition (up to 6 reference images) across 8 fashion categories, with coordinated control over person identity and background. Fourth, to overcome the latency bottlenecks of commercial deployment, our system is heavily optimized for inference speed, delivering near real-time generation for a seamless user experience. These capabilities are enabled by an integrated system design spanning end-to-end model architecture, a scalable data engine, robust infrastructure, and a multi-stage training paradigm. Extensive evaluation and large-scale product deployment demonstrate that Tstars-Tryon1.0 achieves leading overall performance. To support future research, we also release a comprehensive benchmark. The model has been deployed at an industrial scale on the Taobao App, serving millions of users with tens of millions of requests.

cs.CV

A Unified Benchmark Study of Shock-Like Problems in Two-Dimensional Steady Electrohydrodynamic Flow Based on LSTM-PINN

Accurately resolving steady electrohydrodynamic (EHD) flows presents a formidable computational challenge due to the strong nonlinear coupling between charged-particle density, velocity fields, and electric potential. These interactions frequently induce sharp transition layers, crossing fronts, and multiscale spatial structures, which notoriously degrade the predictive accuracy of standard mesh-free solvers like Physics-Informed Neural Networks (PINNs). To systematically address this bottleneck, we formulate a unified four-variable operator framework and develop a comprehensive benchmark suite for two-dimensional steady EHD shock-like problems. The benchmark comprises eight rigorously designed cases featuring diverse front geometries, such as oblique, curved, and intersecting layers, alongside complex multiscale patterns. Under strictly identical configurations, including governing equations, source terms, sampling strategies, and loss formulations, we evaluate a Standard MLP-based PINN, a Residual Attention PINN (ResAtt-PINN), and an LSTM-PINN that leverages pseudo-sequential spatial encoding. Extensive numerical experiments demonstrate that the LSTM-PINN consistently achieves the highest predictive accuracy across all eight cases. It successfully reconstructs sharp gradients and intricate multiscale structures where other architectures fail or over-smooth. Furthermore, the LSTM backbone efficiently captures long-range spatial correlations while maintaining an exceptionally low computational overhead and GPU memory footprint. These findings not only establish the LSTM-PINN as a robust and efficient solver for strongly coupled PDEs with shock-like features, but also provide the computational physics community with a standardized, reproducible benchmark for future algorithmic evaluations.

physics.comp-ph

Widefield Nanodiamond Quantum Sensing Based on Light-Sheet Microscopy

Nanodiamonds containing nitrogen-vacancy (NV) centers are promising quantum sensors for biological applications thanks to their sub-micron spatial resolution, biocompatibility, and versatile multi-modal responses. However, the optically detected magnetic resonance (ODMR) measurement requires laser irradiation, creating a trade-off between high-throughput and low phototoxicity for applications in live cells. Here to address this challenge we develop a widefield quantum sensing method based on light-sheet microscopy (LSM), in which the sample is illuminated by a vertically movable laser sheet and the fluorescence is collected along the vertical axis that is orthogonal to the light sheet. This LSM-ODMR system is demonstrated to feature high throughput sensing due to the wide-field configuration, fast three-dimensional imaging and sensing due to the vertical mobility of the light sheet, enhanced sensitivity due to suppression of out-of-focus background fluorescence, and low phototoxicity for bio-sensing due to elimination of out-of-focus illumination. This LSM-based widefield nanodiamond sensing provides an approach for biological studies with low phototoxicity, offering three-dimensional and multi-modal sensing capability.

physics.app-ph

From Woofs to Words: Towards Intelligent Robotic Guide Dogs with Verbal Communication

Assistive robotics is an important subarea of robotics that focuses on the well-being of people with disabilities. A robotic guide dog is an assistive quadruped robot that helps visually impaired people in obstacle avoidance and navigation. Enabling language capabilities for robotic guide dogs goes beyond naively adding an existing dialog system onto a mobile robot. The novel challenges include grounding language in the dynamically changing environment and improving spatial awareness for the human handler. To address those challenges, we develop a novel dialog system for robotic guide dogs that uses LLMs to verbalize both navigational plans and scenes. The goal is to enable verbal communication for collaborative decision-making within the handler-robot team. In experiments, we conducted a human study to evaluate different verbalization strategies and a simulation study to assess the efficiency and accuracy in navigation tasks.

cs.RO

Rooftop Wind Field Reconstruction Using Sparse Sensors: From Deterministic to Generative Learning Methods

Real-time rooftop wind-speed distribution is important for the safe operation of drones and urban air mobility systems, wind control systems, and rooftop utilization. However, rooftop flows show strong nonlinearity, separation, and cross-direction variability, which make flow field reconstruction from sparse sensors difficult. This study develops a learning-from-observation framework using wind-tunnel experimental data obtained by Particle Image Velocimetry (PIV) and compares Kriging interpolation with three deep learning models: UNet, Vision Transformer Autoencoder (ViTAE), and Conditional Wasserstein GAN (CWGAN). We evaluate two training strategies, single wind-direction training (SDT) and mixed wind-direction training (MDT), across sensor densities from 5 to 30, test robustness under sensor position perturbations of plus or minus 1 grid, and optimize sensor placement via Proper Orthogonal Decomposition with QR decomposition. Results show that deep learning methods can reconstruct rooftop wind fields from sparse sensor data effectively. Compared with Kriging interpolation, the deep learning models improved SSIM by up to 32.7%, FAC2 by 24.2%, and NMSE by 27.8%. Mixed wind-direction training further improved performance, with gains of up to 173.7% in SSIM, 16.7% in FAC2, and 98.3% in MG compared with single-direction training. The results also show that sensor configuration, optimization, and training strategy should be considered jointly for reliable deployment. QR-based optimization improved robustness by up to 27.8% under sensor perturbations, although with metric-dependent trade-offs. Training on experimental rather than simulated data also provides practical guidance for method selection and sensor placement in different scenarios.

cs.CV

SilentLedger: Privacy-Preserving Auditing for Blockchains with Complete Non-Interactivity

Privacy-preserving blockchain systems are essential for protecting transaction data, yet they must also provide auditability that enables auditors to recover participant identities and transaction amounts when warranted. Existing designs often compromise the independence of auditing and transactions, introducing extra interactions that undermine usability and scalability. Moreover, many auditable solutions depend on auditors serving as validators or recording nodes, which introduces risks to both data security and system reliability. To overcome these challenges, we propose SilentLedger, a privacy-preserving transaction system with auditing and complete non-interactivity. To support public verification of authorization, we introduce a renewable anonymous certificate scheme with formal semantics and a rigorous security model. SilentLedger further employs traceable transaction mechanisms constructed from established cryptographic primitives, enabling users to transact without interaction while allowing auditors to audit solely from on-chain data. We formally prove security properties including authenticity, anonymity, confidentiality, and soundness, provide a concrete instantiation, and evaluate performance under a standard 2-2 transaction model. Our implementation and benchmarks demonstrate that SilentLedger achieves superior performance compared with state-of-the-art solutions.

cs.CR

Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks

Privacy-preserving federated learning (PPFL) aims to train a global model for multiple clients while maintaining their data privacy. However, current PPFL protocols exhibit one or more of the following insufficiencies: considerable degradation in accuracy, the requirement for sharing keys, and cooperation during the key generation or decryption processes. As a mitigation, we develop the first protocol that utilizes neural networks to implement PPFL, as well as incorporating an Aggregatable Hybrid Encryption scheme tailored to the needs of PPFL. We name these networks as Homomorphic Adversarial Networks (HANs) which demonstrate that neural networks are capable of performing tasks similar to multi-key homomorphic encryption (MK-HE) while solving the problems of key distribution and collaborative decryption. Our experiments show that HANs are robust against privacy attacks. Compared with non-private federated learning, experiments conducted on multiple datasets demonstrate that HANs exhibit a negligible accuracy loss (at most 1.35%). Compared to traditional MK-HE schemes, HANs increase encryption aggregation speed by 6,075 times while incurring a 29.2 times increase in communication overhead.

cs.CR

Towards Understanding and Enhancing Security of Proof-of-Training for DNN Model Ownership Verification

The great economic values of deep neural networks (DNNs) urge AI enterprises to protect their intellectual property (IP) for these models. Recently, proof-of-training (PoT) has been proposed as a promising solution to DNN IP protection, through which AI enterprises can utilize the record of DNN training process as their ownership proof. To prevent attackers from forging ownership proof, a secure PoT scheme should be able to distinguish honest training records from those forged by attackers. Although existing PoT schemes provide various distinction criteria, these criteria are based on intuitions or observations. The effectiveness of these criteria lacks clear and comprehensive analysis, resulting in existing schemes initially deemed secure being swiftly compromised by simple ideas. In this paper, we make the first move to identify distinction criteria in the style of formal methods, so that their effectiveness can be explicitly demonstrated. Specifically, we conduct systematic modeling to cover a wide range of attacks and then theoretically analyze the distinctions between honest and forged training records. The analysis results not only induce a universal distinction criterion, but also provide detailed reasoning to demonstrate its effectiveness in defending against attacks covered by our model. Guided by the criterion, we propose a generic PoT construction that can be instantiated into concrete schemes. This construction sheds light on the realization that trajectory matching algorithms, previously employed in data distillation, possess significant advantages in PoT construction. Experimental results demonstrate that our scheme can resist attacks that have compromised existing PoT schemes, which corroborates its superiority in security.

cs.CR

MetMamba: Regional Weather Forecasting with Spatial-Temporal Mamba Model

Deep Learning based Weather Prediction (DLWP) models have been improving rapidly over the last few years, surpassing state of the art numerical weather forecasts by significant margins. While much of the optimization effort is focused on training curriculum to extend forecast range in the global context, two aspects remains less explored: limited area modeling and better backbones for weather forecasting. We show in this paper that MetMamba, a DLWP model built on a state-of-the-art state-space model, Mamba, offers notable performance gains and unique advantages over other popular backbones using traditional attention mechanisms and neural operators. We also demonstrate the feasibility of deep learning based limited area modeling via coupled training with a global host model.

physics.ao-ph

FairRelay: Fair and Cost-Efficient Peer-to-Peer Content Delivery through Payment Channel Networks

Peer-to-Peer (P2P) content delivery, known for scalability and resilience, offers a decentralized alternative to traditional centralized Content Delivery Networks (CDNs). A significant challenge in P2P content delivery remains: the fair compensation of relayers for their bandwidth contributions. Existing solutions employ blockchains for payment settlements, however, they are not practical due to high on-chain costs and over-simplified network assumptions. In this paper, we introduce FairRelay, a fair and cost-efficient protocol that ensures all participants get fair payoff in complex content delivery network settings. We introduce a novel primitive, Enforceable Accumulative Hashed TimeLock Contract (Enforceable A-HTLC), designed to guarantee payment atomicity - ensuring all participants receive their payments upon successful content delivery. The fairness of FairRelay is proved using the Universal Composability (UC) framework. Our evaluation demonstrates that, in optimistic scenarios, FairRelay employs zero on-chain costs. In pessimistic scenarios, the on-chain dispute costs for relayers and customers are constant, irrespective of the network complexity. Specifically, empirical results indicate that the on-chain dispute costs for relayers and customers are 24,902 gas (equivalent to 0.01 USD on Optimism L2) and 290,797 gas (0.07 USD), respectively. In a 10-hop relay path, FairRelay introduces less than 1.5% additional overhead compared to pure data transmission, showcasing the efficiency of FairRelay.

cs.CR

RMGN: A Regional Mask Guided Network for Parser-free Virtual Try-on

Virtual try-on(VTON) aims at fitting target clothes to reference person images, which is widely adopted in e-commerce.Existing VTON approaches can be narrowly categorized into Parser-Based(PB) and Parser-Free(PF) by whether relying on the parser information to mask the persons' clothes and synthesize try-on images. Although abandoning parser information has improved the applicability of PF methods, the ability of detail synthesizing has also been sacrificed. As a result, the distraction from original cloth may persistin synthesized images, especially in complicated postures and high resolution applications. To address the aforementioned issue, we propose a novel PF method named Regional Mask Guided Network(RMGN). More specifically, a regional mask is proposed to explicitly fuse the features of target clothes and reference persons so that the persisted distraction can be eliminated. A posture awareness loss and a multi-level feature extractor are further proposed to handle the complicated postures and synthesize high resolution images. Extensive experiments demonstrate that our proposed RMGN outperforms both state-of-the-art PB and PF methods.Ablation studies further verify the effectiveness ofmodules in RMGN.

cs.CV

iQIYI-VID: A Large Dataset for Multi-modal Person Identification

Person identification in the wild is very challenging due to great variation in poses, face quality, clothes, makeup and so on. Traditional research, such as face recognition, person re-identification, and speaker recognition, often focuses on a single modal of information, which is inadequate to handle all the situations in practice. Multi-modal person identification is a more promising way that we can jointly utilize face, head, body, audio features, and so on. In this paper, we introduce iQIYI-VID, the largest video dataset for multi-modal person identification. It is composed of 600K video clips of 5,000 celebrities. These video clips are extracted from 400K hours of online videos of various types, ranging from movies, variety shows, TV series, to news broadcasting. All video clips pass through a careful human annotation process, and the error rate of labels is lower than 0.2\%. We evaluated the state-of-art models of face recognition, person re-identification, and speaker recognition on the iQIYI-VID dataset. Experimental results show that these models are still far from being perfect for the task of person identification in the wild. We proposed a Multi-modal Attention module to fuse multi-modal features that can improve person identification considerably. We have released the dataset online to promote multi-modal person identification research.

cs.CV

Optical monitoring of BL Lac object S5 0716+714 and FSRQ 3C273 from 2000 to 2014

Using the 1.56m telescope at the Shanghai Observatory (ShAO), China, we monitored two sources, BL Lac object S5 0716+714 and Flat Spectrum Radio Quasar (FSRQ) 3C 273. For S5 0716+714, we report 4969 sets of CCD (Charge-coupled Device) photometrical optical observations (1369 for V band, 1861 for R band and 1739 for I band) in the monitoring time from Dec.4, 2000 to Apr.5, 2014. For 3C 273, we report 460 observations (138 for V band, 146 for R band and 176 for I band) in the monitoring time from Mar. 28, 2006 to Apr. 9, 2014. The observations provide us with a large amount of data to analyze the short-term and long-term optical variabilities. Based on the variable timescales, we can estimate the central black hole mass and the Doppler factor. An abundance of multi-band observations can help us to analyze the relations between the brightness and spectrum. We use Gaussian fitting to analyze the intra-day light curves and obtain the intra-day variability (IDV) timescales. We use the discrete correlation function (DCF) method and Jurkevich method to analyze the quasi-periodic variability. Based on the VRI observations, we use the linear fitting to analyze the relations between brightness and spectrum. The two sources both show IDV properties for S5 0716+714. The timescales are in the range from 17.3 minutes to 4.82 hours; for 3C273, the timescale is 35.6 minutes. Based on the periodic analysis methods, we find the periods P(V) = 24.24 days, P(R)=24.12 days, P(I)=24.82 days for S5 0716+714, and P = 12.99, 21.76 yr for 3C273. The two sources displayed the "bluer-when-brighter" spectral evolution properties. S5 0716+714 and 3C 273 are frequently studied objects. The violent optical variability and IDV may come from the jet. Gaussian fitting can be used to analyze IDVs. The relations between brightness (flux density) and spectrum are strongly influenced by the frequency.

astro-ph.GA

The Ratio of the Core to the Extended Emissions in the Comoving Frame for Blazars

In a two-component jet model, the emissions are the sum of the core and extended emissions: $S^{\rm ob}=S_{\rm core}^{\rm ob}+S_{\rm ext}^{\rm ob}$, with the core emissions, $S_{\rm core}^{\rm ob}= f S_{\rm ext}^{\rm ob}\delta^{q}$, being a function of the Doppler factor, $\delta$, the extended emission, $S_{\rm ext}^{\rm ob}$, jet type dependent factor, $q$, and the ratio of the core to the extended emissions in the comoving frame, $f$. The $f$ is an unobservable but important parameter. Following our previous work, we collect 65 blazars with available Doppler factor, $\delta$, superluminal velocity, $\beta_{app}$, and core-dominance parameter, $R$, calculate the ratio, $f$, and peform statistical analyses. We find that the ratio, $f$, in BL Lacs is on average larger than that in FSRQs. We suggest that the difference of the ratio $f$ between FSRQs and BL Lacs is one of the possible reasons that cause the difference of other observed properties between them. We also find some significant correlations between $\log f$ and other parameters, including intrinsic (de-beamed) peak frequency, $\log \nu _{\rm p}^{\rm in}$, intrinsic polarization, $\log P^{\rm in}$, and core-dominance parameter, $\log R$, for the whole sample. In addition, we show that the ratio, $f$, can be estimated by $R$.

astro-ph.HE