SearcharxivSearch

arXiv subjects

Anqi Chen

Publications and source records attributed to Anqi Chen.

13 recordsLinked to original sources

ERNIE-Image Technical Report

We introduce ERNIE-Image, an open-source text-to-image generation model built upon an 8B single-stream DiT architecture. ERNIE-Image aims to bridge the gap between current open-source models and leading closed-source systems through more effective mining of large-scale pre-training data and improved supervision quality throughout training. During pre-training, we adopt a bottom-up data construction pipeline that combines fine-grained image categorization, rich caption annotation, aesthetic assessment, and hierarchical sampling. This strategy reduces data noise while preserving long-tail concepts and detailed real-world knowledge, providing a stronger foundation for complex generation tasks. In the post-training stage, we use a top-down data construction pipeline for high-demand scenarios, diversify prompt annotations to better match real user inputs, and apply a stabilized DPO strategy to align the model with human aesthetic preferences. We further train ERNIE-Image-Turbo for efficient 8-NFE generation and propose MT-DMD to mitigate capability drift during distillation. To make the model easier to use in practical scenarios, we equip it with a lightweight Prompt Enhancer that expands concise user intents into structured visual descriptions. In addition, we develop ERNIE-Image-Aes, an industrial-grade aesthetic model, together with ERNIE-Image-Aes-1K, a human-annotated benchmark for realistic aesthetic evaluation. Extensive qualitative and quantitative experiments show that ERNIE-Image achieves leading performance among open-source models and approaches top-tier commercial models in instruction following, text rendering, and aesthetic quality. We release the trained models and aesthetic resources to facilitate further academic research and technical progress in the AIGC community.

cs.CV

Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in practical deployment necessitates robust confidence estimation. Prior works have predominantly focused on text-only LLMs, often relying on computationally expensive self-consistency sampling. In this paper, we extend this to multimodal settings and conduct a comprehensive evaluation of MLLMs' response confidence estimation. Our analysis reveals a significant instinct-reflection misalignment: the model's implicit token-level support frequently diverges from its verbal self-assessment confidence. To address this misalignment, we propose a monotone confidence fusion framework to merge dual-channel signals and cross-channel consistency to estimate correctness. Subsequently, an order-preserving mean alignment step is applied to correct global bias, which improves calibration while preserving the risk-coverage trade-off for selective prediction. Experiments on diverse open-source and closed-source MLLMs show that our method consistently yields more reliable confidence estimates and improves both calibration and failure prediction. Code will be available at https://github.com/Yunkaidang/Instinct-vs.-Reflection.

cs.CV

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resolution, introducing substantial redundancy and irrelevant information. A common practice is to identify key image regions and refer to their high-resolution counterparts during reasoning, typically trained with external visual supervision. However, such visual supervision cues require costly grounding labels from human annotators. Meanwhile, it remains an open question how to enhance a model's grounding abilities to support reasoning without relying on additional annotations. In this paper, we propose High-resolution Annotation-free Reasoning Technique (HART), a closed-loop framework that enables LMMs to focus on and self-verify key regions of high-resolution visual inputs. HART incorporates a post-training paradigm in which we design Advantage Preference Group Relative Policy Optimization (AP-GRPO) to encourage accurate localization of key regions without external visual annotations. Notably, HART provides explainable reasoning pathways and enables efficient optimization of localization. Extensive experiments on MME-RealWorld-Lite, TreeBench, V* Bench, HR-Bench-4K/8K, and MMStar demonstrate that HART improves performance across a wide range of high-resolution visual tasks, consistently outperforming strong baselines.

cs.CV

Cross-Service Token: Finding Attacks in 5G Core Networks

5G marks a major departure from previous cellular architectures, by transitioning from a monolithic design of the core network to a Service-Based Architecture (SBA) where services are modularized as Network Functions (NFs) which communicate with each other via standard-defined HTTP-based APIs called Service-Based Interfaces (SBIs). These NFs are deployed in private and public cloud infrastructure, and an access control framework based on OAuth restricts how they communicate with each other and obtain access to resources. Given the increased vulnerabilities of clouds to insiders, it is important to study the security of the 5G Core services for vulnerabilities that allow attackers to use compromised NFs to obtain unauthorized access to resources. We present FivGeeFuzz, a grammar-based fuzzing framework designed to uncover security flaws in 5G core SBIs. FivGeeFuzz automatically derives grammars from 3GPP API specifications to generate malformed, unexpected, or semantically inconsistent inputs, and it integrates automated bug detection with manual validation and root-cause analysis. We evaluate our approach on free5GC, the only open-source 5G core implementing Release 17-compliant SBIs with an access control mechanism. Using FivGeeFuzz, we discovered 8 previously unknown vulnerabilities in free5GC, leading to runtime crashes, improper error handling, and unauthorized access to resources, including a very severe attack we call Cross-Service Token Attack. All bugs were confirmed by the free5GC team, 7 have already been patched, and the remaining one has a patch under development.

cs.CR

Mechanical performance of hybrid polymer-lipid vesicles with leaflet asymmetry engineered using microfluidics

Lipid vesicles consist of aqueous cores surrounded by a bilayer of phospholipids. Hybrid polymer-lipid vesicles incorporate both polymers and lipids, offering promising properties for developing pharmaceuticals, biosensors, and artificial cells. The hybrid vesicles can be symmetric, in which their two leaflets contain identical compositions, or asymmetric, in which the leaflets possess dissimilar compositions and can lead to dramatically modified properties. However, methods to produce both symmetric and asymmetric hybrid vesicles result in heterogenous compositions and sizes, making it challenging to quantify the effect of asymmetry and limiting applications. Here, we use a microfluidic approach to produce hybrid vesicles containing symmetric or asymmetric leaflets with precisely engineered compositions. We find the vesicles with asymmetric leaflets are significantly stiffer and tougher than those with symmetric leaflets; moreover, the lateral diffusivity of lipids is greatly decreased. The structure for improved toughness consists of an inner leaflet that is a stretchable lipid leaflet and an outer leaflet that is a fully continuous polymer leaflet. This technique of precisely engineering asymmetric structures may be applied to hybrid vesicles composed of block copolymers and phospholipids dissolvable in chloroform and hexane, further expanding their applications.

cond-mat.soft

Cargo Delivery to Cells Using Laser-Irradiated Carbon-Black-Loaded PDMS

Effective intracellular delivery is essential for successful gene editing of cells. Spatially selective delivery to cells that is simultaneously precise, consistent, and non-destructive remains challenging using conventional state-of-the-art techniques. Here, we introduce a carrier-free method for spatiotemporal delivery of fluorescently labeled cargo into both adherent and suspension cells using carbon-black-embedded polydimethylsiloxane (PDMS) substrates irradiated by nanosecond laser pulses. This low-cost, biocompatible material, coupled with an optical approach, enables scalable, spatially selective, and sequential delivery of multiple cargo molecules, including FITC-dextran and siRNA, to a broad range of cells. Notably, we achieved siRNA delivery into the cytoplasm of hard-to-transfect K562 cells with 45% efficiency, while maintaining nearly 100% cell viability.

cond-mat.mtrl-sci

Cross-subject Brain Functional Connectivity Analysis for Multi-task Cognitive State Evaluation

Cognition refers to the function of information perception and processing, which is the fundamental psychological essence of human beings. It is responsible for reasoning and decision-making, while its evaluation is significant for the aviation domain in mitigating potential safety risks. Existing studies tend to use varied methods for cognitive state evaluation yet have limitations in timeliness, generalisation, and interpretability. Accordingly, this study adopts brain functional connectivity with electroencephalography signals to capture associations in brain regions across multiple subjects for evaluating real-time cognitive states. Specifically, a virtual reality-based flight platform is constructed with multi-screen embedded. Three distinctive cognitive tasks are designed and each has three degrees of difficulty. Thirty subjects are acquired for analysis and evaluation. The results are interpreted through different perspectives, including inner-subject and cross-subject for task-wise and gender-wise underlying brain functional connectivity. Additionally, this study incorporates questionnaire-based, task performance-based, and physiological measure-based approaches to fairly label the trials. A multi-class cognitive state evaluation is further conducted with the active brain connections. Benchmarking results demonstrate that the identified brain regions have considerable influences in cognition, with a multi-class accuracy rate of 95.83% surpassing existing studies. The derived findings bring significance to understanding the dynamic relationships among human brain functional regions, cross-subject cognitive behaviours, and decision-making, which have promising practical application values.

cs.HC

Runtime Stealthy Perception Attacks against DNN-based Adaptive Cruise Control Systems

Adaptive Cruise Control (ACC) is a widely used driver assistance technology for maintaining the desired speed and safe distance to the leading vehicle. This paper evaluates the security of the deep neural network (DNN) based ACC systems under runtime stealthy perception attacks that strategically inject perturbations into camera data to cause forward collisions. We present a context-aware strategy for the selection of the most critical times for triggering the attacks and a novel optimization-based method for the adaptive generation of image perturbations at runtime. We evaluate the effectiveness of the proposed attack using an actual vehicle, a publicly available driving dataset, and a realistic simulation platform with the control software from a production ACC system, a physical-world driving simulator, and interventions by the human driver and safety features such as Advanced Emergency Braking System (AEBS). Experimental results show that the proposed attack achieves 142.9 times higher success rate in causing hazards and 82.6% higher evasion rate than baselines, while being stealthy and robust to real-world factors and dynamic changes in the environment. This study highlights the role of human drivers and basic safety mechanisms in preventing attacks.

cs.CR

Communicating Inferred Goals with Passive Augmented Reality and Active Haptic Feedback

Robots learn as they interact with humans. Consider a human teleoperating an assistive robot arm: as the human guides and corrects the arm's motion, the robot gathers information about the human's desired task. But how does the human know what their robot has inferred? Today's approaches often focus on conveying intent: for instance, upon legible motions or gestures to indicate what the robot is planning. However, closing the loop on robot inference requires more than just revealing the robot's current policy: the robot should also display the alternatives it thinks are likely, and prompt the human teacher when additional guidance is necessary. In this paper we propose a multimodal approach for communicating robot inference that combines both passive and active feedback. Specifically, we leverage information-rich augmented reality to passively visualize what the robot has inferred, and attention-grabbing haptic wristbands to actively prompt and direct the human's teaching. We apply our system to shared autonomy tasks where the robot must infer the human's goal in real-time. Within this context, we integrate passive and active modalities into a single algorithmic framework that determines when and which type of feedback to provide. Combining both passive and active feedback experimentally outperforms single modality baselines; during an in-person user study, we demonstrate that our integrated approach increases how efficiently humans teach the robot while simultaneously decreasing the amount of time humans spend interacting with the robot. Videos here: https://youtu.be/swq_u4iIP-g

cs.RO

Dark Energy Survey Year 1 Results: Constraining Baryonic Physics in the Universe

Measurements of large-scale structure are interpreted using theoretical predictions for the matter distribution, including potential impacts of baryonic physics. We constrain the feedback strength of baryons jointly with cosmology using weak lensing and galaxy clustering observables (3$\times$2pt) of Dark Energy Survey (DES) Year 1 data in combination with external information from baryon acoustic oscillations (BAO) and Planck cosmic microwave background polarization. Our baryon modeling is informed by a set of hydrodynamical simulations that span a variety of baryon scenarios; we span this space via a Principal Component (PC) analysis of the summary statistics extracted from these simulations. We show that at the level of DES Y1 constraining power, one PC is sufficient to describe the variation of baryonic effects in the observables, and the first PC amplitude ($Q_1$) generally reflects the strength of baryon feedback. With the upper limit of $Q_1$ prior being bound by the Illustris feedback scenarios, we reach $\sim 20\%$ improvement in the constraint of $S_8=σ_8(Ω_{\rm m}/0.3)^{0.5}=0.788^{+0.018}_{-0.021}$ compared to the original DES 3$\times$2pt analysis. This gain is driven by the inclusion of small-scale cosmic shear information down to 2.5 arcmin, which was excluded in previous DES analyses that did not model baryonic physics. We obtain $S_8=0.781^{+0.014}_{-0.015}$ for the combined DES Y1+Planck EE+BAO analysis with a non-informative $Q_1$ prior. In terms of the baryon constraints, we measure $Q_1=1.14^{+2.20}_{-2.80}$ for DES Y1 only and $Q_1=1.42^{+1.63}_{-1.48}$ for DESY1+Planck EE+BAO, allowing us to exclude one of the most extreme AGN feedback hydrodynamical scenario at more than $2 σ$.

astro-ph.CO

Superconvergence of ultra-weak discontinuous Galerkin methods for the linear Schr\"odinger equation in one dimension

We analyze the superconvergence properties of ultra-weak discontinuous Galerkin (UWDG) methods with various choices of flux parameters for one-dimensional linear Schr\"odinger equation. In our previous work [10], stability and optimal convergence rate are established for a large class of flux parameters. Depending on the flux choices and if the polynomial degree $k$ is even or odd, in this paper, we prove $2k$ or $(2k-1)$-th order superconvergence rate for cell averages and numerical flux of the function, as well as $(2k-1)$ or $(2k-2)$-th order for numerical flux of the derivative. In addition, we prove superconvergence of $(k+2)$ or $(k+3)$-th order of the DG solution towards a special projection. At a class of special points, the function values and the first and second order derivatives of the DG solution are superconvergent with order $k+2, k+1, k$, respectively. The proof relies on the correction function techniques initiated in [8], and applied to [6] for direct DG (DDG) methods for diffusion problems. Compared with [6], Schr\"odinger equation poses unique challenges for superconvergence proof because of the lack of the dissipation mechanism from the equation. One major highlight of our proof is that we introduce specially chosen test functions in the error equation and show the superconvergence of the second derivative and jump across the cell interfaces of the difference between numerical solution and projected exact solution. This technique was originally proposed in [12] and is essential to elevate the convergence order for our analysis. Finally, by negative norm estimates, we apply the post-processing technique and show that the accuracy of our scheme can be enhanced to order $2k.$ Theoretical results are verified by numerical experiments.

math.NA

Sparse Grid Central Discontinuous Galerkin Method for Linear Hyperbolic Systems in High Dimensions

In this paper, we develop sparse grid central discontinuous Galerkin (CDG) scheme for linear hyperbolic systems with variable coefficients in high dimensions. The scheme combines the CDG framework with the sparse grid approach, with the aim of breaking the curse of dimensionality. A new hierarchical representation of piecewise polynomials on the dual mesh is introduced and analyzed, resulting in a sparse finite element space that can be used for non-periodic problems. Theoretical results, such as $L^2$ stability and error estimates are obtained for scalar problems. CFL conditions are studied numerically comparing discontinuous Galerkin (DG), CDG, sparse grid DG and sparse grid CDG methods. Numerical results including scalar linear equations, acoustic and elastic waves are provided.

math.NA

An Ultra-Weak Discontinuous Galerkin Method for Schr\"odinger Equation in One Dimension

In this paper, we develop an ultra-weak discontinuous Galerkin (DG) method to solve the one-dimensional nonlinear Schr\"odinger equation. Stability conditions and error estimates are derived for the scheme with a general class of numerical fluxes. The error estimates are based on detailed analysis of the projection operator associated with each individual flux choice. Depending on the parameters, we find out that in some cases, the projection can be defined element-wise, facilitating analysis. In most cases, the projection is global, and its analysis depends on the resulting $2\times2$ block-circulant matrix structures. For a large class of parameter choices, optimal $\textit{a priori}$ $L^2$ error estimates can be obtained. Numerical examples are provided verifying theoretical results.

math.NA