SearcharxivSearch

arXiv subjects

Yifeng Xu

Publications and source records attributed to Yifeng Xu.

At least 19 recordsLinked to original sources

Dual Variational Neural Network for the $p$-Laplace Problem

The reliable and accurate numerical approximation of the $p$-Laplacian is particularly challenging in the extreme regimes $p \to 1^{+}$ and $p \gg 1$, where the operator becomes either highly singular or strongly degenerate, often causing severe instability in standard numerical methods. To address these difficulties, we propose a novel deep learning based framework, termed the dual variational neural network, for $p$-Laplace problems. The approach is based on a mixed formulation and an $L^q$-based Helmholtz decomposition, which decouples the original problem into two convex subproblems: a linear Poisson problem for the irrotational component and an unconstrained minimization problem over divergence-free fields for the solenoidal component. Following the decomposition, we employ two neural networks using a gradient--curl representation to approximate the flux, and further establish an error analysis of the neural approximation. The analysis relies on fundamental vector inequalities together with tools from statistical learning theory. Numerical experiments demonstrate robust convergence of the proposed method in challenging settings, including the extreme cases $p \to 1^{+}$ and $p \gg 1$, as well as the $p(x)$-Laplace equation.

math.NA

Adaptive Crouzeix-Raviart finite elements for the first eigenpair of $p$-Laplacian

In this paper, we propose and analyze an adaptive Crouzeix-Raviart finite element method for computing the first Dirichlet eigenpair of the $p$-Laplacian problem. We prove that the sequence of error estimators produced by the adaptive algorithm has a vanishing limit and that, starting from a fine initial mesh, the relevant sequence of approximate eigenvalues converges to the first eigenvalue and the distance in a mesh-dependent broken norm between discrete eigenfunctions and the set composed of relevant continuous eigenfunctions also tends to zero. The analysis hinges on establishing a compactness property for Crouzeix-Raviart finite elements over a sequence of adaptively generated meshes, which represents key theoretical challenges and novelties. We present numerical results to illustrate the advantage of the proposed algorithm.

math.NA

ASC-SW: A Lightweight Atrous Strip Convolution Network for DLOs Segmentation on Edge mobile Robots

Detecting deformable linear objects (DLOs), such as floor cables, is essential for safe mobile robot navigation but remains challenging due to oblique viewpoints, thin structures, and limited edge-device resources. Existing DLO segmentation methods are primarily designed for manipulator platforms with fixed top-down views and often require heavy models, limiting their deployment on mobile robots. We formulate a cross-view DLO segmentation problem, where models trained on manipulator-view data must generalize to mobile robot perspectives. To address this, we propose ASC-SW, a lightweight and geometry-aware segmentation framework. The core network, ASCNet, introduces Atrous Strip Convolution, combining directional strip filtering with dilated receptive fields to enhance sensitivity to elongated structures at low computational cost. An Atrous Strip Convolution Spatial Pyramid Pooling module enables multi-scale anisotropic feature aggregation, while a temporal Sliding Window refinement suppresses viewpoint-induced false positives. Evaluated on real-world mobile robot data, ASC-SW achieves 74.1% mIoU at 261 FPS and remains deployable on edge devices.

cs.RO

Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Visual generation and understanding are two deeply interconnected aspects of human intelligence, yet they have been traditionally treated as separate tasks in machine learning. In this paper, we propose Jodi, a diffusion framework that unifies visual generation and understanding by jointly modeling the image domain and multiple label domains. Specifically, Jodi is built upon a linear diffusion transformer along with a role switch mechanism, which enables it to perform three particular types of tasks: (1) joint generation, where the model simultaneously generates images and multiple labels; (2) controllable generation, where images are generated conditioned on any combination of labels; and (3) image perception, where multiple labels can be predicted at once from a given image. Furthermore, we present the Joint-1.6M dataset, which contains 200,000 high-quality images collected from public sources, automatic labels for 7 visual domains, and LLM-generated captions. Extensive experiments demonstrate that Jodi excels in both generation and understanding tasks and exhibits strong extensibility to a wider range of visual domains. Code is available at https://github.com/VIPL-GENUN/Jodi.

cs.CV

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion

Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on distilling multiple terrain-specific teacher policies into a unified student policy. However, while such distillation captures basic locomotion primitives, it struggles to organically compose these skills to adapt to complex environments, resulting in poor generalization to novel composite terrains unseen during training. To overcome this, we present DreamPolicy, a unified framework that integrates offline data with a diffusion-based world model, enabling a single policy to master both known and unseen terrains. Central to our approach is a terrain-aware world model, driven by an autoregressive diffusion world model trained on aggregated rollouts from specialized policies. This model synthesizes physically plausible future trajectories, which serve as dynamic objectives for a conditioned policy, thereby bypassing manual reward engineering. Unlike distillation, our world model captures generalizable locomotion skills, allowing for robust zero-shot transfer to unseen composite terrains. DreamPolicy naturally scales with data availability. As the offline dataset expands, the diffusion world model continuously acquires richer skills. Experiments demonstrate that DreamPolicy outperforms the strongest baseline by up to 27\% on unseen terrains and 38\% on combined terrains. By unifying world model-based planning and policy learning, DreamPolicy breaks the "one task, one policy" bottleneck and establishes a scalable, data-driven paradigm for generalist humanoid control.

cs.RO

Convergence Analysis of an Adaptive Nonconforming FEM for Phase-Field Dependent Topology Optimization in Stokes Flow

In this work, we develop an adaptive nonconforming finite element algorithm for the numerical approximation of phase-field parameterized topology optimization governed by the Stokes system. We employ the conforming linear finite element space to approximate the phase field, and the nonconforming linear finite elements (Crouzeix-Raviart elements) and piecewise constants to approximate the velocity field and the pressure field, respectively. We establish the convergence of the adaptive method, i.e., the sequence of minimizers contains a subsequence that converges to a solution of the first-order optimality system, and the associated subsequence of discrete pressure fields also converges. The analysis relies crucially on a new discrete compactness result of nonconforming linear finite elements over a sequence of adaptively generated meshes. We present numerical results for several examples to illustrate the performance of the algorithm, including a comparison with the uniform refinement strategy.

math.NA

On the Crouzeix-Raviart Finite Element Approximation of Phase-Field Dependent Topology Optimization in Stokes Flow

In this work, we investigate a nonconforming finite element approximation of phase-field parameterized topology optimization governed by the Stokes flow. The phase field, the velocity field and the pressure field are approximated by conforming linear finite elements, nonconforming linear finite elements (Crouzeix-Raviart elements) and piecewise constants, respectively. When compared with the standard conforming counterpart, the nonconforming FEM can provide an approximation with fewer degrees of freedom, leading to improved computational efficiency. We establish the convergence of the resulting numerical scheme in the sense that the sequences of phase-field functions and discrete velocity fields contain subsequences that converge to a minimizing pair of the continuous problem in the $H^1$-norm and a mesh-dependent norm, respectively. We present extensive numerical results to illustrate the performance of the approach, including a comparison with the popular Taylor-Hood elements.

math.NA

Adaptive Approximations of Inclusions in a Semilinear Elliptic Problem Related to Cardiac Electrophysiology

In this work, we investigate the numerical reconstruction of inclusions in a semilinear elliptic equation arising in the mathematical modeling of cardiac ischemia. We propose an adaptive finite element method for the resulting constrained minimization problem that is relaxed by a phase-field approach. The \textit{a posteriori} error estimators of the adaptive algorithm consist of three components, i.e., the state variable, the adjoint variable and the complementary relation. Moreover, using tools from adaptive finite element analysis and nonlinear optimization, we establish the strong convergence for a subsequence of adaptively generated discrete solutions to a solution of the continuous optimality system. Several numerical examples are presented to illustrate the convergence and efficiency of the adaptive algorithm

math.NA

RGBSQGrasp: Inferring Local Superquadric Primitives from Single RGB Image for Graspability-Aware Bin Picking

Bin picking is a challenging robotic task due to occlusions and physical constraints that limit visual information for object recognition and grasping. Existing approaches often rely on known CAD models or prior object geometries, restricting generalization to novel or unknown objects. Other methods directly regress grasp poses from RGB-D data without object priors, but the inherent noise in depth sensing and the lack of object understanding make grasp synthesis and evaluation more difficult. Superquadrics (SQ) offer a compact, interpretable shape representation that captures the physical and graspability understanding of objects. However, recovering them from limited viewpoints is challenging, as existing methods rely on multiple perspectives for near-complete point cloud reconstruction, limiting their effectiveness in bin-picking. To address these challenges, we propose \textbf{RGBSQGrasp}, a grasping framework that leverages superquadric shape primitives and foundation metric depth estimation models to infer grasp poses from a monocular RGB camera -- eliminating the need for depth sensors. Our framework integrates a universal, cross-platform dataset generation pipeline, a foundation model-based object point cloud estimation module, a global-local superquadric fitting network, and an SQ-guided grasp pose sampling module. By integrating these components, RGBSQGrasp reliably infers grasp poses through geometric reasoning, enhancing grasp stability and adaptability to unseen objects. Real-world robotic experiments demonstrate a 92% grasp success rate, highlighting the effectiveness of RGBSQGrasp in packed bin-picking environments.

cs.RO

DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing

We introduce DRESS, a novel approach for generating stylized large language model (LLM) responses through representation editing. Existing methods like prompting and fine-tuning are either insufficient for complex style adaptation or computationally expensive, particularly in tasks like NPC creation or character role-playing. Our approach leverages the over-parameterized nature of LLMs to disentangle a style-relevant subspace within the model's representation space to conduct representation editing, ensuring a minimal impact on the original semantics. By applying adaptive editing strengths, we dynamically adjust the steering vectors in the style subspace to maintain both stylistic fidelity and semantic integrity. We develop two stylized QA benchmark datasets to validate the effectiveness of DRESS, and the results demonstrate significant improvements compared to baseline methods such as prompting and ITI. In short, DRESS is a lightweight, train-free solution for enhancing LLMs with flexible and effective style control, making it particularly useful for developing stylized conversational agents. Codes and benchmark datasets are available at https://github.com/ArthurLeoM/DRESS-LLM.

cs.CL

CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation

Recently, large-scale diffusion models have made impressive progress in text-to-image (T2I) generation. To further equip these T2I models with fine-grained spatial control, approaches like ControlNet introduce an extra network that learns to follow a condition image. However, for every single condition type, ControlNet requires independent training on millions of data pairs with hundreds of GPU hours, which is quite expensive and makes it challenging for ordinary users to explore and develop new types of conditions. To address this problem, we propose the CtrLoRA framework, which trains a Base ControlNet to learn the common knowledge of image-to-image generation from multiple base conditions, along with condition-specific LoRAs to capture distinct characteristics of each condition. Utilizing our pretrained Base ControlNet, users can easily adapt it to new conditions, requiring as few as 1,000 data pairs and less than one hour of single-GPU training to obtain satisfactory results in most scenarios. Moreover, our CtrLoRA reduces the learnable parameters by 90% compared to ControlNet, significantly lowering the threshold to distribute and deploy the model weights. Extensive experiments on various types of conditions demonstrate the efficiency and effectiveness of our method. Codes and model weights will be released at https://github.com/xyfJASON/ctrlora.

cs.CV

Adaptive Finite Element Method for a Nonlinear Helmholtz Equation with High Wave Number

A nonlinear Helmholtz (NLH) equation with high frequencies and corner singularities is discretized by the linear finite element method (FEM). After deriving some wave-number-explicit stability estimates and the singularity decomposition for the NLH problem, a priori stability and error estimates are established for the FEM on shape regular meshes including the case of locally refined meshes. Then a posteriori upper and lower bounds using a new residual-type error estimator, which is equivalent to the standard one, are derived for the FE solutions to the NLH problem. These a posteriori estimates have confirmed a significant fact that is also valid for the NLH problem, namely the residual-type estimator seriously underestimates the error of the FE solution in the preasymptotic regime, which was first observed by Babuška et al. [Int J Numer Methods Eng 40 (1997)] for a one-dimensional linear problem. Based on the new a posteriori error estimator, both the convergence and the quasi-optimality of the resulting adaptive finite element algorithm are proved the first time for the NLH problem, when the initial mesh size lying in the preasymptotic regime. Finally, numerical examples are presented to validate the theoretical findings and demonstrate that applying the continuous interior penalty (CIP) technique with appropriate penalty parameters can reduce the pollution errors efficiently. In particular, the nonlinear phenomenon of optical bistability with Gaussian incident waves is successfully simulated by the adaptive CIPFEM.

math.NA

An Adaptive Phase-Field Method for Structural Topology Optimization

In this work, we develop an adaptive algorithm for the efficient numerical solution of the minimum compliance problem in topology optimization. The algorithm employs the phase field approximation and continuous density field. The adaptive procedure is driven by two residual type a posteriori error estimators, one for the state variable and the other for the first-order optimality condition of the objective functional. The adaptive algorithm is provably convergent in the sense that the sequence of numerical approximations generated by the adaptive algorithm contains a subsequence convergent to a solution of the continuous first-order optimality system. We provide several numerical simulations to show the distinct features of the algorithm.

math.OC

Adaptive finite element approximations of the first eigenpair associated with $p$-Laplacian

In this paper, we propose an adaptive finite element method for computing the first eigenpair of the $p$-Laplacian problem. We prove that starting from a fine initial mesh our proposed adaptive algorithm produces a sequence of discrete first eigenvalues that converges to the first eigenvalue of the continuous problem and the distance between discrete eigenfunctions and the normalized eigenfunction set corresponding to the first eigenvalue in $W^{1,p}$-norm also tends to zero. Extensive numerical examples are provided to show the effectiveness and efficiency.

math.NA

Adaptive Computation of Elliptic Eigenvalue Topology Optimization with a Phase-Field Approach

In this paper, we discuss adaptive approximations of an elliptic eigenvalue optimization problem in a phase-field setting by a conforming finite element method. An adaptive algorithm is proposed and implemented in several two dimensional numerical examples for illustration of efficiency and accuracy. Theoretical findings consist in the vanishing limit of a subsequence of estimators and the convergence of the relevant subsequence of adaptively-generated solutions to a solution to the continuous optimality system.

math.NA

Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption Detection

Effective image restoration with large-size corruptions, such as blind image inpainting, entails precise detection of corruption region masks which remains extremely challenging due to diverse shapes and patterns of corruptions. In this work, we present a novel method for automatic corruption detection, which allows for blind corruption restoration without known corruption masks. Specifically, we develop a hierarchical contrastive learning framework to detect corrupted regions by capturing the intrinsic semantic distinctions between corrupted and uncorrupted regions. In particular, our model detects the corrupted mask in a coarse-to-fine manner by first predicting a coarse mask by contrastive learning in low-resolution feature space and then refines the uncertain area of the mask by high-resolution contrastive learning. A specialized hierarchical interaction mechanism is designed to facilitate the knowledge propagation of contrastive learning in different scales, boosting the modeling performance substantially. The detected multi-scale corruption masks are then leveraged to guide the corruption restoration. Detecting corrupted regions by learning the contrastive distinctions rather than the semantic patterns of corruptions, our model has well generalization ability across different corruption patterns. Extensive experiments demonstrate following merits of our model: 1) the superior performance over other methods on both corruption detection and various image restoration tasks including blind inpainting and watermark removal, and 2) strong generalization across different corruption patterns such as graffiti, random noise or other image content. Codes and trained weights are available at https://github.com/xyfJASON/HCL .

cs.CV

Fed-TDA: Federated Tabular Data Augmentation on Non-IID Data

Non-independent and identically distributed (non-IID) data is a key challenge in federated learning (FL), which usually hampers the optimization convergence and the performance of FL. Existing data augmentation methods based on federated generative models or raw data sharing strategies for solving the non-IID problem still suffer from low performance, privacy protection concerns, and high communication overhead in decentralized tabular data. To tackle these challenges, we propose a federated tabular data augmentation method, named Fed-TDA. The core idea of Fed-TDA is to synthesize tabular data for data augmentation using some simple statistics (e.g., distributions of each column and global covariance). Specifically, we propose the multimodal distribution transformation and inverse cumulative distribution mapping respectively synthesize continuous and discrete columns in tabular data from a noise according to the pre-learned statistics. Furthermore, we theoretically analyze that our Fed-TDA not only preserves data privacy but also maintains the distribution of the original data and the correlation between columns. Through extensive experiments on five real-world tabular datasets, we demonstrate the superiority of Fed-TDA over the state-of-the-art in test performance and communication efficiency.

cs.LG

Adaptive Reconstruction for Electrical Impedance Tomography with a Piecewise Constant Conductivity

In this work we propose and analyze a numerical method for electrical impedance tomography of recovering a piecewise constant conductivity from boundary voltage measurements. It is based on standard Tikhonov regularization with a Modica-Mortola penalty functional and adaptive mesh refinement using suitable a posteriori error estimators of residual type that involve the state, adjoint and variational inequality in the necessary optimality condition and a separate marking strategy. We prove the convergence of the adaptive algorithm in the following sense: the sequence of discrete solutions contains a subsequence convergent to a solution of the continuous necessary optimality system. Several numerical examples are presented to illustrate the convergence behavior of the algorithm.

math.NA