SearcharxivSearch

arXiv subjects

Yuqing Chen

Publications and source records attributed to Yuqing Chen.

At least 19 recordsLinked to original sources

DualEraser: Joint Video Object and Effect Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver

Video object removal frequently struggles to eliminate target objects and their associated complex physical effects (e.g., smoke and light) in real-world scenes. We attribute this challenge to a fundamental semantic--pixel conflict, which manifests at two aspects: condition-level modality dissonance and optimization-level objective entanglement. In terms of conditioning, modality dissonance emerges from single-modality information incompleteness and cross-modal dominance imbalance. During optimization, two conflicting objectives---high-level semantic erasure and pixel-level background preservation---are inextricably entangled within a single model. To address these conflicts, we propose DualEraser, a novel framework for joint video object and effect removal. First, a Bipartite Text prompt and a Multi-Conditional Capability Elicitation (MCCE) mechanism explicitly inject effect semantics and further leverage multimodal priors to address the limitations of individual modalities. Second, a Learnable Deep CFG Fusion (LD-CFG) module adaptively balances the relative dominance between the text and mask conditions. Finally, we introduce a decoupled expert architecture comprising a Locator for semantic erasure and a Preserver for background pixel alignment to break the objective entanglement. Extensive experiments demonstrate that DualEraser achieves state-of-the-art quantitative performance on standard benchmarks (e.g., gains of 2.16 dB and 1.44 dB on ROSE and VOR-Eval, respectively), while enabling robust removal of complex effects in open-world videos. https://cyqii.github.io/DualEraser.github.io/

cs.CV

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing

Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing tasks, which complicates model design and demands significant training resources. To address this, we propose O-DisCo-Edit, a unified framework that incorporates a novel object distortion control (O-DisCo). This signal, based on random and adaptive noise, flexibly encapsulates a wide range of editing cues within a single representation. Paired with a "copy-form" preservation module for preserving non-edited regions, O-DisCo-Edit enables efficient, high-fidelity editing through an effective training paradigm. Extensive experiments and comprehensive human evaluations consistently demonstrate that O-DisCo-Edit surpasses both specialized and multitask state-of-the-art methods across various video editing tasks. https://cyqii.github.io/O-DisCo-Edit.github.io/

cs.CV

Design of an all-facet illuminator for high NA EUV lithography exposure tool based on deep reinforcement learning

Using the illuminator for high numerical aperture (NA) extreme ultraviolet (EUV) exposure tool in EUV lithography can lead to support volume production of sub-2 nm logic nodes and leading-edge DRAM nodes. However, the typical design method of the illuminator has issues with the transmission owing to the limitation of optical structure that cannot further reduce process parameter k1, and uniformity due to the restriction of matching method that can only consider one factor affecting uniformity. The all-facet illuminator can improve transmission by removing relay system. Deep reinforcement learning (RL) can improve the uniformity by considering multiple factors. In this paper, a design method of the all-facet illuminator for high NA EUV lithography exposure tool and a matching method based on deep RL for the double facets are proposed. The all-facet illuminator is designed using matrix optics, and removing relay system to achieve high transmission. The double facets is matched using the deep RL framework, which includes the policy network with improved trainability and low computational demands, and the reward function with great optimization direction and fast convergence rate, enabling to rapidly generate multiple matching results with high uniformity. An all-facet illuminator for a 0.55 NA EUV lithography exposure tool is designed by the proposed method. Simulation results indicate that the transmission is greater than 35%, and uniformity exceed 99% under multiple illumination pupil shapes.

physics.optics

Controllable Segmentation-Based Text-Guided Style Editing

We present a novel approach for controllable, region-specific style editing driven by textual prompts. Building upon the state-space style alignment framework introduced by \emph{StyleMamba}, our method integrates a semantic segmentation model into the style transfer pipeline. This allows users to selectively apply text-driven style changes to specific segments (e.g., ``turn the building into a cyberpunk tower'') while leaving other regions (e.g., ``people'' or ``trees'') unchanged. By incorporating region-wise condition vectors and a region-specific directional loss, our method achieves high-fidelity transformations that respect both semantic boundaries and user-driven style descriptions. Extensive experiments demonstrate that our approach can flexibly handle complex scene stylizations in real-world scenarios, improving control and quality over purely global style transfer methods.

cs.GR

Integrated Hierarchical Decision-Making in Inverse Kinematic Planning and Control

This work presents a novel and efficient nonlinear programming framework that tightly integrates hierarchical decision-making with whole-body inverse kinematic planning and control. Decision-making plays a central role in many aspects of robotics, from sparse inverse kinematic control with a minimal number of joints, to inverse kinematic planning while simultaneously selecting a discrete end-effector location from multiple candidates. Current approaches often rely on heavy computations using mixed-integer nonlinear programming, separate decision-making from inverse kinematics (some times approximated by reachability methods), or employ efficient but less versatile $\ell_1$-norm formulations of linear sparse programming, without addressing the underlying nonlinear problem formulations. In contrast, the proposed sparse hierarchical nonlinear programming solver is efficient, versatile, and accurate by exploiting sparse hierarchical structure and leveraging the $\ell_0$-norm which is rarely used in robotics. The solver efficiently tackles complex nonlinear hierarchical decision-making problems previously unaddressed in the literature, such as inverse kinematic planning with simultaneous prioritized selection of end-effector locations from a large set of candidates, or inverse kinematic control with simultaneous selection of bi-manual grasp locations on a randomly rotated box.

cs.RO

VMGNet: A Low Computational Complexity Robotic Grasping Network Based on VMamba with Multi-Scale Feature Fusion

While deep learning-based robotic grasping technology has demonstrated strong adaptability, its computational complexity has also significantly increased, making it unsuitable for scenarios with high real-time requirements. Therefore, we propose a low computational complexity and high accuracy model named VMGNet for robotic grasping. For the first time, we introduce the Visual State Space into the robotic grasping field to achieve linear computational complexity, thereby greatly reducing the model's computational cost. Meanwhile, to improve the accuracy of the model, we propose an efficient and lightweight multi-scale feature fusion module, named Fusion Bridge Module, to extract and fuse information at different scales. We also present a new loss function calculation method to enhance the importance differences between subtasks, improving the model's fitting ability. Experiments show that VMGNet has only 8.7G Floating Point Operations and an inference time of 8.1 ms on our devices. VMGNet also achieved state-of-the-art performance on the Cornell and Jacquard public datasets. To validate VMGNet's effectiveness in practical applications, we conducted real grasping experiments in multi-object scenarios, and VMGNet achieved an excellent performance with a 94.4% success rate in real-world grasping tasks. The video for the real-world robotic grasping experiments is available at https://youtu.be/S-QHBtbmLc4.

cs.RO

High-performance automated abstract screening with large language model ensembles

Large language models (LLMs) excel in tasks requiring processing and interpretation of input text. Abstract screening is a labour-intensive component of systematic review involving repetitive application of inclusion and exclusion criteria on a large volume of studies identified by a literature search. Here, LLMs (GPT-3.5 Turbo, GPT-4 Turbo, GPT-4o, Llama 3 70B, Gemini 1.5 Pro, and Claude Sonnet 3.5) were trialled on systematic reviews in a full issue of the Cochrane Library to evaluate their accuracy in zero-shot binary classification for abstract screening. Trials over a subset of 800 records identified optimal prompting strategies and demonstrated superior performance of LLMs to human researchers in terms of sensitivity (LLM-max = 1.000, human-max = 0.775), precision (LLM-max = 0.927, human-max = 0.911), and balanced accuracy (LLM-max = 0.904, human-max = 0.865). The best performing LLM-prompt combinations were trialled across every replicated search result (n = 119,691), and exhibited consistent sensitivity (range 0.756-1.000) but diminished precision (range 0.004-0.096). 66 LLM-human and LLM-LLM ensembles exhibited perfect sensitivity with a maximal precision of 0.458, with less observed performance drop in larger trials. Significant variation in performance was observed between reviews, highlighting the importance of domain-specific validation before deployment. LLMs may reduce the human labour cost of systematic review with maintained or improved accuracy and sensitivity. Systematic review is the foundation of evidence synthesis across academic disciplines, including evidence-based medicine, and LLMs may increase the efficiency and quality of this mode of research.

cs.CL

When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction

Recent advances in large language models (LLMs) have enabled intelligent tutoring systems, yet the development of LLM-based Virtual Student Agents (LVSAs) remains underexplored. Such agents are essential for teacher-facing applications, where simulating diverse learner traits can support adaptive instruction and pedagogical skill development. However, current methods lack principled personality modeling, scalable evaluation of behavioral consistency, and empirical validation in interactive teaching settings. We propose the SOEI framework, a structured pipeline comprising Scene, Object, Evaluation, and Interaction, for constructing and evaluating personality-aligned LVSAs in classroom scenarios. Leveraging Chinese language instruction as a cognitively and emotionally rich testbed, we generate five LVSAs based on Big Five traits through LoRA fine-tuning and expert-informed prompt design. Their behavioral realism and personality coherence are assessed using a hybrid human & GPT-4 evaluation and a multi-dimensional annotation protocol. Through controlled experiments with real pre-service teachers, we demonstrate that LVSAs can elicit adaptive teaching strategies and maintain trait-consistent behavior across multi-turn dialogues. Our results provide: (1) an educationally and psychologically grounded generation pipeline for LLM-based student agents; (2) a hybrid, scalable evaluation framework for behavioral realism; and (3) empirical insights into the pedagogical utility of LVSAs in shaping instructional adaptation. By embedding LVSAs into both generative modeling and human-in-the-loop teaching, SOEI bridges AI for Education (AI4Edu) and Education for AI (Edu4AI), positioning classroom interaction as a rigorous testbed for controllability, personality alignment, and human-likeness in large language models.

cs.CV

Visual Harmony: Text-Visual Interplay in Circular Infographics

Infographics are visual representations designed for efficient and effective communication of data and knowledge. One crucial aspect of infographic design is the interplay between text and visual elements, particularly in circular visualizations where the textual descriptions can either be embedded within the graphics or placed adjacent to the visual representation. While several studies have examined text layout design in visualizations in general, the text-visual interplay in infographics and its subsequent perceptual effects remain underexplored. To address this, our study investigates how varying text placement and descriptiveness impact pleasantness, comprehension and overall memorability in the infographics viewing experience. We recruited 30 participants and presented them with a collection of 15 infographics across a diverse set of topics, including media and public events, health and nutrition, science and research, and sustainability. The text placement (embed, side-to-side) and descriptiveness (simplistic, normal, descriptive) were systematically manipulated, resulting in a total of six experimental conditions. Our key findings indicate that text placement can significantly influence the memorability of infographics, whereas descriptiveness can significantly impact the pleasantness of the viewing experience. Embedding text placement and simplistic text can potentially contribute to more effective infographic designs. These results offer valuable insights for infographic designers, contributing to the creation of more effective and memorable visual representations.

cs.HC

To better understand realized ecosystem services: An integrated analysis framework of supply, demand, flow and use

Realized ecosystem services (ES) are the actual use of ES by societies, which is more directly linked to human well-being than potential ES. However, there is a lack of a general analysis framework to understand how much ES was realized. In this study, we first proposed a Supply-Demand-Flow-Use (SDFU) framework that integrates the supply, demand, flow, and use of ES and differentiates these concepts into different aspects (e.g., potential vs. actual ES demand, export and import flows of supply, etc.). Then, we applied the framework to three examples of ES that can be found in typical urban green parks (i.e., wild berry supply, pollination, and recreation). We showed how the framework could assess the actual use of ES and identify the supply-limited, demand-limited, and supply-demand-balanced types of realized ES. We also discussed the scaling features, temporal dynamics, and spatial characteristics of realized ES, as well as some critical questions for future studies. Although facing challenges, we believe that the applications of the SDFU framework can provide a systematic way to accurately assess the actual use of ES and better inform management and policy-making for sustainable use of nature's benefits. Therefore, we hope that our study will stimulate more research on realized ES and contribute to a deeper understanding of their roles in enhancing human well-being.

econ.GN

A Compact Variable Stiffness Actuator for Agile Legged Locomotion

The legged robots with variable stiffness actuators (VSAs) can achieve energy-efficient and versatile locomotion. However, equipping legged robots with VSAs in real-world application is usually restricted by (i) the redundant mechanical structure design, (ii) limited stiffness variation range and speed, (iii) high energy consumption in stiffness modulation, and (iv) the lack of online stiffness control structure with high performance. In this paper, we present a novel Variable-Length Leaf-Spring Actuator (VLLSA) designed for legged robots that aims to address the aforementioned limitations. The design is based on leaf-spring mechanism and we improve the structural design to make the proposed VSA (i) compact and lightweight in mechanical structure, (ii) precise in theoretical modeling, and (iii) capable of modulating stiffness with wide range, fast speed, low energy consumption and high control performance. Hardware experiments including in-place and forward hopping validate advantages of the proposed VLLSA.

cs.RO

Design and Control of a Bio-inspired Wheeled Bipedal Robot

Wheeled bipedal robots (WBRs) have the capability to execute agile and versatile locomotion tasks. This paper focuses on improving the dynamic performance of WBRs through innovations in both hardware and software development. Inspired by the human barbell squat, a bionic mechanical design is proposed and implemented as shown in Fig. 1. It distributes the weight onto its hip and knee joints to improve the effectiveness of joint motors while maintaining a relatively large workspace of the base link. Meanwhile, a novel model-based controller is devised, synthesizing height-variable wheeled linear inverted pendulum (HV-wLIP) model, Control Lyapunov Function (CLF) and whole-body dynamics for theoretically guaranteed stability and efficient computation. Compared with other alternatives, as a more accurate approximation of the WBR dynamics, the HV-wLIP can enable more agile response and provide theory basis for WBR controller design. Experimental results demonstrate that the robot could perform human-like deep squat, and is capable of maintaining tracking CoM velocity while manipulating base states. Furthermore, it exhibited robustness against external disturbances and unknown terrains even in the wild.

cs.RO

Seismic Inversion by Multi-dimensional Newtonian Machine

Newtonian machine learning (NML) is a wave-equation inversion method that inverts single-dimensional latent space (LS) features of the seismic data for retrieving the subsurface background velocity model. The single-dimensional LS features mainly contain the kinematic information of the seismic data, which are automatically extracted from the seismic signal by using an autoencoder network. Because its LS feature dimension is too small to preserve the dynamic information, such as the waveform variations, of the seismic data. Therefore the NML inversion is not able to recover the high-wavenumber velocity details. To mitigate this problem, we propose to invert multi-dimensional LS features, which can fully represent the entire characters of the seismic data. We denote this method as multi-dimensional Newtonian machine learning (MNML). In MNML, we define a new multi-variable connective function that works together with the multi-variable implicit function theorem to connect the velocity perturbations to the multi-dimensional LS feature perturbations. Numerical tests show that (1) the multi-dimensional LS features can preserve more data information than the single-dimensional LS features; (2) a higher resolution velocity model can be recovered by inverting the multi-dimensional LS features, and the inversion quality is comparable to that of FWI; (3) the MNML method requires a much smaller storage space than conventional FWI because only the low-dimensional representations of the high-dimensional seismic data are needed to be stored. The disadvantage of MNML is that it can more easily get stuck in local minima compared to the NML method. So we suggest a multiscale inversion approach that inverts for higher dimensional LS features as the iteration count increase.

physics.geo-ph

Seismic Inversion by Hybrid Machine Learning

We present a new seismic inversion method that uses deep learning (DL) features for the subsurface velocity model estimation. The DL feature is a low-dimensional representation of the high-dimensional seismic data, which is automatically generated by a convolutional autoencoder (CAE) and preserved in the latent space. The low-dimensional DL feature contains the key information of the input seismic data. Therefore, instead of directly comparing the waveform differences between the observed and predicted data, such as full-waveform inversion (FWI). We measure their DL feature differences in the latent space of a CAE. The advantage of this low-dimensional comparison is that it is less prone to the cycle-skipping problem compared to FWI. The reason is that the DL features mainly contain the kinematic information, such as traveltime, of the input seismic data when the latent space dimension is small. However, more dynamic information, such as the waveform variations, can be preserved in the DL feature when the latent space dimension becomes larger. Therefore we propose a multiscale inversion approach that starts with inverting the low-dimensional DL features for the low-wavenumber information of the subsurface model. Then recover its high-wavenumber details through inverting the high-dimensional DL features. However, there is no governing equation that contains both the velocity and DL feature terms in the same equation. Therefore we use the automatic differentiation (AD) to numerically connect the perturbation of DL features to the velocity perturbation. In another word, we connect a deep learning network with the wave-equation inversion by using the AD. We denote this hybrid connection as hybrid machine learning (HML) inversion. Here, the AD replaces the complex math derivations of the gradient with a black box so anyone can do HML without having a deep geophysical background.

physics.geo-ph

Deep Convolutional Neural Network and Sparse Least Squares Migration

We recast the forward pass of a multilayered convolutional neural network (CNN) as the solution to the problem of sparse least squares migration (LSM). The CNN filters and feature maps are shown to be analogous, but not equivalent, to the migration Green's functions and the quasi-reflectivity distribution, respectively. This provides a physical interpretation of the filters and feature maps in deep CNN in terms of the operators for seismic imaging. Motivated by the connection between sparse LSM and CNN, we propose the neural network version of sparse LSM. Unlike the standard LSM method that finds the optimal reflectivity image, neural network LSM (NNLSM) finds both the optimal quasi-reflectivity image and the quasi-migration Green's functions. These quasi-migration-Green's functions are also denoted as the convolutional filters in a CNN and are similar to migration Green's functions. The advantage of NNLSM over standard LSM is that its computational cost is significantly less and it can be used for denoising coherent and incoherent noise in migration images. Its disadvantage is that the NNLSM quasi-reflectivity image is only an approximation to the actual reflectivity distribution. However, the quasi-reflectivity image can be used as a superresolution attribute image for high-resolution delineation of geologic bodies.

physics.geo-ph

Core-Shell Nanofiber Containing Large Amount of Flame Retardants via Coaxial Dual-Nozzle Electrospinning as Battery Separators

Lithium-ion batteries have attracted enormous interests recently as promising power sources. However, the safety issue associated with the employment of highly flammable liquid electrolyte impedes the further development of next-generation lithium-ion batteries. Recently, researchers reported the use of electrospun core-shell fiber as the battery separator consisting of polymer layer as protective shell and flame retardants loaded inside as core. In case of a typical battery shorting, the protective polymer shell melts during thermal-runaway and the flame retardants inside would be released to suppress the combustion of the electrolyte. Due to the use of a single precursor solution for electrospinning containing both polymer and flame retardants, the weight ratio of flame retardants is limited and dependent. Herein, we developed a dual-nozzle, coaxial electrospinning approach to fabricate the core-shell nanofiber with a greatly enhanced flame retardants weight percentage in the final fibers. The weight ratio of flame retardants of triphenyl phosphate in the final composite reaches over 60 wt.%. The LiFePO4-based cell using this composite nanofiber as battery separator exhibits excellent flame-retardant property without compromising the cycling stability or rate performances. In addition, this functional nanofiber can also be coated onto commercial separators instead of being used directly as separators.

physics.app-ph

Seismic Inversion by Newtonian Machine Learning

We present a wave-equation inversion method that inverts skeletonized data for the subsurface velocity model. The skeletonized representation of the seismic traces consists of the low-rank latent-space variables predicted by a well-trained autoencoder neural network. The input to the autoencoder is the recorded common shot gathers, and the implicit function theorem is used to determine the perturbation of the skeletonized data with respect to the velocity perturbation. The final velocity model is the one that best predicts the observed latent-space parameters. Empirical results suggest that the cycle-skipping problem is largely mitigated compared to the conventional full waveform inversion (FWI) method by replacing the waveform differences by those of the latent-space parameters. The advantage of this method over other skeletonized data methods is that no manual picking of important features is required because the skeletal data are automatically selected by the autoencoder. The most significant contribution of this paper is that it provides a general framework for using solutions to the governing PDE to invert skeletal data generated by any type of a neural network. The governing equation can be that for gravity, seismic waves, electromagnetic fields, and magnetic fields. The input data can be the records from different types of data and their skeletal features, as long as the model parameters are sensitive to their perturbations. The skeletal data can be the latent space variables of an autoencoder, a variational autoencoder, or a feature map from a convolutional neural network (CNN), or principal component analysis (PCA) features. In other words, we have combined the best features of Newtonian physics and the pattern matching capabilities of machine learning to invert seismic data by Newtonian machine learning.

physics.geo-ph

The selection of LEGUE disk targets for LAMOST's pilot survey

We describe the target selection algorithm for the low latitude disk portion of the LAMOST Pilot Survey, which aims to test systems in preparation for the LAMOST spectroscopic survey. We use the PPMXL (Roeser et al. 2010) astrometric catalog, which provides positions, proper motions, B/R/I magnitudes (mostly) from USNO-B (Monet et al. 2003) and J/H/Ks from The Two Micron All Sky Survey (2MASS, see Skrutskie et al. 2006) as well. We chose 8 plates along the Galactic plane, in the region $0^\circ<α<67^\circ$ and $42^\circ<δ<59^\circ$, that cover 22 known open clusters with a range of ages. Adjacent plates may have small overlapping. Each plate covers an area $2.5^\circ$ in radius,with central star (for Shack-Hartmann guider) brighter than $\sim8^{\rm th}$ magnitude. For each plate, we create an input catalog in the magnitude range $11.3<Imag<16.3$ and $Bmag$ available from PPMXL. The stars are selected to satisfy the requirements of the fiber positioning system and have a uniform distribution in the $I$ vs. $B-I$ color-magnitude diagram. Our final input catalog consists of 12,000 objects on each of 8 plates that are observable during the winter observing season in Xinglong Station of the National Astronomical Observatory of China.

astro-ph.GA