SearcharxivSearch

arXiv subjects

Wendong Zhang

Publications and source records attributed to Wendong Zhang.

At least 19 recordsLinked to original sources

DREAM Technical Report

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

cs.IR

MetaStrategy: Generative Ranking with Executable LLM Strategies

Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.

cs.IR

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality. While recent dense-feedback methods move beyond scalar supervision, their heatmap-centric representations still formulate diagnosis as pixel-field regression, making it difficult to localize variable-cardinality defects and bind semantic reasons to individual failures. To address this representation bottleneck, we propose Structured Defect Grounding (SDG), which casts T2I diagnosis as structured set prediction by modeling each defect as a (location, type, reason, importance) tuple. To make this formulation trainable and measurable, we introduce SDG-30K, a 30K-image dataset with box-grounded annotations across four modern T2I generators, together with a dedicated evaluation protocol, SDG-Eval. Building on this structured representation, we further present a diagnosis-to-alignment framework in which a Vision-Language Model (VLM) serves as the SDG detector, and BoxFlow-GRPO converts predicted defect sets into box-derived, importance-weighted spatial rewards for diffusion model alignment. Extensive experiments show that our SDG detector outperforms leading proprietary VLMs on structured defect grounding, while SDG-guided rewards consistently improve T2I alignment and support localized image refinement. These results establish SDG as a unified, instance-level interface for diagnosing, evaluating, and enhancing modern generative models.

cs.CV

Continual Visual Reinforcement Learning with A Life-Long World Model

Learning physical dynamics in a series of non-stationary environments is a challenging but essential task for model-based reinforcement learning (MBRL) with visual inputs. It requires the agent to consistently adapt to novel tasks without forgetting previous knowledge. In this paper, we present a new continual learning approach for visual dynamics modeling and explore its efficacy in visual control. The key assumption is that an ideal world model can provide a non-forgetting environment simulator, which enables the agent to optimize the policy in a multi-task learning manner based on the imagined trajectories from the world model. To this end, we first introduce the life-long world model, which learns task-specific latent dynamics using a mixture of Gaussians and incorporates generative experience replay to mitigate catastrophic forgetting. Then, we further address the value estimation challenge for previous tasks with the exploratory-conservative behavior learning approach. Our model remarkably outperforms the straightforward combinations of existing continual learning and visual RL algorithms on DeepMind Control Suite and Meta-World benchmarks with continual visual control tasks.

cs.LG

FPGA: Flexible Portrait Generation Approach

Portrait Fidelity Generation is a prominent research area in generative models.Current methods face challenges in generating full-body images with low-resolution faces, especially in multi-ID photo phenomenon.To tackle these issues, we propose a comprehensive system called FPGA and construct a million-level multi-modal dataset IDZoom for training.FPGA consists of Multi-Mode Fusion training strategy (MMF) and DDIM Inversion based ID Restoration inference framework (DIIR). The MMF aims to activate the specified ID in the specified facial region. The DIIR aims to address the issue of face artifacts while keeping the background.Furthermore, DIIR is plug-and-play and can be applied to any diffusion-based portrait generation method to enhance their performance. DIIR is also capable of performing face-swapping tasks and is applicable to stylized faces as well.To validate the effectiveness of FPGA, we conducted extensive comparative and ablation experiments. The experimental results demonstrate that FPGA has significant advantages in both subjective and objective metrics, and achieves controllable generation in multi-ID scenarios. In addition, we accelerate the inference speed to within 2.5 seconds on a single L20 graphics card mainly based on our well designed reparameterization method, RepControlNet.

cs.CV

Modeling and Simulation of 2D Transducers Based on Suspended Graphene-Based Heterostructures in Nanoelectromechanical Pressure Sensors

Graphene-based 2D heterostructures exhibit excellent mechanical and electrical properties, which are expected to exhibit better performances than graphene for nanoelectromechanical pressure sensors. Here, we built the pressure sensor models based on suspended heterostructures of graphene/h-BN, graphene/MoS2, and graphene/MoSe2 by using COMSOL Multiphysics finite element software. We found that suspended circular 2D membranes show the best sensitivity to pressures compared to rectangular and square ones. We simulated the deflections, strains, resonant frequencies, and Young's moduli of suspended graphene-based heterostructures under the conditions of different applied pressures and geometrical sizes, built-in tensions, and the number of atomic layers of 2D membranes. The Young's moduli of 2D heterostructures of graphene, graphene/h-BN, graphene/MoS2, and graphene/MoSe2 were estimated to be 1.001TPa, 921.08 GPa, 551.11 GPa, and 475.68 GPa, respectively. We also discuss the effect of highly asymmetric cavities on device performance. These results would contribute to the understanding of the mechanical properties of graphene-based heterostructures and would be helpful for the design and manufacture of high-performance NEMS pressure sensors.

cond-mat.mes-hall

Fast response and highly sensitive flexible humidity sensor based on nanocomposite film of MoS2 and graphene oxide

Graphene oxide (GO)-based humidity sensors are attracting widespread attention due to their high responsivity and low cost. However, GO-based humidity sensors generally suffer from slow response and recovery as well as poor stability,etc. Here, we reported a flexible resistive humidity sensor based on a MoS2/GO composite film that was fabricated by mixing different volumes of MoS2 and GO dispersions with adjustable volume ratios. The MoS2/GO composite film has been used as a sensing layer on screen-printed interdigital electrodes. The results show that the best device performance was achieved at a dispersion volume of 0.05 mL with the MoS2/GO volume ratio of 5:1, featuring high responsivity (~98%), fast response/recovery time (1.3/12.1 s), excellent stability and low cost. Further, the humidity sensor exhibits good linearity over a wide humidity range (33% RH-98% RH) at room temperature and can be fabricated easily and feasibly. The application of the humidity sensors we prepared in human respiration detection and human fingertip proximity detection has been demonstrated. These findings indicate the great potential of the composite of MoS2/GO in developing the next generation of high-performance humidity sensors.

cond-mat.mes-hall

Recent Advances in Graphene-Based Humidity Sensors with the Focus of Structural Design: A Review

The advent of the 5G era means that the concepts of robot, VR/AR, UAV, smart home, smart healthcare based on IoT (Internet of Things) have gradually entered human life. Since then, intelligent life has become the dominant direction of social development. Humidity sensors, as humidity detection tools, not only convey the comfort of human living environment, but also display great significance in the fields of meteorology, medicine, agriculture and industry. Graphene-based materials exhibit tremendous potential in humidity sensing owing to their ultra-high specific surface area and excellent electron mobility under room temperature for application in humidity sensing. This review begins with the introduction of examples of various synthesis strategies of graphene, followed by the device structure and working mechanism of graphene-based humidity sensor. In addition, several different structural design methods of graphene are summarized, demonstrating the structural design of graphene can not only optimize the performance of graphene, but also bring significant advantages in humidity sensing. Finally, key challenges hindering the further development and practical application of high-performance graphene-based humidity sensors are discussed, followed by presenting the future perspectives.

cond-mat.mes-hall

Recent Advances in Graphene-Based Pressure Sensors: A Review

In recent years, pressure sensors have been widely used as crucial technology components in industrial, healthcare, consumer electronics, and automotive safety applications. With the development of intelligent technologies, there is a growing demand for pressure sensors with higher sensitivity, smaller size, and wider detection range. Graphene and its derivatives, as novel emerging materials in recent years, have received widespread attention from researchers due to their unique mechanical and electrical properties, and are considered as promising sensing materials for the high-performance pressure sensors. In general, graphene-based pressure sensors can be classified into flexible pressure sensors and gas pressure sensors. In this paper, we firstly introduce the basic properties of graphene and its derivatives and then review the research progress of both graphene-based flexible pressure sensors and graphene-based gas pressure sensors respectively, focusing on different sensing mechanisms. Finally, the application prospects of graphene-based pressure sensors as well as future challenges are discussed.

cond-mat.mes-hall

Graphene MEMS and NEMS

Graphene is being increasingly used as an interesting transducer membrane in micro- and nanoelectromechanical systems (MEMS and NEMS, respectively) due to its atomical thickness, extremely high carrier mobility, high mechanical strength and piezoresistive electromechanical transductions. NEMS devices based on graphene feature increased sensitivity, reduced size, and new functionalities. In this review, we discuss the merits of graphene as a functional material for MEMS and NEMS, the related properties of graphene, the transduction mechanisms of graphene MEMS and NEMS, typical transfer methods for integrating graphene with MEMS substrates, methods for fabricating suspended graphene, and graphene patterning and electrical contact. Consequently, we provide an overview of devices based on suspended and nonsuspended graphene structures. Finally, we discuss the potential and challenges of applications of graphene in MEMS and NEMS. Owing to its unique features, graphene is a promising material for emerging MEMS, NEMS and sensor applications.

cond-mat.mes-hall

Humidity Sensing Properties of Different Atomic Layers of Graphene on SiO2/Si Substrate

Graphene has the great potential to be used for humidity sensing due to ultrahigh surface area and conductivity. However, the impact of different atomic layers of graphene on SiO2/Si substrate on the humidity sensing have not been studied yet. In this paper, we fabricated three types of humidity sensors on SiO2/Si substrate based on one to three atomic layers of graphene, in which the sensing areas of graphene are 75 μm * 72 μm and 45 μm * 72 μm, respectively. We studied the impact of both the number of atomic layers of graphene and the sensing areas of graphene on the responsivity and response/recovery time of the prepared graphene-based humidity sensors. We found the relative resistance change of the prepared devices decreased with the increase of number of atomic layers of graphene under the same change of relative humidity. Further, devices based on tri-layer graphene showed the fastest response/recovery time while devices based on double-layer graphene showed the slowest response/recovery time. Finally, we chose the devices based on double-layer graphene that have relatively good responsivity and stability for application in respiration monitoring and contact-free finger monitoring.

cond-mat.mes-hall

Four ribbons of double-layer graphene suspending masses for NEMS applications

Graphene ribbons with a suspended proof mass for nanomechanical systems have been rarely studied. Here, we report three types of nanomechanical devices consisting of graphene ribbons (two ribbons, four ribbons-cross and four ribbons-parallel) with suspended Si proof masses and studied their mechanical properties. The resonance frequencies and built-in stresses of three types of devices ranged from tens of kHz to hundreds of kHz, and from 82.61 MPa to 545.73 MPa, respectively, both of which decrease with the increase of the size of proof mass. The devices with four graphene ribbons featured higher resonance frequencies and spring constants, but lower built-in stresses than two ribbon devices under otherwise identical conditions. The Young's modulus and fracture strain of double-layer graphene were measured to be 0.34 TPa and 1.13% respectively, by using the experimental data and finite element analysis (FEA) simulations. Our studies would lay the foundation for understanding of mechanical properties of graphene ribbons with a suspended proof mass and their potential applications in nanoelectromechanical systems.

cond-mat.mes-hall

Improving Masked Autoencoders by Learning Where to Mask

Masked image modeling is a promising self-supervised learning method for visual data. It is typically built upon image patches with random masks, which largely ignores the variation of information density between them. The question is: Is there a better masking strategy than random sampling and how can we learn it? We empirically study this problem and initially find that introducing object-centric priors in mask sampling can significantly improve the learned representations. Inspired by this observation, we present AutoMAE, a fully differentiable framework that uses Gumbel-Softmax to interlink an adversarially-trained mask generator and a mask-guided image modeling process. In this way, our approach can adaptively find patches with higher information density for different images, and further strike a balance between the information gain obtained from image reconstruction and its practical training difficulty. In our experiments, AutoMAE is shown to provide effective pretraining models on standard self-supervised benchmarks and downstream tasks.

cs.CV

Fully Context-Aware Image Inpainting with a Learned Semantic Pyramid

Restoring reasonable and realistic content for arbitrary missing regions in images is an important yet challenging task. Although recent image inpainting models have made significant progress in generating vivid visual details, they can still lead to texture blurring or structural distortions due to contextual ambiguity when dealing with more complex scenes. To address this issue, we propose the Semantic Pyramid Network (SPN) motivated by the idea that learning multi-scale semantic priors from specific pretext tasks can greatly benefit the recovery of locally missing content in images. SPN consists of two components. First, it distills semantic priors from a pretext model into a multi-scale feature pyramid, achieving a consistent understanding of the global context and local structures. Within the prior learner, we present an optional module for variational inference to realize probabilistic image inpainting driven by various learned priors. The second component of SPN is a fully context-aware image generator, which adaptively and progressively refines low-level visual representations at multiple scales with the (stochastic) prior pyramid. We train the prior learner and the image generator as a unified model without any post-processing. Our approach achieves the state of the art on multiple datasets, including Places2, Paris StreetView, CelebA, and CelebA-HQ, under both deterministic and probabilistic inpainting setups.

cs.CV

MMDR: A Result Feature Fusion Object Detection Approach for Autonomous System

Object detection has been extensively utilized in autonomous systems in recent years, encompassing both 2D and 3D object detection. Recent research in this field has primarily centered around multimodal approaches for addressing this issue.In this paper, a multimodal fusion approach based on result feature-level fusion is proposed. This method utilizes the outcome features generated from single modality sources, and fuses them for downstream tasks.Based on this method, a new post-fusing network is proposed for multimodal object detection, which leverages the single modality outcomes as features. The proposed approach, called Multi-Modal Detector based on Result features (MMDR), is designed to work for both 2D and 3D object detection tasks. Compared to previous multimodal models, the proposed approach in this paper performs feature fusion at a later stage, enabling better representation of the deep-level features of single modality sources. Additionally, the MMDR model incorporates shallow global features during the feature fusion stage, endowing the model with the ability to perceive background information and the overall input, thereby avoiding issues such as missed detections.

cs.CV

Obstacle-Transformer: A Trajectory Prediction Network Based on Surrounding Trajectories

Recurrent Neural Network, Long Short-Term Memory, and Transformer have made great progress in predicting the trajectories of moving objects. Although the trajectory element with the surrounding scene features has been merged to improve performance, there still exist some problems to be solved. One is that the time series processing models will increase the inference time with the increase of the number of prediction sequences. Another lies in which the features can not be extracted from the scene's image and point cloud in some situations. Therefore, this paper proposes an Obstacle-Transformer to predict trajectory in a constant inference time. An ``obstacle'' is designed by the surrounding trajectory rather than images or point clouds, making Obstacle-Transformer more applicable in a wider range of scenarios. Experiments are conducted on ETH and UCY data sets to verify the performance of our model.

cs.CV

HiFuse: Hierarchical Multi-Scale Feature Fusion Network for Medical Image Classification

Medical image classification has developed rapidly under the impetus of the convolutional neural network (CNN). Due to the fixed size of the receptive field of the convolution kernel, it is difficult to capture the global features of medical images. Although the self-attention-based Transformer can model long-range dependencies, it has high computational complexity and lacks local inductive bias. Much research has demonstrated that global and local features are crucial for image classification. However, medical images have a lot of noisy, scattered features, intra-class variation, and inter-class similarities. This paper proposes a three-branch hierarchical multi-scale feature fusion network structure termed as HiFuse for medical image classification as a new method. It can fuse the advantages of Transformer and CNN from multi-scale hierarchies without destroying the respective modeling so as to improve the classification accuracy of various medical images. A parallel hierarchy of local and global feature blocks is designed to efficiently extract local features and global representations at various semantic scales, with the flexibility to model at different scales and linear computational complexity relevant to image size. Moreover, an adaptive hierarchical feature fusion block (HFF block) is designed to utilize the features obtained at different hierarchical levels comprehensively. The HFF block contains spatial attention, channel attention, residual inverted MLP, and shortcut to adaptively fuse semantic information between various scale features of each branch. The accuracy of our proposed model on the ISIC2018 dataset is 7.6% higher than baseline, 21.5% on the Covid-19 dataset, and 10.4% on the Kvasir dataset. Compared with other advanced models, the HiFuse model performs the best. Our code is open-source and available from https://github.com/huoxiangzuo/HiFuse.

eess.IV

Continual Predictive Learning from Videos

Predictive learning ideally builds the world model of physical processes in one or more given environments. Typical setups assume that we can collect data from all environments at all times. In practice, however, different prediction tasks may arrive sequentially so that the environments may change persistently throughout the training procedure. Can we develop predictive learning algorithms that can deal with more realistic, non-stationary physical environments? In this paper, we study a new continual learning problem in the context of video prediction, and observe that most existing methods suffer from severe catastrophic forgetting in this setup. To tackle this problem, we propose the continual predictive learning (CPL) approach, which learns a mixture world model via predictive experience replay and performs test-time adaptation with non-parametric task inference. We construct two new benchmarks based on RoboNet and KTH, in which different tasks correspond to different physical robotic environments or human actions. Our approach is shown to effectively mitigate forgetting and remarkably outperform the naïve combinations of previous art in video prediction and continual learning.

cs.CV