SearcharxivSearch

arXiv subjects

Yifeng Zhou

Publications and source records attributed to Yifeng Zhou.

14 recordsLinked to original sources

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.

cs.CL

Asymmetric Generative Recommendation via Kronecker Residual Bridge and Multi-Faceted Hierarchical Quantization

Generative Recommendation (GenRec) models reformulate recommendation as a sequence generation task, representing items as discrete Semantic IDs used symmetrically as both inputs and prediction targets. We identify a critical dual-stage information bottleneck in this design: (1) the Input Bottleneck, where lossy quantization degrades fine-grained semantics, while popularity bias skews learned representations toward frequent items, and (2) the Output Bottleneck, where imprecise discrete targets limit supervision quality. To address these issues, we propose AsymRec, an asymmetric continuous-discrete framework that decouples input and output representations. Specifically, Kronecker Residual Bridge (KRB) maps continuous embeddings into the Transformer's hidden space via a Kronecker projection with a residual pathway, preserving semantic richness and improving generalization to infrequent items. Multi-faceted Hierarchical Quantization (MHQ) constructs high-capacity, structured discrete targets through multi-view and multi-level quantization with semantic regularization, preventing dimensional collapse while retaining fine-grained distinctions. Extensive experiments demonstrate that AsymRec consistently outperforms state-of-the-art generative recommenders by an average of 18.7%. Our project page is available at https://github.com/huangb23/AsymRec.

cs.IR

TokenFormer: Unify the Multi-Field and Sequential Recommendation Worlds

Recommender systems have historically developed along two largely independent paradigms: feature interaction models for modeling correlations among multi-field categorical features, and sequential models for capturing user behavior dynamics from historical interaction sequences. Although recent trends attempt to bridge these paradigms within shared backbones, we empirically reveal that naive unifying these two branches may lead to a failure mode of Sequential Collapse Propagation (SCP). That is, the interaction with those dimensionally ill non-sequence fields leads to the dimensional collapse of the sequence features. To overcome this challenge, we propose TokenFormer, a unified recommendation architecture with the following innovations. First, we introduce a Bottom-Full-Top-Sliding (BFTS) attention scheme, which applies full self-attention in the lower layers and shrinking-window sliding attention in the upper layers. Second, we introduce a Non-Linear Interaction Representation (NLIR) that applies one-sided non-linear multiplicative transformations to the hidden states. Extensive experiments on public benchmarks and Tencent's advertising platform demonstrate state-of-the-art performance, while detailed analysis confirm that TokenFormer significantly improves dimensional robustness and representation discriminability under unified modeling.

cs.IR

EvolvingGS: High-Fidelity Streamable Volumetric Video via Evolving 3D Gaussian Representation

We have recently seen great progress in 3D scene reconstruction through explicit point-based 3D Gaussian Splatting (3DGS), notable for its high quality and fast rendering speed. However, reconstructing dynamic scenes such as complex human performances with long durations remains challenging. Prior efforts fall short of modeling a long-term sequence with drastic motions, frequent topology changes or interactions with props, and resort to segmenting the whole sequence into groups of frames that are processed independently, which undermines temporal stability and thereby leads to an unpleasant viewing experience and inefficient storage footprint. In view of this, we introduce EvolvingGS, a two-stage strategy that first deforms the Gaussian model to coarsely align with the target frame, and then refines it with minimal point addition/subtraction, particularly in fast-changing areas. Owing to the flexibility of the incrementally evolving representation, our method outperforms existing approaches in terms of both per-frame and temporal quality metrics while maintaining fast rendering through its purely explicit representation. Moreover, by exploiting temporal coherence between successive frames, we propose a simple yet effective compression algorithm that achieves over 50x compression rate. Extensive experiments on both public benchmarks and challenging custom datasets demonstrate that our method significantly advances the state-of-the-art in dynamic scene reconstruction, particularly for extended sequences with complex human performances.

cs.CV

MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLLMs' ability in industrial anomaly detection has not been systematically studied. To bridge this gap, we present MMAD, the first-ever full-spectrum MLLMs benchmark in industrial Anomaly Detection. We defined seven key subtasks of MLLMs in industrial inspection and designed a novel pipeline to generate the MMAD dataset with 39,672 questions for 8,366 industrial images. With MMAD, we have conducted a comprehensive, quantitative evaluation of various state-of-the-art MLLMs. The commercial models performed the best, with the average accuracy of GPT-4o models reaching 74.9%. However, this result falls far short of industrial requirements. Our analysis reveals that current MLLMs still have significant room for improvement in answering questions related to industrial anomalies and defects. We further explore two training-free performance enhancement strategies to help models improve in industrial scenarios, highlighting their promising potential for future research.

cs.AI

CAR: Controllable Autoregressive Modeling for Visual Generation

Controllable generation, which enables fine-grained control over generated outputs, has emerged as a critical focus in visual generative models. Currently, there are two primary technical approaches in visual generation: diffusion models and autoregressive models. Diffusion models, as exemplified by ControlNet and T2I-Adapter, offer advanced control mechanisms, whereas autoregressive models, despite showcasing impressive generative quality and scalability, remain underexplored in terms of controllability and flexibility. In this study, we introduce Controllable AutoRegressive Modeling (CAR), a novel, plug-and-play framework that integrates conditional control into multi-scale latent variable modeling, enabling efficient control generation within a pre-trained visual autoregressive model. CAR progressively refines and captures control representations, which are injected into each autoregressive step of the pre-trained model to guide the generation process. Our approach demonstrates excellent controllability across various types of conditions and delivers higher image quality compared to previous methods. Additionally, CAR achieves robust generalization with significantly fewer training resources compared to those required for pre-training the model. To the best of our knowledge, we are the first to propose a control framework for pre-trained autoregressive visual generation models.

cs.CV

Einasto profile as the halo model solution coupled to the depletion radius

We constrain the halo profiles outside the halo boundaries by solving for the matching profiles required by the halo model. In the halo model framework, the matter distribution in the universe can be decomposed into the spatial distribution of halos convolved with their internal structures. This leads to a set of linear equations in Fourier space which uniquely determines the matching halo profiles for any given halo catalog. In this work, we construct three halo catalogs with different boundary definitions, and solve for the matching profiles in each case using measurements of halo-matter and halo-halo power spectra. Our results show that for a given halo field, there is always a set of matching profiles to accurately reconstruct the input statistics of the matter field, even though it might be complex to model the profiles analytically. Comparing the solutions from different halo catalogs, we find their mass distributions inside the inner depletion radii are nearly identical, while they deviate from each other on larger scales, with a larger boundary resulting in a more extended profile. For the depletion radius based catalog, the numerical solution agrees well with the Einasto profile. Coupling the Einasto profile with the depletion catalog, the resulting halo model can simultaneously predict the halo-matter power spectra to $10\%$ and matter-matter power spectrum to $5\%$, improving over conventional models in both the interpretability and versatility. The conditions and limitation of using the Navarro-Frenk-White profile in the halo model are also discussed.

astro-ph.CO

Decision Boundary-aware Knowledge Consolidation Generates Better Instance-Incremental Learner

Instance-incremental learning (IIL) focuses on learning continually with data of the same classes. Compared to class-incremental learning (CIL), the IIL is seldom explored because IIL suffers less from catastrophic forgetting (CF). However, besides retaining knowledge, in real-world deployment scenarios where the class space is always predefined, continual and cost-effective model promotion with the potential unavailability of previous data is a more essential demand. Therefore, we first define a new and more practical IIL setting as promoting the model's performance besides resisting CF with only new observations. Two issues have to be tackled in the new IIL setting: 1) the notorious catastrophic forgetting because of no access to old data, and 2) broadening the existing decision boundary to new observations because of concept drift. To tackle these problems, our key insight is to moderately broaden the decision boundary to fail cases while retain old boundary. Hence, we propose a novel decision boundary-aware distillation method with consolidating knowledge to teacher to ease the student learning new knowledge. We also establish the benchmarks on existing datasets Cifar-100 and ImageNet. Notably, extensive experiments demonstrate that the teacher model can be a better incremental learner than the student model, which overturns previous knowledge distillation-based methods treating student as the main role.

cs.LG

A physical and concise halo model based on the depletion radius

We develop a self-consistent and accurate halo model by partitioning matter according to the depletion radii of haloes. Unlike conventional models that define haloes with the virial radius while relying on a separate exclusion radius or ad-hoc fixes to account for halo exclusion, our model distributes mass across all scales self-consistently and accounts for both the virialized and non-virialized matter distribution around each halo. Using a cosmological simulation, we show that our halo definition leads to very simple and intuitive model components, with the one-halo term given by the Einasto profile with no truncation needed, and the halo-halo correlation function following a universal power-law form down to the halo boundary. The universal halo-halo correlation also allows us to easily model the distribution of unresolved haloes as well as diffuse matter. Convolving the halo profile with the halo-halo correlation function, we obtain a complete description of the halo-matter correlation across all scales, which self-consistently accounts for halo exclusion at the transition scale. Mass conservation is explicitly maintained in our model, and the scale dependence of the classical halo bias is easily reproduced. Our model can successfully reconstruct the halo-matter correlation function within an accuracy of $9\%$ for halo virial masses in the range of $10^{11.5}h^{-1}{\rm M}_{\odot}<M_{\rm vir}<10^{15.35}h^{-1}{\rm M}_{\odot}$ at $z=0$, and covers the radial range of $0.01h^{-1}{\rm Mpc}<r<20h^{-1}{\rm Mpc}$. We also show that our model profile can accurately predict the characteristic depletion radius at the minimum bias and the splash-back radius at the steepest density slope locations.

astro-ph.CO

Joint Learning Content and Degradation Aware Feature for Blind Super-Resolution

To achieve promising results on blind image super-resolution (SR), some attempts leveraged the low resolution (LR) images to predict the kernel and improve the SR performance. However, these Supervised Kernel Prediction (SKP) methods are impractical due to the unavailable real-world blur kernels. Although some Unsupervised Degradation Prediction (UDP) methods are proposed to bypass this problem, the \textit{inconsistency} between degradation embedding and SR feature is still challenging. By exploring the correlations between degradation embedding and SR feature, we observe that jointly learning the content and degradation aware feature is optimal. Based on this observation, a Content and Degradation aware SR Network dubbed CDSR is proposed. Specifically, CDSR contains three newly-established modules: (1) a Lightweight Patch-based Encoder (LPE) is applied to jointly extract content and degradation features; (2) a Domain Query Attention based module (DQA) is employed to adaptively reduce the inconsistency; (3) a Codebook-based Space Compress module (CSC) that can suppress the redundant information. Extensive experiments on several benchmarks demonstrate that the proposed CDSR outperforms the existing UDP models and achieves competitive performance on PSNR and SSIM even compared with the state-of-the-art SKP methods.

cs.CV

Thunder: Thumbnail based Fast Lightweight Image Denoising Network

To achieve promising results on removing noise from real-world images, most of existing denoising networks are formulated with complex network structure, making them impractical for deployment. Some attempts focused on reducing the number of filters and feature channels but suffered from large performance loss, and a more practical and lightweight denoising network with fast inference speed is of high demand. To this end, a \textbf{Thu}mb\textbf{n}ail based \textbf{D}\textbf{e}noising Netwo\textbf{r}k dubbed Thunder, is proposed and implemented as a lightweight structure for fast restoration without comprising the denoising capabilities. Specifically, the Thunder model contains two newly-established modules: (1) a wavelet-based Thumbnail Subspace Encoder (TSE) which can leverage sub-bands correlation to provide an approximate thumbnail based on the low-frequent feature; (2) a Subspace Projection based Refine Module (SPR) which can restore the details for thumbnail progressively based on the subspace projection approach. Extensive experiments have been carried out on two real-world denoising benchmarks, demonstrating that the proposed Thunder outperforms the existing lightweight models and achieves competitive performance on PSNR and SSIM when compared with the complex designs.

cs.CV

Fast Camera Image Denoising on Mobile GPUs with Deep Learning, Mobile AI 2021 Challenge: Report

Image denoising is one of the most critical problems in mobile photo processing. While many solutions have been proposed for this task, they are usually working with synthetic data and are too computationally expensive to run on mobile devices. To address this problem, we introduce the first Mobile AI challenge, where the target is to develop an end-to-end deep learning-based image denoising solution that can demonstrate high efficiency on smartphone GPUs. For this, the participants were provided with a novel large-scale dataset consisting of noisy-clean image pairs captured in the wild. The runtime of all models was evaluated on the Samsung Exynos 2100 chipset with a powerful Mali GPU capable of accelerating floating-point and quantized neural networks. The proposed solutions are fully compatible with any mobile GPU and are capable of processing 480p resolution images under 40-80 ms while achieving high fidelity results. A detailed description of all models developed in the challenge is provided in this paper.

eess.IV

Optimal estimation of functionals of high-dimensional mean and covariance matrix

Motivated by portfolio allocation and linear discriminant analysis, we consider estimating a functional $\mathbf{\mu}^T \mathbf{\Sigma}^{-1} \mathbf{\mu}$ involving both the mean vector $\mathbf{\mu}$ and covariance matrix $\mathbf{\Sigma}$. We study the minimax estimation of the functional in the high-dimensional setting where $\mathbf{\Sigma}^{-1} \mathbf{\mu}$ is sparse. Akin to past works on functional estimation, we show that the optimal rate for estimating the functional undergoes a phase transition between regular parametric rate and some form of high-dimensional estimation rate. We further show that the optimal rate is attained by a carefully designed plug-in estimator based on de-biasing, while a family of naive plug-in estimators are proved to fall short. We further generalize the estimation problem and techniques that allow robust inputs of mean and covariance matrix estimators. Extensive numerical experiments lend further supports to our theoretical results.

math.ST

Characterization and Optical Properties of Erbium doped As2S3 Films Prepared by Multi-layer Magnetron Sputtering

As2S3 film doped with erbium is prepared using multi-layer magnetron sputtering. The optical properties were measured by reflectance spectroscopy, and its chemical composition is examined by x-ray photoelectron, Rutherford backscattering, and Raman spectroscopy. The results show that the refractive index and absorption coefficient follow closely to a sputtered As2S3 film, and there are no detectable Er-S clusters and photo-induced As2O3 in the film. Rutherford backscattering spectroscopy shows that the film is homogeneous, and revealed the concentration level of erbium, and the stoichiometry of the film. The deposition method was used to fabricate an integrated Erdoped As2S3 Mach-Zehnder Interferometer and the presence of active erbium ions in the waveguide is evident from the green luminescence it emitted when it was pumped by 1488 nm diode laser. This method is attractive because the doping process can produce an Er:As2S3 film that is close to the ideal stoichiometry of As2S3 with lower risk of photo-decomposed As2O3 crystals developing on the surface when the as-deposited film is exposed to the environment.

cond-mat.mtrl-sci