SearcharxivSearch

arXiv subjects

Wenqi Zhu

Publications and source records attributed to Wenqi Zhu.

At least 19 recordsLinked to original sources

PCap: Personalized Retrieval-Stage Diversity Capping in Facebook Marketplace

We propose a personalized capping framework (PCap) to improve the diversity in Facebook Marketplace by introducing user-level diversity constraints at the retrieval stage. PCap models individual diversity preferences using Shannon entropy-based scoring, segments users into diversity buckets, and applies personalized category caps during multi-source candidate retrieval. To navigate the high-dimensional parameter space of per-bucket caps, we leverage an automated online optimization method called Parameter Tuning Sequence. Large-scale online experiments demonstrate that PCap significantly improves users' browsing experience shown in engagement metrics. This work provides practical insights into integrating personalized diversity into industrial retrieval systems.

cs.IR

A Globally Convergent Third-Order Newton Method via Unified Semidefinite Programming Subproblems

We propose the Adaptive Levenberg-Marquardt Third-Order Newton Method (ALM-TON) method for unconstrained nonconvex optimization; to our knowledge, the framework provides the first globally convergent realization of the unregularized third-order Newton method. Unlike the standard Adaptive Regularization framework with third-order models (AR3), which enforces global behavior through a quartic term, ALMTON employs an adaptive Levenberg-Marquardt (quadratic) regularization. This choice preserves a cubic model at every iteration, so that every subproblem is a tractable semidefinite programming (SDP). Algorithmically, ALMTON follows a mixed-mode strategy: it attempts an unregularized thirdorder step whenever the cubic Taylor model admits a strict local minimizer with adequate curvature, and activates (or increases) quadratic regularization only when needed to ensure that the model is well posed and the step is globally reliable. For the Heuristic strategy, under the stated assumptions and an exact local-minimizer oracle, we prove finite termination at an $ε$-approximate first-order stationary point with $O\left(ε^{-2}\right)$ worst-case evaluation complexity. Moreover, if an accepted iterate enters the stated neighborhood of a positive-definite local minimizer, subsequent nonterminal steps recover the unregularized third-order Newton recursion and its cubic local rate. Under a common post hoc terminal audit over 4,500 deterministic starts on five two-dimensional nonconvex problems, both ALMTON variants satisfy the terminal criterion on $99.91 \%$ of the instances, compared with $55.42 \%$ for the unregularized third-order Newton method. This robustness gain comes at substantial SDP cost: AR$2$ is faster, so the results support robust globalization of the unregularized cubic model rather than overall empirical superiority.

math.OC

A Homogeneous Tensor Framework for High-Order Trust-Region and Spherical Polynomial Optimization

High-order methods can improve worst-case evaluation complexity, but for orders $p\geq3$ their Taylor subproblems are nonconvex polynomial optimization problems and are generally difficult to solve. We develop a radius-controlled boundary approach based on homogeneous tensor representations. By augmenting the step with a constant coordinate, any $p$th-order Taylor polynomial can be represented exactly as an order-$p$ homogeneous tensor form; at a prescribed radius, the boundary model is a spherical polynomial optimization problem. The representation applies to arbitrary $p$, while the algorithmic development focuses on the cubic case $p=3$. For an inhomogeneous cubic on the sphere, we introduce a quadratic shift and prove, under an explicit shift bound, equivalence with a three-block multilinear formulation at global optimality. This motivates a proximal alternating minimization (PAM) method with closed-form block updates; its objective values decrease and every accumulation point is stationary. We embed the boundary-step mechanism in an Adaptive Homogeneous Tensor Method (Ada--HTM). Under explicit smoothness, safeguarded-decrease, weak-curvature nondegeneracy, and local-refinement conditions, Ada--HTM attains the adaptive-regularization-type (AR$p$-type) evaluation complexity $\mathcal{O}(ε^{-(p+1)/p})$ for first-order stationarity. Numerically, PAM matches order-$2$ moment--sum-of-squares (SOS) certificates on the structured cubic instances for which certification is tractable, scales particularly well for low-rank tensors, and makes Ada--HTM competitive with trust-region and cubic-regularization methods, with its largest gains on ill-conditioned and badly-scaled problems.

math.OC

Global Optimality Characterizations and Algorithms for Minimizing Quartically-Regularized Third-Order Taylor Polynomials

High-order methods for convex and nonconvex optimization, particularly $p$th-order Adaptive Regularization Methods (AR$p$), have attracted significant research interest by naturally incorporating high-order Taylor models into adaptive regularization frameworks, resulting in algorithms with faster global and local convergence rates than first- and second-order methods. This paper establishes global optimality conditions for general, nonconvex cubic polynomials with quartic regularization. These criteria generalise existing results, recovering the optimality results for regularized quadratic polynomials, and can be further simplified in the low-rank and diagonal tensor cases. Under suitable assumptions on the Taylor polynomial, we derive a lower bound for the regularization parameter such that the necessary and sufficient criteria coincide, establishing a connection between this bound and the subproblem's convexification and sum-of-squares (SoS) convexification techniques. Leveraging the optimality characterization, we develop a Diagonal Tensor Method (DTM) for minimizing quartically-regularized cubic Taylor polynomials by iteratively minimizing a sequence of local models that incorporate both diagonal cubic terms and quartic regularization (DTM model). We show that the DTM algorithm is provably convergent, with a global evaluation complexity of $\mathcal{O}(ε^{-3/2})$. Furthermore, when special structure is present (such as low rank or diagonal), DTM can exactly solve the given problem (in one iteration). In our numerical experiments, we propose practical DTM variants that exploit local problem information for model construction, which we then show to be competitive with cubic regularization and other subproblem solvers, with superior performance on problems with special structure.

math.OC

A scalable infrastructure for strontium optical clocks with integrated photonics

Optical atomic clocks provide exceptionally accurate and precise signals for timekeeping and precision measurements, but they require high-power, free-space laser configurations that limit scalability. We introduce and explore a scalable infrastructure for strontium (Sr) optical-lattice clocks that incorporates co-design of atomic-beam slowing and a magneto-optical trap (MOT) from an effusion source, generation of complex, three-dimensional free-space laser configurations with a photonic integrated circuit (PIC) and metasurface (MS) optics, and laser stabilization to a frequency-comb supercontinuum generated with integrated nonlinear photonics. With these elements, we realize MOTs of all stable strontium isotopes ($^{84}$Sr, $^{86}$Sr, $^{87}$Sr, $^{88}$Sr) with populations commensurate with natural abundances, demonstrating precise beam control and robustness. Access to laser-cooled alkaline-earth atoms with scalable integrated photonics enables system engineering for optical clocks, quantum sensing, and quantum information, and our experiments demonstrate extensible technologies that advance toward a Sr optical clock largely free of bulk optics.

physics.app-ph

Sufficiently Regularized Nonnegative Quartic Polynomials are Sum-of-Squares

A polynomial that is nonnegative need not be a sum of squares of polynomials. This classical gap, identified by Hilbert in 1888, lies at the heart of why the global optimization of multivariate quartic polynomials is NP-hard. Yet we show that this gap is closed when using (sufficient) regularization, which fundamentally alters the algebraic structure of the problem. Namely, we investigate a class of quartically-regularized cubic polynomials which arise naturally in polynomial optimization and higher-order tensor methods for nonconvex problems. We show that, under mild assumptions and for sufficiently large Euclidean quartic regularization, the shifted nonnegative polynomial becomes a sum of squares, yielding an exact semidefinite programming (SDP) formulation at the zeroth level of the Lasserre hierarchy. We further derive explicit bounds on the regularization parameter that guarantee this property. Beyond this asymptotic regime, we identify structured subclasses for which SoS exactness holds for all regularization levels, including quadratic-quartic models and a class of low-rank-type cubic tensors. In contrast, we show that separable quartic regularized polynomials -- including classical tensor models proposed by Schnabel (1991) -- do not, in general, induce SoS representations, even under arbitrarily large regularization. Our results reveal a sharp structural boundary between tractable and intractable regimes in polynomial optimization. In particular, they explain why Euclidean quartic regularization plays a significant role: in addition to regularising the model, it can induce exact SoS certificates and exact SDP representations.

math.OC

Recover Cell Tensor: Diffusion-Equivalent Tensor Completion for Fluorescence Microscopy Imaging

Fluorescence microscopy (FM) imaging is a fundamental technique for observing live cell division, one of the most essential processes in the cycle of life and death. Observing 3D live cells requires scanning through the cell volume while minimizing lethal phototoxicity. That limits acquisition time and results in sparsely sampled volumes with anisotropic resolution and high noise. Existing image restoration methods, primarily based on inverse problem modeling, assume known and stable degradation processes and struggle under such conditions, especially in the absence of high-quality reference volumes. In this paper, from a new perspective, we propose a novel tensor completion framework tailored to the nature of FM imaging, which inherently involves nonlinear signal degradation and incomplete observations. Specifically, FM imaging with equidistant Z-axis sampling is essentially a tensor completion task under a uniformly random sampling condition. On one hand, we derive the theoretical lower bound for exact cell tensor completion, validating the feasibility of accurately recovering 3D cell tensor. On the other hand, we reformulate the tensor completion problem as a mathematically equivalent score-based generative model. By incorporating structural consistency priors, the generative trajectory is effectively guided toward denoised and geometrically coherent reconstructions. Our method demonstrates state-of-the-art performance on SR-CACO-2 and three real \textit{in vivo} cellular datasets, showing substantial improvements in both signal-to-noise ratio and structural fidelity.

eess.IV

Efficient Implementation of Third-Order Tensor Methods with Adaptive Regularization for Unconstrained Optimization

High-order tensor methods that employ local Taylor models of degree $p$ within adaptive regularization frameworks (AR$p$) have recently received significant attention, due to their optimal/improved global and local rates of convergence, for both convex and nonconvex optimization problems. In this paper, we showcase the numerical performance of standard second- and third-order variants ($p=2,3$) and propose novel techniques for key algorithmic aspects when $p\geq 3$. In particular, we extend the interpolation-based updating strategy for the regularization parameter introduced in [Gould, Porcelli and Toint, Comput Optim Appl (2012) 53:1--22] for $p=2$, to the case when $p \geq 3$. We identify fundamental differences between the different local minima of the regularised subproblems for $p=2$ and $p \geq 3$ and their effect on algorithm performance. For $p\geq 3$, we introduce a novel pre-rejection technique that rejects poor/unsuccessful subproblem minimizers prior to any function evaluation. Numerical studies showcase the efficiency improvements generated by our proposed modifications of the AR$3$ algorithm. We also assess numerically, the effect of different subproblem termination conditions and the choice of the initial regularization parameter on the overall algorithm performance. Finally, we benchmark our best-performing AR$3$ variants, as well as those in [Birgin et al., Optim Lett (2020) 14:815--838], against second-order ones (AR$2$). Encouraging results on standard test problems are obtained, confirming that AR$3$ variants can be made to outperform second-order variants in terms of objective evaluations, derivative evaluations, and number of subproblem solves. We provide an efficient, extensive and modular software package in MATLAB that includes many AR$2$ and AR$3$ variants, including Hessian- and tensor-free ones, allowing ease of use and experimentation for interested users.

math.OC

Dexbotic: Open-Source Vision-Language-Action Toolbox

In this paper, we present Dexbotic, an open-source Vision-Language-Action (VLA) model toolbox based on PyTorch. It aims to provide a one-stop VLA research service for professionals in the field of embodied intelligence. It offers a codebase that supports multiple mainstream VLA policies simultaneously, allowing users to reproduce various VLA methods with just a single environment setup. The toolbox is experiment-centric, where the users can quickly develop new VLA experiments by simply modifying the Exp script. Moreover, we provide much stronger pretrained models to achieve great performance improvements for state-of-the-art VLA policies. Dexbotic will continuously update to include more of the latest pre-trained foundation models and cutting-edge VLA models in the industry.

cs.RO

Coupling a Fabry-Pérot Cavity to a Single-Mode Optical Fiber Using a Metalens

Efficient coupling of light from an optical cavity to a single-mode fiber is required in a range of quantum technologies. In this work we consider the coupling of a high-finesse macroscopic Fabry-Pérot (FP) cavity to a single-mode fiber using a metalens. We perform sensitivity analysis with respect to longitudinal and transverse misalignment errors. We then detail a fiber-coupled cavity at 1650 nm using a monolithic cryo-compatible assembly incorporating a metalens.

physics.optics

Convergence and Near-optimal Sampling for Multivariate Function Approximations in Irregular Domains via Vandermonde with Arnoldi

Vandermonde matrices are usually exponentially ill-conditioned and often result in unstable approximations. In this paper, we introduce and analyze the \textit{multivariate Vandermonde with Arnoldi (V+A) method}, which is based on least-squares approximation together with a Stieltjes orthogonalization process, for approximating continuous, multivariate functions on $d$-dimensional irregular domains. The V+A method addresses the ill-conditioning of the Vandermonde approximation by creating a set of discrete orthogonal bases with respect to a discrete measure. The V+A method is simple and general, relying only on the domain's sample points. This paper analyzes the sample complexity of {the least-squares approximation that uses the V+A method}. We show that, for a large class of domains, this approximation gives a well-conditioned and near-optimal $N$-dimensional least-squares approximation using $M=O(N^2)$ equispaced sample points or $M=O(N^2\log N)$ random sample points, independently of $d$. We provide a comprehensive analysis of the error estimates and the rate of convergence of the least-squares approximation that uses the V+A method. Based on the multivariate V+A techniques, we propose a new variant of the weighted V+A least-squares algorithm that uses only $M=O(N\log N)$ sample points to achieve a near-optimal approximation. {Our initial numerical results validate that the V+A least-squares approximation method provides well-conditioned and near-optimal approximations for multivariate functions on (irregular) domains. Additionally, the (weighted) least-squares approximation that uses the V+A method performs competitively with state-of-the-art orthogonalization techniques and can serve as a practical tool for selecting near-optimal distributions of sample points in irregular domains.

math.NA

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level grounding in videos. To address this, we introduce VideoGLaMM, a LMM designed for fine-grained pixel-level grounding in videos based on user-provided textual inputs. Our design seamlessly connects three key components: a Large Language Model, a dual vision encoder that emphasizes both spatial and temporal details, and a spatio-temporal decoder for accurate mask generation. This connection is facilitated via tunable V-L and L-V adapters that enable close Vision-Language (VL) alignment. The architecture is trained to synchronize both spatial and temporal elements of video content with textual instructions. To enable fine-grained grounding, we curate a multimodal dataset featuring detailed visually-grounded conversations using a semiautomatic annotation pipeline, resulting in a diverse set of 38k video-QA triplets along with 83k objects and 671k masks. We evaluate VideoGLaMM on three challenging tasks: Grounded Conversation Generation, Visual Grounding, and Referring Video Segmentation. Experimental results show that our model consistently outperforms existing approaches across all three tasks.

cs.CV

SoS1: O1 and R1-Like Reasoning LLMs are Sum-of-Square Solvers

Large Language Models (LLMs) have achieved human-level proficiency across diverse tasks, but their ability to perform rigorous mathematical problem solving remains an open challenge. In this work, we investigate a fundamental yet computationally intractable problem: determining whether a given multivariate polynomial is nonnegative. This problem, closely related to Hilbert's Seventeenth Problem, plays a crucial role in global polynomial optimization and has applications in various fields. First, we introduce SoS-1K, a meticulously curated dataset of approximately 1,000 polynomials, along with expert-designed reasoning instructions based on five progressively challenging criteria. Evaluating multiple state-of-the-art LLMs, we find that without structured guidance, all models perform only slightly above the random guess baseline 50%. However, high-quality reasoning instructions significantly improve accuracy, boosting performance up to 81%. Furthermore, our 7B model, SoS-7B, fine-tuned on SoS-1K for just 4 hours, outperforms the 671B DeepSeek-V3 and GPT-4o-mini in accuracy while only requiring 1.8% and 5% of the computation time needed for letters, respectively. Our findings highlight the potential of LLMs to push the boundaries of mathematical reasoning and tackle NP-hard problems.

cs.LG

Tensor-based Dinkelbach method for computing generalized tensor eigenvalues and its applications

In this paper, we propose a novel tensor-based Dinkelbach--Type method for computing extremal tensor generalized eigenvalues. We show that the extremal tensor generalized eigenvalue can be reformulated as a critical subproblem of the classical Dinkelbach--Type method, which can subsequently be expressed as a multilinear optimization problem (MOP). The MOP is solved under a spherical constraint using an efficient proximal alternative minimization method, in which we rigorously establish the global convergence. Additionally, the equivalent MOP is reformulated as an unconstrained optimization problem, allowing for the analysis of the Kurdyka-Lojasiewicz (KL) exponent and providing an explicit expression for the convergence rate of the proposed algorithm. Preliminary numerical experiments on solving extremal tensor generalized eigenvalues and minimizing high-order trust-region subproblems are provided, validating the efficacy and practical utility of the proposed method.

math.NA

Second-order methods for quartically-regularised cubic polynomials, with applications to high-order tensor methods

There has been growing interest in high-order tensor methods for nonconvex optimization, with adaptive regularization, as they possess better/optimal worst-case evaluation complexity globally and faster convergence asymptotically. These algorithms crucially rely on repeatedly minimizing nonconvex multivariate Taylor-based polynomial sub-problems, at least locally. Finding efficient techniques for the solution of these sub-problems, beyond the second-order case, has been an open question. This paper proposes a second-order method, Quadratic Quartic Regularisation (QQR), for efficiently minimizing nonconvex quartically-regularized cubic polynomials, such as the AR$p$ sub-problem [3] with $p=3$. Inspired by [35], QQR approximates the third-order tensor term by a linear combination of quadratic and quartic terms, yielding (possibly nonconvex) local models that are solvable to global optimality. In order to achieve accuracy $ε$ in the first-order criticality of the sub-problem in finitely many iterations, we show that the error in the QQR method decreases either linearly or by at least $\mathcal{O}(ε^{4/3})$ for locally convex iterations, while in the nonconvex case, by at least $\mathcal{O}(ε)$; thus improving, on these types of iterations, the general cubic-regularization bound. Preliminary numerical experiments indicate that two QQR variants perform competitively with state-of-the-art approaches such as ARC (also known as AR$p$ with $p=2$), achieving either a lower objective value or iteration counts.

math.OC

CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation

Open-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Language-Image Pre-training (CLIP) has shown robust zero-shot classification ability in image-level open-vocabulary tasks. In this paper, we propose a simple encoder-decoder network, called CLIP-VIS, to adapt CLIP for open-vocabulary video instance segmentation. Our CLIP-VIS adopts frozen CLIP and introduces three modules, including class-agnostic mask generation, temporal topK-enhanced matching, and weighted open-vocabulary classification. Given a set of initial queries, class-agnostic mask generation introduces a pixel decoder and a transformer decoder on CLIP pre-trained image encoder to predict query masks and corresponding object scores and mask IoU scores. Then, temporal topK-enhanced matching performs query matching across frames using the K mostly matched frames. Finally, weighted open-vocabulary classification first employs mask pooling to generate query visual features from CLIP pre-trained image encoder, and second performs weighted classification using object scores and mask IoU scores. Our CLIP-VIS does not require the annotations of instance categories and identities. The experiments are performed on various video instance segmentation datasets, which demonstrate the effectiveness of our proposed method, especially for novel categories. When using ConvNeXt-B as backbone, our CLIP-VIS achieves the AP and APn scores of 32.2% and 40.2% on the validation set of LV-VIS dataset, which outperforms OV2Seg by 11.1% and 23.9% respectively. We will release the source code and models at https://github.com/zwq456/CLIP-VIS.git.

cs.CV

MLC-GCN: Multi-Level Generated Connectome Based GCN for AD Analysis

Alzheimer's Disease (AD) is a currently incurable neurodegeneartive disease. Accurately detecting AD, especially in the early stage, represents a high research priority. AD is characterized by progressive cognitive impairments that are related to alterations in brain functional connectivity (FC). Based on this association, many studies have been published over the decades using FC and machine learning to differentiate AD from healthy aging. The most recent development in this detection method highlights the use of graph neural network (GNN) as the brain functionality analysis. In this paper, we proposed a stack of spatio-temporal feature extraction and graph generation based AD classification model using resting state fMRI. The proposed multi-level generated connectome (MLC) based graph convolutional network (GCN) (MLC-GCN) contains a multi-graph generation block and a GCN prediction block. The multi-graph generation block consists of a hierarchy of spatio-temporal feature extraction layers for extracting spatio-temporal rsfMRI features at different depths and building the corresponding connectomes. The GCN prediction block takes the learned multi-level connectomes to build and optimize GCNs at each level and concatenates the learned graphical features as the final predicting features for AD classification. Through independent cohort validations, MLC-GCN shows better performance for differentiating MCI, AD, and normal aging than state-of-art GCN and rsfMRI based AD classifiers. The proposed MLC-GCN also showed high explainability in terms of learning clinically reasonable connectome node and connectivity features from two independent datasets. While we only tested MLC-GCN on AD, the basic rsfMRI-based multi-level learned GCN based outcome prediction strategy is valid for other diseases or clinical outcomes.

cs.LG

Laser cooling $^{88}$Sr to microkelvin temperature with an integrated-photonics system

We report on experiments generating a magneto-optical trap (MOT) of 88-strontium ($^{88}$Sr) atoms at microkelvin temperature, using integrated-photonics devices. With metasurface optics integrated on a fused-silica substrate, we generate six-beam, circularly polarized, counter-propagating MOTs on the blue broad-line, 461 nm, and red narrow-line, 689 nm, Sr cooling transitions without bulk optics. By use of a diverging beam configuration, we create up to 10 mm diameter MOT beams at the trapping location. To frequency stabilize and linewidth narrow the cooling lasers, we use fiber-packaged, integrated nonlinear waveguides to spectrally broaden a frequency comb. The ultra-coherent supercontinuum of the waveguides covers 650 nm to 2500 nm, enabling phase locks of the cooling lasers to hertz level linewidth. Our work highlights the possibility to simplify the preparation of an ultracold 88Sr gas for an optical-lattice clock with photonic devices. By implementing a timing sequence for control of the MOT lasers and the quadrupole magnetic-field gradient, we collect atoms directly from a thermal beam into the blue MOT and continuously cool into a red MOT with dynamic detuning and intensity control. There, the red MOT temperature is as low as $2~μ$K and the overall transfer efficiency up to 16%. We characterize this sequence, including an intermediate red MOT with modulated detuning. Our experiments demonstrate an integrated photonics system capable of cooling alkaline-earth gases to microkelvin temperature with sufficient transfer efficiencies for adoption in scalable optical clocks and quantum sensors.

physics.atom-ph