SearcharxivSearch

arXiv subjects

Qi Ding

Publications and source records attributed to Qi Ding.

At least 19 recordsLinked to original sources

SlimPer: Make Personalization Model Slim and Smart

Transformer-style architectures are increasingly adopted for industrial recommendation systems, yet they inherit a design premise misaligned with the task: generative models rely on per-token autoregressive prediction, which justifies maintaining large intermediate tensors that scale with sequence length. In contrast, recommendation systems produce a single set of relevance scores for each pair without token-level supervision. Leveraging this observation, we propose SlimPer, which reformulates personalized ranking as iterative refinement of a compact, unified knowledge base. At each layer, the model selectively queries raw multi-modal user-side tokens, computes explicit relevance matching scores, and refines the knowledge base, all in O(N) per-layer cost with a fixed-size intermediate representation. As a result, model depth is decoupled from user history length, enabling deeper relevance understanding without proportional growth in compute or memory; request-only optimization further trims memory by sharing a single copy of user-side tokens across all candidate items. SlimPer unifies sparse, dense, and sequence features within a single backbone and provides inherent interpretability through its attention mechanism. Deployed on Instagram Reels and Feed, SlimPer yields measurable improvements in user engagement while streamlining the overall system and enabling effective modeling of 10k+ fine-grained user history events.

cs.IR

Task-Differentiated Atomic Skill Expansion and Routing for Continual Learning Across Highly Heterogeneous Tasks

Continual learning (CL) is commonly studied under the assumption that sequential tasks are semantically related or structurally similar. However, in highly heterogeneous settings, where tasks differ substantially in reasoning patterns and input-output formats, existing methods often suffer from catastrophic forgetting and inefficient capacity allocation. To address this challenge, we propose Task-differentiated Atomic Skill Expansion and Routing (\texttt{TASER}), a CL framework that jointly determines how many new atomic skills to introduce for each task and which skills to activate. The framework first uses atomic skill incremental learning to dynamically expand capacity based on task divergence and model uncertainty. It then applies orthogonality-enhanced skill detection to ensure these skills remain semantically distinct and independently reusable. Finally, a skill dynamic routing mechanism composes task-relevant skills through lightweight task-conditioned gating. We further introduce \texttt{HeteroCLBench}, a highly heterogeneous benchmark for CL, comprising 19 diverse tasks across 9 cognitive dimensions under a standardized sequential protocol. Experiments on \texttt{HeteroCLBench} show that \texttt{TASER} consistently outperforms strong baselines by improving plasticity and reducing catastrophic forgetting.

cs.LG

Integrability of Lawson-Osserman Cone and its Applications

In this paper, we characterize all eigenfunctions corresponding to nonpositive eigenvalues of the Jacobi operator of the link $M$ of the Lawson-Osserman cone $\mathbf{C}$ in $\mathbb{R}^7$. In particular, we prove that $\mathbf{C}$ is integrable, i.e., all Jacobi fields on $\mathbf{C}$ of homogeneous degree 1 and 0, are generated by rotations and translations in $\mathbb{R}^7$. As applications, we prove that $M$ is rigid as minimal submanifolds in $\mathbb{S}^6$, and derive the optimal decay order for minimal submanifolds in $\mathbb{R}^7$ asymptotic to $\mathbf{C}$ at infinity.

math.DG

Benchmark Shadows: Data Alignment, Parameter Footprints, and Generalization in Large Language Models

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To investigate this, we design controlled data interventions that isolate distributional effects under fixed training settings. We find that benchmark-aligned data improves narrow evaluation metrics while limiting broader representational development, whereas coverage-expanding data leads to more distributed parameter adaptation and better generalization. We further introduce parameter-space diagnostics based on spectral and rank analyses, which reveal distinct structural signatures of these regimes. Similar patterns are observed across diverse open-source model families, including multimodal models as a key case study, suggesting that these effects extend beyond controlled settings. A case study on prompt repetition shows that not all data artifacts induce regime shifts. These results indicate that benchmark performance alone is insufficient to characterize model capability, and highlight the importance of data distribution in shaping learning dynamics.

cs.LG

Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less predictable than that of text-only LLMs. Increasing model size and task diversity often yields diminishing returns. In this work, we argue that the primary bottleneck in multimodal scaling is not task format, but knowledge density in training data. We first show that task-specific supervision such as Visual Question Answering (VQA) contributes little incremental semantic information beyond image captions: VQA signals can be reconstructed from captions with negligible performance loss. We then demonstrate that increasing knowledge density -- through structured caption enrichment and cross-modal knowledge injection -- leads to consistent performance improvements across multimodal and downstream benchmarks. Across controlled experiments, performance correlates more strongly with semantic coverage than with task diversity. These findings suggest that current MLLMs fail to scale primarily because training data lacks sufficient knowledge coverage. We advocate for knowledge-centric multimodal training as a principled foundation for scalable multimodal models.

cs.CL

Topology of complete minimal submanifolds in $\mathbb{R^{n+m}}$ with finite total curvature

In [CKM17], Chodosh, Ketover, and Maximo proved finite diffeomorphism theorems for complete embedded minimal hypersurfaces of dimension $\leqslant$ 6 with finite index and bounded volume growth ratio. In this paper, we adapt their method to study finite diffeomorphism types for complete immersed minimal submanifolds of arbitrary codimension in Euclidean space with finite total curvature and Euclidean volume growth.

math.DG

Frequency- and Amplitude-Modulated Gates for Universal Quantum Control

Achieving high-fidelity single- and two-qubit gates is essential for executing arbitrary digital quantum algorithms and for building error-corrected quantum computers. We propose a theoretical framework for implementing quantum gates using frequency- and amplitude-modulated microwave control, which extends conventional amplitude modulation by introducing frequency modulation as an additional degree of control. Our approach operates on fixed-frequency qubits, converting the need for qubit frequency tunability into drive frequency modulation. Using Floquet theory, we analyze and design these drives for optimal fidelity within specified criteria. Our framework spans adiabatic to nonadiabatic gates within the Floquet framework, ensuring broad applicability across gate types and control schemes. Using typical transmon qubit parameters in numerical simulations, we demonstrate a universal gate set-including the X, Hadamard, phase, and CZ gates-with control error well below 0.1% and gate times of 25-40 ns for single-qubit operations and 125-135 ns for two-qubit operations. Furthermore, we show an always-on CZ gate tailored for driven qubits, which has gate times of 80-90 ns.

quant-ph

ZZ-Free Two-Transmon CZ Gate Mediated by a Fluxonium Coupler

Eliminating residual ZZ interactions in a two-qubit system is essential for reducing coherent errors during quantum operations. In a superconducting circuit platform, coupling two transmon qubits via a transmon coupler has been shown to effectively suppress residual ZZ interactions. However, in such systems, perfect cancellation usually requires the qubit-qubit detuning to be smaller than the individual qubit anharmonicities, which exacerbates frequency crowding and microwave crosstalk. To address this limitation, we introduce TFT (Transmon-Fluxonium-Transmon) architecture, wherein two transmon qubits are coupled via a fluxonium qubit. The coupling mediated by the fluxonium eliminates residual ZZ interactions even for transmons detuned larger than their anharmonicities. We experimentally identified zero-ZZ interaction points at qubit-qubit detunings of 409 MHz and 616 MHz from two distinct TFT devices. We then implemented an adiabatic, coupler-flux-biased controlled-Z gate on both devices, achieving CZ gate fidelities of 99.64(6)% and 99.68(8)%.

quant-ph

Application of Optimal Control to Time-Resolution Protocol for Quantum Sensing

Time-resolution protocol of quantum sensing aims to measure the fast temporal variation of an external field and demands a high field sensitivity in a short interrogation time $\tau$. Since any operation that evolves the quantum state takes time and is counted as part of the interrogation, evaluating the performance of time-resolution protocol requires a complete end-to-end description of the measurement process. In particular, the initial state has to be one of the sensor qubit's eigenstates in the absence of external fields, and the final projective measurements must be performed in the same eigenstate basis. Building upon prior works which proposed limits for time-resolved sensing using a quantum sensor, we apply optimal control theory to optimize the time-resolution protocol. Our analysis indicates that there exists a critical interrogation time $T^*$: when $\tau T^*$ the optimal protocol involves a singular control during the interrogation. In the short-$\tau$ regime, which is relevant to high time resolution, we propose a ``detune protocol'' that involves only smooth control during the entire interrogation. As the discontinuities of control pose the main obstacles to experimental realization, we expect the presented detune protocol to be practically useful. In the long-$\tau$ regime, the optimal protocol closely resembles the Ramsey sequence; protocols based on maximizing Quantum Fisher Information are constructed to highlight the difference between the theoretically optimal and practically implementable measurements. Effective use of the time-resolution protocol requires a setup where the unknown time-domain signal of interest can be identically and repeatedly generated. As a potentially relevant application, we outline the calibration of baseband flux pulse distortion in the control of superconducting qubits.

quant-ph

BlueLM-2.5-3B Technical Report

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our knowledge, this is the first 3B-scale MLLM to support both thinking and non-thinking modes, while also enabling explicit control over thinking token budget. BlueLM-2.5-3B is developed through diversified data curation, key data resampling, hybrid heterogeneous reinforcement learning, and a high-performance training infrastructure. Our model achieves superior multimodal capacity while preserving competitive pure-text performance with only 2.9 billion parameters. We conduct comprehensive evaluations across a broad range of multimodal and text-only benchmarks. In thinking mode, BlueLM-2.5-3B achieves comparable performance to Qwen3-4B on text-only benchmarks, and trails the larger Kimi-VL-A3B-16B by only about 5% on average across multimodal evaluations. In non-thinking mode, it outperforms Qwen2.5-VL-3B on the majority of multimodal benchmarks. Additionally, BlueLM-2.5-3B exhibits exceptional data efficiency. All of the aforementioned performance is achieved with substantially less total training data than Qwen2.5-VL-3B and Qwen3-4B. We hope our work contributes to the advancement of high-performance, on-device MLLMs and provides meaningful insights to the research community.

cs.AI

Strong coupling and interfering resonances in isolated van der Waals nanoresonators

The study of strong light-matter interaction in van der Waals materials is at the forefront of current research in physics and chemistry, and it can be enhanced dramatically by employing resonances. Here we present the first observation of quasi-bound states in the continuum (qBICs) realized via polaritonic interfering resonances in isolated WS$_2$ nanodisks. We experimentally validate the existence of polaritonic qBICs driven by intrinsic coupling of Mie resonances and excitons. The system exhibits exceptionally strong light-matter interaction with a measured Rabi splitting exceeding 310 meV - the largest reported value among all transition metal dichalcogenide (TMDC) self-hybridized systems to date. The giant coupling strength stems from qBIC-induced in-plane field enhancement, which strongly interacts with in-plane excitonic dipoles while suppressing radiative losses. Polarization-controlled measurements further demonstrate selective excitation of qBIC through switching incident polarization to specific orthogonal configurations. The observed polarization-dependent coupling provides an additional degree of freedom to control over the hybrid states' spectral characteristics and spatial field distributions. Our demonstrations provide a pathway for engineering high-quality light-matter hybrid states in compact nanostructures, with potential applications in on-chip photonics, polaritonics, and quantum optics.

physics.optics

Time-optimal single-scalar control on a qubit of unitary dynamics

Optimal control theory is applied to analyze the time-optimal solution with a single scalar control knob in a two-level quantum system without quantum decoherence. Emphasis is \change{placed} on the dependence on the maximum control strength $u_\text{max}$. General constraints on the optimal protocol are derived and used to rigorously parameterize the time-optimal solution. Two concrete problems are investigated. For generic state preparation problems, both multiple bang-bang and bang-singular-bang are legitimate and should be considered. Generally, the optimal is bang-bang for small $u_\text{max}$, and there exists a state-dependent critical amplitude above which singular control emerges. For the X-gate operation of a qubit, the optimal protocol \change{is exclusively} multiple bang-bang. The minimum gate time is about 80\% of that based on the resonant Rabi $\pi$-pulse over a wide range of control strength; in the $u_\text{max} \rightarrow 0$ limit this ratio is derived to be $\pi/4$. To develop practically feasible protocols, we present methods to smooth the abrupt changes in the bang-bang control while preserving perfect gate fidelity. \change{The presence of bang-bang segments in the time-optimal protocol} indicates that the high-frequency components and a full calculation (instead of the commonly adopted Rotating Wave Approximation) are essential for the ultimate quantum speed limit.

quant-ph

Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Language model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient manner. Specifically, we draw inspiration from recent findings showing that compression efficiency (i.e., the normalized loss) of diverse models on certain text correlates strongly with their downstream performance, when the text domain aligns with the downstream benchmarks(Huang et al., 2024). Building on this observation, we hypothesize that data on which model losses are predictive of downstream abilities also contribute effectively to learning, which shares similar intuition with Thrush et al.(2024). To leverage this insight, we introduce predictive data selection (PreSelect), a lightweight and efficient data selection method that requires training and deploying only a fastText-based scorer. Through comprehensive experiments with 1B and 3B parameter models, we demonstrate that models trained on 30B tokens selected with PreSelect surpass the performance of the vanilla baseline trained on 300B tokens, achieving a 10x reduction in compute requirements. Furthermore, PreSelect significantly outperforms other competitive data selection baselines, such as DCLM and FineWeb-Edu on a scale of 3B models trained on 100B tokens. We open-source our trained data selection scorer along with the curated datasets at https://github.com/hkust-nlp/PreSelect.

cs.CL

Hessian estimates for Lagrangian mean curvature equation with Lipschitz critical and supercritical phases

In this paper, we develop a new strategy to study Lagrangain mean curvature equation on open sets of $\mathbb{R}^{n}(n\geq2)$. By establishing an Allard-type regularity theorem, we obtain an interior Hessian estimate of solutions to this equation with prescribed Lipschitz critical and supercritical phases. Here, our condition on the phases is sharp. The proof heavily relies on geometric measure theory, geometry of Lagrangian graphs, and De Giorgi-Nash-Moser iteration. We expect that the techniques and ideas developed here can be used in some other equations.

math.DG

qGDP: Quantum Legalization and Detailed Placement for Superconducting Quantum Computers

Noisy Intermediate-Scale Quantum (NISQ) computers are currently limited by their qubit numbers, which hampers progress towards fault-tolerant quantum computing. A major challenge in scaling these systems is crosstalk, which arises from unwanted interactions among neighboring components such as qubits and resonators. An innovative placement strategy tailored for superconducting quantum computers can systematically address crosstalk within the constraints of limited substrate areas. Legalization is a crucial stage in placement process, refining post-global-placement configurations to satisfy design constraints and enhance layout quality. However, existing legalizers are not supported to legalize quantum placements. We aim to address this gap with qGDP, developed to meticulously legalize quantum components by adhering to quantum spatial constraints and reducing resonator crossing to alleviate various crosstalk effects. Our results indicate that qGDP effectively legalizes and fine-tunes the layout, addressing the quantum-specific spatial constraints inherent in various device topologies. By evaluating diverse NISQ benchmarks. qGDP consistently outperforms state-of-the-art legalization engines, delivering substantial improvements in fidelity and reducing spatial violation, with average gains of 34.4x and 16.9x, respectively.

quant-ph

Pulse Design of Baseband Flux Control for Adiabatic Controlled-Phase Gates in Superconducting Circuits

Despite progress towards achieving low error rates with superconducting qubits, error-prone two-qubit gates remain a bottleneck for realizing large-scale quantum computers. Therefore, a systematic framework to design high-fidelity gates becomes imperative. One type of two-qubit gate in superconducting qubits is the controlled-phase (CPHASE) gate, which utilizes a conditional interaction between higher energy levels of the qubits controlled by a baseband flux pulse on one of the qubits or a tunable coupler. In this work, we study an adiabatic implementation of CPHASE gates and formulate the design of the control trajectory for the gate as a pulse-design problem. We show in simulation that the Chebyshev-based trajectory can, in certain cases, enable gates with gate infidelity lower by an average of 23.3% when compared to the widely used Slepian-based trajectory.

quant-ph

Qplacer: Frequency-Aware Component Placement for Superconducting Quantum Computers

Noisy Intermediate-Scale Quantum (NISQ) computers face a critical limitation in qubit numbers, hindering their progression towards large-scale and fault-tolerant quantum computing. A significant challenge impeding scaling is crosstalk, characterized by unwanted interactions among neighboring components on quantum chips, including qubits, resonators, and substrate. We motivate a general approach to systematically resolving multifaceted crosstalks in a limited substrate area. We propose Qplacer, a frequency-aware electrostatic-based placement framework tailored for superconducting quantum computers, to alleviate crosstalk by isolating these components in spatial and frequency domains alongside compact substrate design. Qplacer commences with a frequency assigner that ensures frequency domain isolation for qubits and resonators. It then incorporates a padding strategy and resonator partitioning for layout flexibility. Central to our approach is the conceptualization of quantum components as charged particles, enabling strategic spatial isolation through a 'frequency repulsive force' concept. Our results demonstrate that Qplacer carefully crafts the physical component layout in mitigating various crosstalk impacts while maintaining a compact substrate size. On various device topologies and NISQ benchmarks, Qplacer improves fidelity by an average of 36.7x and reduces spatial violations (susceptible to crosstalk) by an average of 12.76x, compared to classical placement engines. Regarding area optimization, compared to manual designs, Qplacer can reduce the required layout area by 2.14x on average

quant-ph

Liouville theorem for minimal graphs over manifolds of nonnegative Ricci curvature

Let $\Sigma$ be a complete Riemannian manifold of nonnegative Ricci curvature. We prove a Liouville-type theorem: every smooth solution $u$ to minimal hypersurface equation on $\Sigma$ is a constant provided $u$ has sublinear growth for its negative part. Here, the sublinear growth condition is sharp. Our proof relies on a gradient estimate for minimal graphs over $\Sigma$ with small linear growth of the negative parts of graphic functions via iteration.

math.DG