SearcharxivSearch

arXiv subjects

Wei Dai

Publications and source records attributed to Wei Dai.

At least 19 recordsLinked to original sources

Nanocavity Confinement by Orthogonal Valley- and SSH- Topological Interfaces In Glide-Symmetric Photonic Crystal Structures

Valley photonic crystals enable valley-dependent transport and chirality-selective emission, but incorporating wavelength-scale localization remains challenging. Existing valley-photonic-crystal cavities rely on finite defects or local lattice modifications that require structure-specific optimization and offer limited continuous control. Here, we theoretically and experimentally demonstrate two-dimensional nanocavity confinement using two orthogonal domain walls in a glide-symmetric valley photonic crystal. A valley domain wall confines the guided interface mode transversely, while an SSH-like domain wall localizes it longitudinally. Starting from a glide-symmetry-protected Dirac point in a bearded-interface waveguide, controlled displacements of adjacent triangular holes open a topological gap in the continuous guided-mode dispersion. The displacement amplitude $\Delta R$ tunes the gap, mode volume, and intrinsic radiative $Q$ factor. Implemented in a silicon photonic-crystal slab, the structure exhibits localized resonances within the topological mode gap and systematic spectral tuning with $\Delta R$. The maximum measured loaded $Q$ factor is $1.2\times10^{4}$. This approach enables continuously tunable, high-$Q$ nanocavities integrated into topological waveguide networks for compact resonant devices and enhanced light--matter interactions.

physics.optics

Wideband Large-Array Processing and Sparse Design for Angle Imaging

This paper shows that wideband large-array processing can recover a large number of angle pixels with far fewer antenna elements. The key advantage of wideband signaling is that different frequencies induce different virtual arrays, whose union forms a virtual array with a substantially increased number of effective virtual elements. Thus, a sparse physical array can support far more spatial samples than physical antennas. Motivated by this capability, we study the recovery of angular responses across the full field of view $[-90^\circ, 90^\circ)$, discretized according to the improved angular resolution, and refer to this sensing regime as angle imaging. However, the resulting virtual array is inherently irregular, clustered, and does not automatically guarantee stable recovery. To address this challenge, we introduce a coverage criterion that estimates the number of stably recoverable angle pixels, without computationally intensive singular-value-based conditioning tests over candidate image dimensions. For systems satisfying this criterion, we theoretically establish deterministic condition-number bounds that characterize stable angle imaging. Building on this criterion, we derive non-uniform sparse array designs that minimize the number of physical antennas while maintaining recovery over the full field of view. Simulation results show that the proposed criterion provides practical guidance for stable system design, and that the resulting sparse arrays can recover substantially more angle pixels than the number of physical antennas, with representative designs supporting over ten times as many angle pixels as physical antennas.

eess.SP

Optimal Rigidity Results for the $k$-Hessian Equation of Lane--Emden Type

In this paper, we establish optimal Liouville theorems and classification results for the \(k\)-Hessian Lane--Emden equation \[ \sigma_k(-D^2u)=u^p\quad\text{in }\R^n,\qquad -D^2u\in\overline{\Gamma_k},\qquad u\geq 0, \] where \(2\leq k<\frac{n}{2}\) and $p>0$. Let $p_- = \frac{nk}{n-2k}$ and the critical Hessian--Sobolev exponent $p_* = \frac{(n+2)k}{n-2k}$. Phuc and Verbitsky proved nonexistence of positive solutions for \(k 2k\), without any additional assumption. For the limiting case \(n=2k\), we classify finite-mass solutions to the $\frac{n}{2}$-Hessian Liouville equation under a proper asymptotic condition $u(x)\rightarrow-\infty$ as $|x|\rightarrow\infty$. In particular, we provide the fully nonlinear counterparts of the classical Liouville and classification theorems of Gidas--Spruck, Gidas--Ni--Nirenberg, and Caffarelli--Gidas--Spruck.

math.AP

Non-radial solutions for the quasi-linear H\'enon type $N$-Laplacian Liouville equation

In this paper, we investigate the following quasi-linear weighted $N$-Laplacian Liouville equation \begin{equation*}\label{0} -\Delta_N u=|x|^{N\alpha}e^{u}, \qquad x\in \R^N, \end{equation*} where $N \geq 2$. For $\al>0$, by carefully studying the linearized problem and applying the approximation method and bifurcation theory, we prove that, when the parameter $\alpha$ equals to the critical values $\alpha(k):=\frac{\sqrt{k(N-1)(k+N-2)}}{N-1}-1$ for $k \geq 2$, there exist non-radial solutions $u$ (bifurcating from $U_{\alpha(k)}$) to the above quasi-linear H\'enon type Liouville equation such that $u\sim \ln|x|$, $|\nabla u|= O(|x|^{-1})$ at $\infty$ and $\int_{\R^N}|x|^{N\alpha}e^{u}\md x=N\left(\frac{N^2}{N-1}\right)^{N-1}(\alpha+1)^{N-1}\omega_N$. One should note that, $\alpha(k)=k-1$ for $k\geq2$ when $N=2$. Our results successfully extend the existence result of J. Prajapat and G. Tarantello in \cite{PT} concerning the $2$-dimension and Laplacian case (i.e., $N=2$) to the more general $N$-dimension and $N$-Laplacian cases ($N\geq 2$), and extend the results of F. Gladiali, M. Grossi, and S. L. N. Neves in \cite{GGN} and the authors in \cite{DDGL} from $1<p<N$ to the much more complicated limiting case $p=N$. We introduced some new ideas and overcame a series of crucial difficulties, including the nonlinearity nature of the $N$-Laplacian $\Delta_N$, the lack of Green integral representation formula and critical weighted Sobolev embedding inequality, the absence of Kelvin type transforms for linearized/difference equations, the invariance of the total mass under scalings of $u$, and the signs-changing and divergence (to $-\infty$) at $\infty$ of the solutions, which makes the suitable choices of the approximate problems, the (normalized) approximate function sequences and the working space to be quite difficult.

math.AP

SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens

Foundation models such as Segment Anything Model 2 (SAM2) have transformed natural-image and video segmentation, and recent work has begun adapting them to medical imaging. These adaptations, however, are largely general-purpose models that treat MRI as one modality among many; large-scale, MRI-specific modelling and benchmarking remain limited, even though MRI's low soft-tissue contrast leaves many boundaries effectively invisible on individual slices. We present SAMRI-3D, a benchmark and method for 3D MRI segmentation with SAM2. The SAMRI-3D benchmark is the largest MRI-only evaluation to date - 10,392 volumes from 34 datasets (27 public, 7 in-house) spanning 12 anatomical domains and 10+ sequences, with explicit seen/unseen splits. Freezing the image encoder and fine-tuning only the lightweight decoder and memory modules raises mean Dice from 0.58 (zero-shot SAM2) to 0.76, surpassing recent SAM-based medical models (SAMed-2 0.69, Medical-SAM2 0.49, SAM-Med3D 0.37) with strong statistical significance. To target invisible boundaries, we introduce Global Volume Tokens (GVT): persistent memory tokens trained with a Truncated Signed Distance Field (TSDF) reconstruction objective that is discarded at inference (zero added cost). This full model, SAMRI-3D, attains the best accuracy (0.78) and lowest variance across all 34 datasets and, uniquely, shows no drop on 8 held-out datasets (0.79 unseen vs. 0.78 seen); per-sequence analysis confirms the TSDF objective helps most where per-slice contrast is weakest. We will release the benchmark, code, and models in this paper.

cs.CV

Constraint-Anchored Reasoning Traces

Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-source MLLMs, once the first error occurs, the reasoning cascades into failure across all remaining steps in 65% of such cases (a metric we term the snowball rate). Existing mitigations-sampling multiple chains, post-hoc self-verification, or full program synthesis-either lack symbolic grounding, catch errors too late, or sacrifice the flexibility of natural language reasoning. We propose Constraint-Anchored Reasoning Traces (CART), a neuro-symbolic framework that trains MLLMs to interleave natural language reasoning steps with symbolic constraint assertions: lightweight, machine-checkable statements about visual content (e.g., count(red_objects) = 3). A dual-pronged Constraint Propagation Module-combining a learned neural grounding head with Boolean Constraint Propagation-continuously verifies these anchors against extracted visual features and checks their mutual logical consistency. When a contradiction is detected, a backtrack controller halts generation and reverts to the last consistent checkpoint, preventing error propagation. A variable-frequency emission mechanism allows the model to adaptively control anchor density, avoiding trace bloat. We construct 218K training instances by augmenting GQA, CLEVR-CoGenT, and VCR with ground-truth constraint annotations derived from scene graphs, and fine-tune open-source MLLMs (LLaVA-NeXT, Qwen2-VL) via LoRA. On five benchmarks, CART reduces the snowball rate from 0.65 to 0.14, improves GQA accuracy by +4.6 percentage points over trainingonly baselines, and achieves 89.1 F1 on POPE-all with at most 18% inference overhead.

cs.AI

Global existence for quasilinear wave equations on hyperbolic space

The main purpose of this paper is to study the global solvability for a general class of quasilinear shifted wave equations on hyperbolic spaces, for smooth initial data with small amplitude. In contrast to the case of Euclidean spaces, when the space dimension is three, we do not need to assume structural conditions like the null conditions to ensure global existence. To achieve this, we establish the energy and local energy estimates for perturbed wave operators on $\mathbb{R}\times \mathbb{H}^n$. These estimates allow time-dependent metric perturbations and require only suitable smallness together with polynomial decay in the radial variable $r$. As a byproduct, for semilinear problems with power-type nonlinearities and radial data, we also obtain global solutions with low-regularity. In particular, we prove an analog of the radial Glassey conjecture on hyperbolic space.

math.AP

Optimizing Pump Conditions of Parametric Amplifiers for Fast Multiplexed Readout of Superconducting Qubits

Low-noise parametric amplifiers are widely used as the first-stage amplifier in qubit readout chains. The performance of parametric amplifiers depends sensitively on the choice of the pump condition. We propose a strategy for determining the pump condition that is tailored for fast multiplexed readout. Choosing the amplifier pump to maximize the signal-to-noise ratio (SNR) improvement at the readout frequency of the limiting qubit--the qubit that requires the longest readout time to reach a target SNR--minimizes the total multiplexed readout time. We demonstrate our pump calibration strategy experimentally on a five-qubit multiplexed readout chain with a traveling-wave parametric amplifier. Using our strategy, we reduce the multiplexed readout time by 320 ns compared to optimizing the average SNR improvement on all qubits, without degrading the target SNR for any qubit.

quant-ph

B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet

EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical and brain-computer interface tasks. Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision. Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained temporal dynamics, we present B[FM]$^2$(Brain Foundation Model via Flow Matching), whose inductive bias aligns with the data by pretraining directly on the raw signal using continuous-time flow matching without patches, tokenization, or masking. However, multi-channel EEG signals pose an architectural challenge for flow matching: time is densely sampled and highly autocorrelated (thousands of timepoints), while the electrode axis is short (tens of channels) at distinct scalp positions. To address this time-electrode asymmetry, we introduce SplitUNet, a velocity network that factorizes each block into separate 1D temporal and 1D electrode convolutions and downsamples only along time, preserving electrode topology throughout the hierarchy. B[FM]$^2$ sets a new state of the art on 7 of 9 standard downstream EEG classification tasks, using a pretraining budget of only 36,895 segments ($\approx$ 307h), 1-2 orders of magnitude ($\approx$ 30x) less than required by existing EEG foundation models. Further, it generates synthetic EEGs that two board-certified neurologists cannot distinguish from brain data (Cohen's $\kappa =$ -0.096). https://jd730.github.io/projects/BFM2

cs.LG

Readout-Induced Leakage in Superconducting Circuits with Nonlinear Couplings

In superconducting circuits, drive-induced unwanted transitions limit the readout power, thereby constraining readout speed and fidelity. When such transitions excite the qubit into leakage states, they produce correlated errors that are particularly harmful for quantum error correction. Native nonlinear qubit-readout resonator coupling is a promising alternative to conventional linear hybridization because it provides intrinsic Purcell protection and stricter selection rules for multiphoton processes. In realistic devices, however, we show that such a coupling alone neither eliminates nor necessarily suppresses drive-induced transitions. Instead, if not appropriately engineered, these couplings often worsen the situation by introducing additional parasitic processes. Moreover, the rates of these unwanted transitions remain sensitive to the choice of readout frequency, regardless of the coupling mechanism. We demonstrate that readout-induced leakage can thus vary by orders of magnitude even when readout frequencies differ by less than ~7%. Our results establish that the benefits of native nonlinear couplings are realized only through informed device design, including the spectral placement of relevant auxiliary modes and elimination of parasitic ones.

quant-ph

Critical quasi-linear Schr\"{o}dinger system with $p$-Laplacian

In this paper, we mainly consider positive solution to the $D^{1,p}(\R^{N})$-critical quasi-linear Schr\"{o}dinger system with $p$-Laplacian: \begin{equation*}\begin{cases} -\Delta_p u = u^{\alpha}v^{\beta} \, \ \ \ \ \ \text{in}\,\ \ \R^N, \\ -\Delta_p v = u^{\beta}v^{\alpha} \,\ \ \ \ \ \text{in}\,\ \ \R^N, \end{cases}\end{equation*} where $1<p<N$, $N\geq2$, $0\leq \alpha \leq \beta,$ and $u,v\in D^{1,p}(\R^N)$. We establish regularity and the sharp estimates on asymptotic behaviors for any positive solution $(u,v)$. Then, we prove that all positive solutions are radially symmetric and strictly decreasing about some point. Furthermore, we obtain the uniqueness and complete classification of positive solutions. Our results extend the uniqueness results in \cite{LM,QS} for $p=2$ to general cases $1<p<N$.

math.AP

CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that replaces a Namaste with a handshake receives the same score as one that generates the gesture correctly. We propose CultureScore, a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions: Identity (who is represented), Context (culturally localized background), and Behavior (normative gestures and interactions). We operationalize this framework through an evaluation suite spanning 10 countries, yielding 6,174 generated videos across three state-of-the-art models. Our evaluation reveals that no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8\% overall CultureScore, with Behavior the most challenging dimension, which remains below 52.1\% across all models. Furthermore, human preference rankings align directionally with CultureScore but are inverted relative to VideoScore; the highest-scoring model on visual quality was ranked last by annotators, underscoring that cultural faithfulness is an essential criterion for equitable video generation. Data and code are available at https://huggingface.co/datasets/ankurani/CultureScore.

cs.CV

Equivariant Neural Belief Propagation

Probabilistic inference over spatially embedded variables requires beliefs that respect $SE(3)$ symmetry, yet existing equivariant networks produce only scalars and vectors -- not the rank-2 precision tensors needed for anisotropic uncertainty, and single-component messages collapse multi-modal energy landscapes to physically meaningless averages. We introduce Equivariant Neural Belief Propagation (ENBP), a factor-graph framework whose messages are equivariant Gaussian mixture models with sufficient statistics that transform exactly under $SE(3)$. Rank-2 precision matrices are synthesised via equivariant outer products, ingested through differentiable spectral decomposition, and kept tractable by a greedy KL-based mixture reduction that provably commutes with $SE(3)$. On GEOM-QM9 and GEOM-Drugs, ENBP achieves 98.9% conformational coverage at 0.090 $\mathring{A}$ error with sub-second latency -- over $100\times$ faster than diffusion baselines at higher accuracy. On multi-body robotic inference, vanilla loopy BP diverges at 15+ agents while ENBP converges with near-zero collision rates and machine-precision equivariance error (${\sim}10^{-7}$ vs.\ $10^{-1}$ for augmented baselines).

cs.LG

Invariant Gradient Alignment for Robust Reasoning Distillation

Large language models (LLMs) suffer from shortcut learning: they systematically fail on out-of-distribution (OOD) inputs whose semantic surface differs from training data, even when the logical structure is identical. This undermines knowledge distillation pipelines that transfer chain-of-thought reasoning to smaller students. We introduce Invariant Gradient Alignment (IGA), a training framework that aligns gradient updates across semantically diverse but logically isomorphic examples via three innovations: (i) Logical Isomer Sets, groups of problems sharing identical logical structure across distinct semantic domains (mathematics, medicine, law, science); (ii) a differentiable \emph{Continuous Gradient Conflict Mask}, that suppresses parameter dimensions with high cross-domain gradient variance while preserving invariant directions; and (iii) a truncated SVD projection of the masked gradient back onto the LoRA low-rank manifold, maintaining parameter efficiency throughout. Theoretically, IGA yields tighter OOD generalization bounds than ERM, scaling with the number of isomer domains, and converges at the standard SGD rate under mild regularity. Empirically, IGA outperforms eight baselines across four benchmarks with accuracy gains up to 14.3 pp over ERM-SFT and a Logical Consistency Score of 0.031 versus 0.142 -- a fourfold improvement in representational invariance.

cs.LG

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair

Autoregressive chain-of-thought (CoT) reasoning in large language models (LLMs) is fundamentally forward-directed: each step conditions only on prior tokens. This unidirectional inductive bias renders even capable models susceptible to error snowballing, wherein a single logical or arithmetic mistake in an early step irreversibly corrupts the entire reasoning chain. We introduce Teleological Reasoning Infilling (\TRI{}), a training framework that endows decoder-only transformers with a native \emph{goal-conditioned bridging} capability. The key insight is to reframe erroneous reasoning segments as fill-in-the-middle (FIM) tasks: given a verified prefix premise $P$, a verified downstream milestone $S$, and the original query $Q$, the model must synthesise the logical bridge $M$ that connects $P$ to $S$ rigorously and completely. To achieve this with standard causal architectures, we introduce a Prefix-Suffix-Middle (PSM) sequence rearrangement with three non-overlapping sentinel tokens, enabling $M$ to attend to both $P$ and $S$ without any structural modification to the self-attention mechanism. Training proceeds in two stages: (i) Supervised Fine-Tuning (SFT) on symbolically verified $(P, S, M)$ triples extracted from formal mathematics corpora, and (ii) Direct Preference Optimisation (DPO) with a deterministic symbolic verifier (Lean 4 / Python) as the sole reward oracle, eliminating LLM-judge sycophancy. At inference, TRI operates as a surgical repair module within a dual-system loop: a causal draft model generates an initial trace, the verifier pinpoints failures, and TRI infills only the damaged segment, leaving verified sections intact. Comprehensive experiments on three benchmarks demonstrate that TRI achieves state-of-the-art performance across all tasks, while reducing per-problem token expenditure by 31.2%.

cs.CL

In-Context Graphical Inference

Marginal inference in discrete graphical models forces a choice between exactness and scalability: exact algorithms are intractable for high-treewidth graphs, while iterative approximations (Belief Propagation, variational methods) sacrifice convergence guarantees on frustrated topologies. We argue that this dichotomy stems from a mismatched inductive bias: iterative methods abandon the sequential elimination structure that makes exact inference correct. We introduce In-Context Graphical Inference (ICG-I), an autoregressive Graph Transformer that restores this structure by mimicking Variable Elimination with learned, Tensor- Train-compressed intermediate factors, paired with a Dirichlet output layer and Weighted Conformal Prediction for calibrated, distribution-free coverage guarantees under topological shift. We prove that TT compression errors propagate at most lincarly through the autoregressive chain, that the Dirichlet-Multinomial loss is a proper scoring rule, and that WCP maintains coverage with a quantifiable degradation under estimated density ratios. We conducted intensive experiments to evaluate ICG-I and achieved state-of-the-art performance across all benchmarks. ICG-I reduces MAE from 0.041 (best baseline) to 0.020 on standard instances and achieves 0.048 on N=500 frustrated spin glasses where BP diverges entirely.

cs.LG

$D^0$-$D_s^+$ Elliptic-Flow Splitting under Event-Shape Engineering: A Probe of Sequential Charm Hadronization

Recent work has proposed sequential hadronization of open-charm hadrons in the quark-gluon plasma, wherein more tightly bound species such as $D_s^+$ form earlier near $1.2 T_c$ and $D^0$ forms later at $T_c$. That work showed that this mechanism naturally reverses the sign of the $D^0-D_s^+$ elliptic-flow splitting relative to the conventional simultaneous baseline. In this work, we demonstrate that event-shape engineering (ESE) provides a sharper discrimination between the two pictures than inclusive measurements alone. By selecting large-$q_2$ and small-$q_2$ events in 0--10\% and 30--50\% centrality classes in Pb-Pb collisions at $\sqrt{s_{\mathrm{NN}}}=5.02$ TeV, we show that the geometry-driven enhancement of charm-meson $v_2$ can be separated from the hadronization-time response: the positive $\Delta v_2(D^0-D_s^+)$ in the sequential scenario grows systematically with $q_2$, while the corresponding response slope $\chi$ reveals a species-dependent hierarchy $\chi(D^0) > \chi(D_s^+)$ that is robust against the overall flow normalization and absent in the simultaneous baseline. In the simultaneous case, the splitting is near zero or negative and does not follow the same geometry scaling. Notably, the semi-central 30--50\% class emerges as the optimal window, because the non-monotonic interplay between QGP lifetime and initial eccentricity maximizes the late-stage flow conversion. The $q_2$ ratios of the $D_s^+/D^0$ yield ratio remain close to unity, confirming that the splitting is a dynamical flow effect rather than a chemical yield modification. These results establish $\Delta v_2(D^0-D_s^+)$ and the response slope $\chi$ under ESE as complementary differential probes of the space-time structure of charm hadronization near the QCD transition temperature.

hep-ph

Deep Psychovisual Image Representations

Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In contrast, deep learning-based vision models routinely extract and aggregate features using homogeneous stacks of spatial layers, rendering their decision-making processes opaque. In this paper, we propose Deep Visual Coding, a learned frequency-domain representation inspired by 1990s image codes that quantised perceptually salient frequencies, which together with complex-valued image representations produces psychovisual-style abstractions. This approach enables the first psychovisual-based deep learning framework, utilizing data-driven spectral filters that learn to encode task-relevant semantic structures within distinct frequency sub-bands. Salience analyses reveal that our psychovisual models extract highly interpretable object parts compared to the amorphous regions produced by regular Convolutional Neural Networks (CNNs). Furthermore, we find that our models are less depth dependent than CNNs for model scaling, since our complex-valued representations and learned abstractions subsume the role of the deep spatial layers. Together, these findings demonstrate that psychovisual coding provides a promising path toward more efficient and transparent vision models.

cs.CV