SearcharxivSearch

arXiv subjects

Kangning Liu

Publications and source records attributed to Kangning Liu.

At least 19 recordsLinked to original sources

ABJM/BMN Index Matching in the Plane-Wave Limit

We establish coefficientwise equality between the refined supersymmetric indices of ABJM theory and the BMN matrix model in the plane-wave limit at every fixed positive Chern-Simons level k. For multi-membrane families with fixed multiplicities, and hence fixed finite ABJM gauge rank, the equality holds as all longitudinal momenta tend to infinity with their pairwise differences fixed. Scalar-dressed equal-flux ABJM monopoles correspond to BMN vacua of fuzzy-sphere membranes: multiplicities count coincident membranes, whereas flux magnitudes give the momentum per membrane along the M-theory circle. At the index level, the associated Higgs residues yield a diagonal-holonomy integral for the ABJM plane-wave contribution, whose monopole-harmonic kernel agrees with the BMN matrix-harmonic kernel up to the finite BMN upper-spin cutoff. At finite momentum, the full index difference splits into the non-plane-wave remainder and the harmonic tail above this cutoff. Increasing k shifts the former to higher fugacity grade without changing the surviving kernel or the cutoff contribution. This gives a coefficientwise, multi-membrane extension of the earlier single-membrane index relation.

hep-th

Scaling Muon for Diffusion Transformers

The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establish Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing that its optimization and generative quality advantages over AdamW persist across model scales. However, at scale, the 5-step Newton--Schulz iteration (NS5) performed at every optimization step, together with full-momentum materialization, introduces substantial computation and communication overhead that can offset Muon's step-efficiency advantage. We introduce \emph{Periodic Row-wise Muon}, which performs a full NS5 spectral update once every \(K\) steps and applies a low compute and communication cost row-wise constrained update based on the current momentum at the remaining steps. We further co-design a distributed implementation that operates directly on sharded momentum during non-refresh steps and accelerates spectral refreshes through bucketed all-gather and communication--computation overlap. Across all scales, Muon improves the best observed generative quality over AdamW by 12.9--19.1\%. Compared with vanilla Muon, Periodic Row-wise Muon remains within 0.5\% in best generative quality on the 1.3B--4B models and improves it by 4.5\% at 9B. It reduces optimizer time by 46.9--54.3\%, end-to-end step time by 15.7--24.3\%, and logical communication volume by 66.7\%, while reaching its respective best generative quality with 33.7--64.8\% less active training time. These results show that Periodic Row-wise Muon preserves Muon's generative quality advantage while translating it into end-to-end training efficiency for large DiTs.

cs.LG

Mass-Flow Invariance of $Q$-Cohomology in BMN Matrix Quantum Mechanics

We study the dependence of the dynamical supercharges of BMN matrix quantum mechanics on the mass parameter $\mu$. Taking the $\mu$-derivative at fixed canonical matrix variables, we show that the sixteen-component supercharge evolves by the adjoint action of a Hermitian quadratic bosonic operator $\mathcal{K}$, together with the spinor-space factor $i\gamma^{123}$. After projection to a $\gamma^{123}$-eigenspace, this flow integrates to a finite similarity transformation. For the nilpotent component $Q(\mu)=\mathcal Q^4_-(\mu)$, one obtains $Q(\mu)=M(\mu,\mu_0)Q(\mu_0)M(\mu,\mu_0)^{-1}$, giving an algebraic mass-flow non-renormalization statement for the $Q$-cohomology. The corresponding Hilbert-space statement has an analytic qualification, parallel to Witten's argument for supersymmetric quantum mechanics: $M$ is non-unitary and unbounded, so its action on the normalizable domain must be controlled. We formulate a small-step criterion by comparing the quadratic growth of $M$ with the Gaussian falloff of BMN oscillator wavefunctions within each component $\mu>0$ or $\mu<0$. As a concrete check, we evaluate this condition in the $N=2$ theory, whose two vacuum sectors are built on the trivial vacuum and the irreducible fuzzy-sphere vacuum. We also compute the induced $Q_{\rm BPS}$-action on the corresponding BPS letters: in the trivial sector it agrees with the standard BMN-sector BPS-letter differential of $\mathcal{N}=4$ SYM, while in the irreducible sector it vanishes.

hep-th

MAOAM: Unified Object and Material Selection with Vision-Language Models

Selection is a core operation in interactive image editing. To be practical, a user should be able to specify and disambiguate the desired selection region through either text or click-based interactions, and the system should support selecting not only objects but also other criteria, such as materials. Material-based selection is valuable for tasks like re-texturing surfaces or editing instances of a specific material. However, existing vision-language-model (VLM) based selection methods are object-centric and typically support a single interaction modality, limiting their applicability. In this work, we thus present Mask Any Object And Material (MAOAM), a unified selection framework that enables precise object and material-level selection across both text- and click-based interactions. MAOAM leverages a VLM with a segmentation head to produce pixel-accurate masks from user prompts: the VLM interprets the user's selection intent (object or material-level) and encodes visual entities, attributes, and spatial relations, while the segmentation head decodes the output token into a mask. A key challenge is the lack of material selection datasets with text annotations. We propose a scalable data generation pipeline: we collect real and synthetic images with material masks, and leverage VLMs to generate material descriptions with rich visual-semantics. We train MAOAM with a multi-task objective over click and text-based selection, along with an auxiliary VQA task derived from the material descriptions to facilitate deeper material understanding. Despite being trained with uni-modal prompts, our model exhibits an emergent improvement in selection when combining text and clicks at inference, enabling flexible image editing workflows. Experiments demonstrate accurate and coherent selections across diverse objects, materials, and interaction scenarios, highlighting robustness in practice.

cs.CV

Finite-$N$ BMN index across all vacuum sectors

We compute the finite-$N$ Witten index of BMN matrix quantum mechanics after summing over all partition-labeled supersymmetric vacuum sectors. Starting from the unitary-matrix integral for each sector, we develop two complementary evaluation methods: a symmetric-group character expansion, which reduces each fixed fugacity order to a finite combinatorial sum, and a residue expansion in which the contributing poles are organized by rooted trees, with a colored-tree generalization for multi-partition sectors. Where practical, direct integration and extraction of the constant term in the expanded integrand give independent coefficient-by-coefficient checks. We evaluate every vacuum sector for $N\leq 9$. In the equal-fugacity expansion, the coefficients near charges $j\sim N^2$ show entropy growth of order $N^2$, and, in this range, the sector sum does not cancel this growth. The finite-$N$ data also reveal a nontrivial sectoral organization: near $j=N^2$, the sector giving the largest contribution changes with $N$, from single-partition sectors at small rank to double-partition sectors starting at $N=5$. We call this phenomenon dominance switching. These results provide quantitative finite-$N$ input for using the BMN index as a diagnostic of protected plane-wave black-hole sectors and suggest a D2 dressed black-hole interpretation in the controlled type-IIA regime, where D0 black-hole sectors are accompanied by macroscopic spherical D2-brane degrees of freedom, analogous to dual dressed black holes in $AdS_5\times S^5$.

hep-th

$J\bar{J}$-deformation as a Riemann bilinear dressing

We propose a reformulation of the conformal perturbation theory of the correlation functions in $J\bar{J}$-deformed CFTs as a dressing on the deformed operators, that matches both bare and renormalized perturbation theory. The key is to use the Riemann bilinear identity to convert the deformation into a dressing and a large-cycle integral for higher genus. Based on the proposal, we calculate the deformation of partition functions on the torus and higher genus Riemann surfaces, which can be written as kernel integrals that preserve modular invariance or covariance. We also calculate the flow of the conformal weights and conserved charges along the deformation. Based on this flow and the modular $S$-transformation, we propose a criterion for constructing dressed operators. We test our formalism and results by studying the $O(2, 2)$ theories and strings on the TsT background.

hep-th

Inline Critic Steers Image Editing

Instruction-based image editing exhibits heterogeneous difficulty not only across cases but also across regions of an image, motivating refinement approaches that allocate correction to where the model struggles. Existing refinement signals arrive late, after a fully generated image or a completed denoising step. We ask whether such a signal can act within an ongoing forward pass. To investigate this, we probe a frozen image-editing model and find that although generation capability emerges only in the last few layers, the error pattern is already set in early layers (rank correlation \r{ho} = 0.83 with the final-layer error map). Based on this, we introduce Inline Critic, a learnable token that critiques a frozen model's predictions at its intermediate layers and steers its hidden states to refine generation during the forward pass. A three-stage recipe is proposed to stabilize the training from learning how to critique to steering generation. As a result, we achieve state of the art on GEdit-Bench (7.89), a +9.4 gain on RISEBench over the same backbone, and the strongest open-source result on KRIS-Bench (81.92, surpassing GPT-4o). We further provide analyses showing that the critic genuinely shapes the model's attention and prediction updates at subsequent layers.

cs.CV

SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation

Recent advancements in discrete image generation showed that scaling the VQ codebook size significantly improves reconstruction fidelity. However, training generative models with a large VQ codebook remains challenging, typically requiring larger model size and a longer training schedule. In this work, we propose Stochastic Neighbor Cross Entropy Minimization (SNCE), a novel training objective designed to address the optimization challenges of large-codebook discrete image generators. Instead of supervising the model with a hard one-hot target, SNCE constructs a soft categorical distribution over a set of neighboring tokens. The probability assigned to each token is proportional to the proximity between its code embedding and the ground-truth image embedding, encouraging the model to capture semantically meaningful geometric structure in the quantized embedding space. We conduct extensive experiments across class-conditional ImageNet-256 generation, large-scale text-to-image synthesis, and image editing tasks. Results show that SNCE significantly improves convergence speed and overall generation quality compared to standard cross-entropy objectives.

cs.CV

Corrections of an elliptic block in the NS sector

We propose a correction to one of the elliptic blocks in the NS sector of 2d $\mathcal N = 1$ superconformal field theories. We analyze the 4-point block in the pillow geometry to demonstrate the necessity of the correction and verify the formula by numerically checking the crossing symmetries in the $\mathcal N =1 $ super Liouville theory, as well as directly comparing the $c$-recursion and $h$-recursion results.

hep-th

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.

cs.CV

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, their inference speed remains suboptimal due to the need to repeatedly process redundant masked tokens at every sampling step. In this work, we propose Sparse-LaViDa, a novel modeling framework that dynamically truncates unnecessary masked tokens at each inference step to accelerate MDM sampling. To preserve generation quality, we introduce specialized register tokens that serve as compact representations for the truncated tokens. Furthermore, to ensure consistency between training and inference, we design a specialized attention mask that faithfully matches the truncated sampling procedure during training. Built upon the state-of-the-art unified MDM LaViDa-O, Sparse-LaViDa achieves up to a 2x speedup across diverse tasks including text-to-image generation, image editing, and mathematical reasoning, while maintaining generation quality.

cs.CV

VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

Current visual grounding models are either based on a Multimodal Large Language Model (MLLM) that performs auto-regressive decoding, which is slow and risks hallucinations, or on re-aligning an LLM with vision features to learn new special or object tokens for grounding, which may undermine the LLM's pretrained reasoning ability. In contrast, we propose VGent, a modular encoder-decoder architecture that explicitly disentangles high-level reasoning and low-level bounding box prediction. Specifically, a frozen MLLM serves as the encoder to provide untouched powerful reasoning capabilities, while a decoder takes high-quality boxes proposed by detectors as queries and selects target box(es) via cross-attending on encoder's hidden states. This design fully leverages advances in both object detection and MLLM, avoids the pitfalls of auto-regressive decoding, and enables fast inference. Moreover, it supports modular upgrades of both the encoder and decoder to benefit the whole system: we introduce (i) QuadThinker, an RL-based training paradigm for enhancing multi-target reasoning ability of the encoder; (ii) mask-aware label for resolving detection-segmentation ambiguity; and (iii) global target recognition to improve the recognition of all the targets which benefits the selection among augmented proposals. Experiments on multi-target visual grounding benchmarks show that VGent achieves a new state-of-the-art with +20.6% F1 improvement over prior methods, and further boosts gIoU by +8.2% and cIoU by +5.8% under visual reference challenges, while maintaining constant, fast inference latency.

cs.CV

$\mathcal{N}=1$ super complex Liouville string

We study the (type 0B) $\mathcal{N}=1$ supersymmetric complex Liouville string ($\text{S}\mathbb{C}\text{LS}$), a supersymmetric extension of the bosonic complex Liouville string ($\mathbb{C}\text{LS}$). We compute the sphere three-point amplitudes (including NS-NS-NS and NS-R-R types) and find they share the same form as the sphere three-point amplitude of the bosonic $\mathbb{C}\text{LS}$. Analysis of the analytic structure of the NS-NS-NS-NS four-point amplitude and the higher equations of motion also yields results identical to the bosonic case. Based on these findings, we propose that the dual matrix model for the $\text{S}\mathbb{C}\text{LS}$ is the same as that for the bosonic $\mathbb{C}\text{LS}$. We also investigate a related theory $\widehat{\text{S}\mathbb{C}\text{LS}}$, which differs in the gauged worldsheet supersymmetry. A parallel analysis is performed for $\widehat{\text{S}\mathbb{C}\text{LS}}$, and a candidate for its dual matrix model is proposed. We then carry out a partial numerical evaluation of the moduli space integral, which provides further evidence for both the proposals of the dual matrix model regarding $\text{S}\mathbb{C}\text{LS}$ and $\widehat{\text{S}\mathbb{C}\text{LS}}$.

hep-th

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution image generation, Lavida-O presents a single framework that enables image-level understanding, object grounding, image editing, and high-resolution (1024px) text-to-image synthesis. Lavida-O incorporates a novel Elastic Mixture-of-Transformers (Elastic-MoT) architecture that couples a lightweight generation branch with a larger understanding branch, supported by token compression, universal text conditioning and stratified sampling for efficient and high-quality generation. Lavida-O further incorporates planning and iterative self-reflection in image generation and editing tasks, seamlessly boosting generation quality with its understanding capabilities. Lavida-O achieves state-of-the-art performance on a wide range of benchmarks including RefCOCO object grounding, GenEval text-to-image generation, and ImgEdit image editing, outperforming existing autoregressive models and continuous diffusion models such as Qwen2.5-VL and FluxKontext-dev, while offering considerable speedup at inference. These advances establish Lavida-O as a new paradigm for scalable multimodal reasoning and generation.

cs.CV

Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration

Autonomous agents for desktop automation struggle with complex multi-step tasks due to poor coordination and inadequate quality control. We introduce Agentic Lybic, a novel multi-agent system where the entire architecture operates as a finite-state machine (FSM). This core innovation enables dynamic orchestration. Our system comprises four components: a Controller, a Manager, three Workers (Technician for code-based operations, Operator for GUI interactions, and Analyst for decision support), and an Evaluator. The critical mechanism is the FSM-based routing between these components, which provides flexibility and generalization by dynamically selecting the optimal execution strategy for each subtask. This principled orchestration, combined with robust quality gating, enables adaptive replanning and error recovery. Evaluated officially on the OSWorld benchmark, Agentic Lybic achieves a state-of-the-art 57.07% success rate in 50 steps, substantially outperforming existing methods. Results demonstrate that principled multi-agent orchestration with continuous quality control provides superior reliability for generalized desktop automation in complex computing environments.

cs.AI

Symmetries and operators in $T\bar{T}$ deformed CFTs

$T\bar{T}$-deformed CFTs are known to possess nonlocal conformal symmetries that do not act tractably on the undeformed local operators. In this paper, we explicitly construct two distinct classes of operators: (i) dressed operators, which are primary operators with respect to the nonlocal conformal symmetries, and (ii) physical operators, a new type of local operator we introduce. While the dressed operators preserve the conformal symmetry structure, they are themselves nonlocal. The physical operators, by contrast, are local and can be expressed in terms of the dressed operators. We calculate the two-point correlation functions of these operators in momentum space and find that our results align with both string theory predictions and field theory calculations. Additionally, we explore the relationship between physical operators and alternative operator definitions proposed in the literature.

hep-th

Refer to Any Segmentation Mask Group With Vision-Language Prompts

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io.

cs.CV

Multi-modal AI for comprehensive breast cancer prognostication

Treatment selection in breast cancer is guided by molecular subtypes and clinical characteristics. However, current tools including genomic assays lack the accuracy required for optimal clinical decision-making. We developed a novel artificial intelligence (AI)-based approach that integrates digital pathology images with clinical data, providing a more robust and effective method for predicting the risk of cancer recurrence in breast cancer patients. Specifically, we utilized a vision transformer pan-cancer foundation model trained with self-supervised learning to extract features from digitized H&E-stained slides. These features were integrated with clinical data to form a multi-modal AI test predicting cancer recurrence and death. The test was developed and evaluated using data from a total of 8,161 female breast cancer patients across 15 cohorts originating from seven countries. Of these, 3,502 patients from five cohorts were used exclusively for evaluation, while the remaining patients were used for training. Our test accurately predicted our primary endpoint, disease-free interval, in the five evaluation cohorts (C-index: 0.71 [0.68-0.75], HR: 3.63 [3.02-4.37, p<0.001]). In a direct comparison (n=858), the AI test was more accurate than Oncotype DX, the standard-of-care 21-gene assay, achieving a C-index of 0.67 [0.61-0.74] versus 0.61 [0.49-0.73], respectively. Additionally, the AI test added independent prognostic information to Oncotype DX in a multivariate analysis (HR: 3.11 [1.91-5.09, p<0.001)]). The test demonstrated robust accuracy across major molecular breast cancer subtypes, including TNBC (C-index: 0.71 [0.62-0.81], HR: 3.81 [2.35-6.17, p=0.02]), where no diagnostic tools are currently recommended by clinical guidelines. These results suggest that our AI test improves upon the accuracy of existing prognostic tests, while being applicable to a wider range of patients.

cs.AI