Searcharxiv⌕ Search

arXiv subjects

Ming Lu

Publications and source records attributed to Ming Lu.

At least 37 records · Page 2Linked to original sources

Fast Electromagnetic and RF Circuit Co-Simulation for Passive Resonator Field Calculation and Optimization in MRI

Passive resonators have been widely used in MRI to manipulate RF field distributions. However, optimizing these structures using full-wave electromagnetic simulations is computationally prohibitive, particularly for large passive resonator arrays with many degrees of freedom. This work presents a co-simulation framework tailored specifically for the analysis and optimization of passive resonators. The framework performs a single full-wave electromagnetic simulation in which the resonator's lumped components are replaced by ports, followed by circuit-level computations to evaluate arbitrary capacitor and inductor configurations. This allows integration with a genetic algorithm to rapidly optimize resonator parameters and enhance the B1 field in a targeted region of interest (ROI). We validated the method in three scenarios: (1) a single-loop passive resonator on a spherical phantom, (2) a two-loop array on a cylindrical phantom, and (3) a two-loop array on a human head model. In all cases, the co-simulation results showed excellent agreement with full-wave simulations, with relative errors below 1%. The genetic-algorithm-driven optimization, involving tens of thousands of capacitor combinations, completed in under 5 minutes, whereas equivalent full-wave EM sweeps would require an impractically long computation time. This work extends co-simulation methodology to passive resonator design for first time, enabling the fast, accurate, and scalable optimization. The approach significantly reduces computational burden while preserving full-wave accuracy, making it a powerful tool for passive RF structure development in MRI.

physics.med-ph↗

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under partial observations. Reliable grasping depends on both local contact cues and object-level 3D structure. Existing geometry-aware grasping methods recognize the value of reconstruction, but they typically treat geometry as an intermediate prediction rather than a reusable object prior for grasping. In this paper, we present GraspFoM, a unified framework that leverages 3D foundation priors (SAM3D) to build a shared 3D object latent for both reconstruction and grasp pose prediction. Built on this shared object latent, we introduce an anchor-initialized truncated pose-reasoning diffuser that predicts continuous and multimodal grasp poses without directly relying on discrete grasp candidates. We further investigate the interaction between reconstruction and grasping through a reconstruction-aware scorer and a residual latent updater. Reconstruction provides grounded geometric cues, while grasp supervision refines the shared object latent toward grasp-relevant affordances. GraspFoM jointly predicts grasp poses and reconstructs high-fidelity 3D assets in mesh and 3DGS forms. Comprehensive experiments demonstrate that GraspFoM achieves state-of-the-art results on both reconstruction and grasping. Notably, these improvements require only a small number of additional trainable parameters. Component-wise ablation studies also demonstrate the contribution of each component.

cs.RO↗

SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation

While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/.

cs.CV↗

Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.

cs.CV↗

Quiver varieties and dual canonical bases

We survey some recent developments on the theory of dual canonical bases for quantum groups and $\imath$quantum groups. The $\imath$quiver algebras were introduced by Wang and the first author, which are used to give two realizations of quasi-split $\imath$quantum groups of type ADE: one via the $\imath$Hall algebras and the other via the quantum Grothendieck rings of Nakajima-Keller-Scherotzke quiver varieties. The geometric construction of the $\imath$quantum groups produces their dual canonical bases with positivity, generalizing Qin's geometric realization of quantum groups of type ADE. Recently, the authors provided a new construction of the dual canonical basis in the setting of $\imath$Hall algebras, and proved that it is invariant under braid group actions, and obtained the positivity of the transition matrix coefficients from the Hall basis to the dual canonical basis. As quantum groups can be regarded as $\imath$quantum groups of diagonal type, we demonstrate that the dual canonical bases of quantum groups coincide with the double canonical bases defined by Berenstein and Greenstein, and resolve several conjectures therein.

math.QA↗

Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent

Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limitations in dynamic SQL generation, and the need for deep multi-dimensional analysis.In this paper, we propose AIDA(Autonomous Insight Discovery Agent), the first end-to-end framework designed for autonomous exploration in complex business environments. We establish a highly flexible instant retail environment encompassing 200+ metrics and 100+ dimensions, and integrates a proprietary Domain-Specific Language (DSL) that bridges semantic reasoning with precise SQL execution. Our reinforcement learning system subsequently formulates business analysis as a Pareto Principle-guided cumulative reasoning process. Experimental results demonstrate that AIDA significantly outperforms workflow-based agents, and extensive evaluations further reveal that AIDA achieves superior environmental perception and more in-depth analysis from diverse perspectives. Our work ultimately establishes the transformative potential of autonomous intelligence for industrial-scale business intelligence systems.

cs.AI↗

Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning

Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token-level alignment via supervised fine-tuning, which fails to fully capture the potential synergy between comprehension and creation. In this work, we propose Sync-R1, an end-to-end reinforcement learning framework that jointly optimizes personalized understanding and generation within a single, explicit reasoning loop. Through this unified feedback process, Sync-R1 enables personalized comprehension to guide content creation, while the resulting generation quality reciprocally refines understanding within an integrated reward landscape. To efficiently orchestrate this dual-task synergy, we introduce Sync-GRPO, a reinforcement learning method utilizing an ensemble reward system. Furthermore, we propose Dynamic Group Scaling (DGS), which adaptively filters low-potential trajectories to reduce gradient variance and accelerate convergence. To better reflect real-world complexity, we introduce UnifyBench++, featuring denser textual descriptions and richer user contexts. Experimental results demonstrate that Sync-R1 achieves state-of-the-art performance, showcasing superior cross-task reasoning and robust personalization without requiring complex cold-start procedures. The code and the UnifyBench++ dataset will be released at: https://github.com/arctanxarc/UniCTokens.

cs.CV↗

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in representations that lack accurate spatial awareness. Existing implicit spatial grounding methods partially address this by aligning VLA features with those of 3D-aware foundation models, but they rely on empirical layer search and perform alignment on LLM-level visual tokens where spatial structure has already been entangled with linguistic semantics, limiting both generalizability and geometric interpretability. We propose VEGA (Visual Encoder Grounding Alignment), a simple yet effective framework that directly aligns the output of the VLA's visual encoder with spatially-aware features from DINOv2-FiT3D, a DINOv2 model fine-tuned with multi-view consistent 3D Gaussian Splatting supervision. By performing alignment at the visual encoder output level, VEGA grounds spatial awareness before any linguistic entanglement occurs, offering a more interpretable and principled alignment target. The alignment is implemented via a lightweight projector trained with a cosine similarity loss alongside the standard action prediction objective, and is discarded at inference time, introducing no additional computational overhead. Extensive experiments on simulation benchmark and real-world manipulation tasks demonstrate that VEGA consistently outperforms existing implicit spatial grounding baselines, establishing a new state-of-the-art among implicit spatial grounding methods for VLA models.

cs.RO↗

Growth rates of indecomposable summands in tensor powers of representations of quivers

Tensor products of quiver representations have been extensively studied; typical examples include the pointwise tensor product and the tensor product induced by the coalgebra structure of path algebras. In this paper, we investigate the growth rates of the number of indecomposable direct summands in tensor powers of quiver representations with respect to these two typical tensor products.

math.RT↗

Client-Verifiable and Efficient Federated Unlearning in Low-Altitude Wireless Networks

In low-altitude wireless networks (LAWN), federated learning (FL) enables collaborative intelligence among unmanned aerial vehicles (UAVs) and integrated sensing and communication (ISAC) devices while keeping raw sensing data local. Due to the "right to be forgotten" requirements and the high mobility of ISAC devices that frequently enter or leave the coverage region of UAV-assisted servers, the influence of departing devices must be removed from trained models. This necessity motivates the adoption of federated unlearning (FUL) to eliminate historical device contributions from the global model in LAWN. However, existing FUL approaches implicitly assume that the UAV-assisted server executes unlearning operations honestly. Without client-verifiable guarantees, an untrusted server may retain residual device information, leading to potential privacy leakage and undermining trust. To address this issue, we propose VerFU, a privacy-preserving and client-verifiable federated unlearning framework designed for LAWN. It empowers ISAC devices to validate the server-side unlearning operations without relying on original data samples. By integrating linear homomorphic hash (LHH) with commitment schemes, VerFU constructs tamper-proof records of historical updates. ISAC devices ensure the integrity of unlearning results by verifying decommitment parameters and utilizing the linear composability of LHH to check whether the global model accurately removes their historical contributions. Furthermore, VerFU is capable of efficiently processing parallel unlearning requests and verification from multiple ISAC devices. Experimental results demonstrate that our framework efficiently preserves model utility post-unlearning while maintaining low communication and verification overhead.

cs.CR↗

Singular limits for non-isentropic compressible rotating fluids

In this article, we study the singular limit of non-isentropic compressible rotating fluids. We incorporate the capillary effect into both the $α=1$ and $α=0$ cases, and investigate the Navier-Stokes-Korteweg equations involving the terms of low Mach number, low Rossby number and high Reynolds number. When $α=1$, the dispersion estimate of the acoustic wave equation is derived by Rage's theorem. When $α=0$, we obtain the convergence results by error estimate. Moreover, we obtain that the three dimensions compressible Navier-Stokes-Korteweg equations converge to the two dimensions incompressible Euler equations.

math.AP↗

DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression

Diffusion-based image compression has recently shown outstanding perceptual fidelity, yet its practicality is hindered by prohibitive sampling overhead and high memory usage. Most existing diffusion codecs employ U-Net architectures, where hierarchical downsampling forces diffusion to operate in shallow latent spaces (typically with only 8x spatial downscaling), resulting in excessive computation. In contrast, conventional VAE-based codecs work in much deeper latent domains (16x - 64x downscaled), motivating a key question: Can diffusion operate effectively in such compact latent spaces without compromising reconstruction quality? To address this, we introduce DiT-IC, an Aligned Diffusion Transformer for Image Compression, which replaces the U-Net with a Diffusion Transformer capable of performing diffusion in latent space entirely at 32x downscaled resolution. DiT-IC adapts a pretrained text-to-image multi-step DiT into a single-step reconstruction model through three key alignment mechanisms: (1) a variance-guided reconstruction flow that adapts denoising strength to latent uncertainty for efficient reconstruction; (2) a self-distillation alignment that enforces consistency with encoder-defined latent geometry to enable one-step diffusion; and (3) a latent-conditioned guidance that replaces text prompts with semantically aligned latent conditions, enabling text-free inference. With these designs, DiT-IC achieves state-of-the-art perceptual quality while offering up to 30x faster decoding and drastically lower memory usage than existing diffusion-based codecs. Remarkably, it can reconstruct 2048x2048 images on a 16 GB laptop GPU.

eess.IV↗

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. Existing methods primarily focus on metrics related to generation quality and controllability. However, they often overlook the evaluation of downstream perception tasks, which are $\mathbf{really\ crucial}$ for the performance of autonomous driving. Existing methods usually leverage a training strategy that first pretrains on synthetic data and finetunes on real data, resulting in twice the epochs compared to the baseline (real data only). When we double the epochs in the baseline, the benefit of synthetic data becomes negligible. To thoroughly demonstrate the benefit of synthetic data, we introduce Dream4Drive, a novel synthetic data generation framework designed for enhancing the downstream perception tasks. Dream4Drive first decomposes the input video into several 3D-aware guidance maps and subsequently renders the 3D assets onto these guidance maps. Finally, the driving world model is fine-tuned to produce the edited, multi-view photorealistic videos, which can be used to train the downstream perception models. Dream4Drive enables unprecedented flexibility in generating multi-view corner cases at scale, significantly boosting corner case perception in autonomous driving. To facilitate future research, we also contribute a large-scale 3D asset dataset named DriveObj3D, covering the typical categories in driving scenarios and enabling diverse 3D-aware video editing. We conduct comprehensive experiments to show that Dream4Drive can effectively boost the performance of downstream perception models under various training epochs. Page: https://wm-research.github.io/Dream4Drive/ GitHub Link: https://github.com/wm-research/Dream4Drive

cs.CV↗

Dual and double canonical bases of quantum groups

Qin established the geometric realization of entire quantum groups via perverse sheaves, which further give rise to dual canonical bases with integral and positive structure constants for quantum groups of type ADE. In this paper, we prove that the dual canonical bases of (Drinfeld double) quantum groups coincide with Berenstein--Greenstein's double canonical bases, by reinterpreting their intricate algebraic construction via the geometry of NKS quiver varieties. This result settles several conjectures therein, including those on positivity and invariance under braid group actions.

math.QA↗

Dual canonical bases of quantum groups and $\imath$quantum groups I: Hall algebras

The $\imath$quantum groups have two realizations: one via the $\imath$Hall algebras and the other via the quantum Grothendieck rings of quiver varieties, as developed by the first author and Wang. The isoclasses of perverse sheaves provide the dual canonical bases for $\imath$quantum groups of type ADE with integral and positive structure constants. In this paper, we present a new construction of the dual canonical bases in the setting of $\imath$Hall algebras. We also introduce Fourier transforms for $\imath$Hall algebras, and prove the invariance of the dual canonical bases under braid group actions and Fourier transforms. Since quantum groups are $\imath$quantum groups of diagonal type, all results also apply to this classical case.

math.QA↗

Dual canonical bases of quantum groups and $\imath$quantum groups II: geometry

The $\imath$quantum groups admit two realizations: one via the $\imath$Hall algebras and the other via the quantum Grothendieck rings of quiver varieties, as developed by the first author and Wang. Based on these two realizations, we establish the dual canonical bases for $\imath$quantum groups of type ADE in two distinct ways, using perverse sheaves and Hall algebras respectively. In this paper, we prove that these two dual canonical bases coincide, thereby proving their invariance under braid group actions, and that their structural constants are integral and positive. Furthermore, we establish the positivity of the coefficients of the transition matrix from the Hall basis (and PBW basis) to the dual canonical basis.

math.QA↗

MC-LLaVA: Multi-Concept Personalized Vision-Language Model

Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies have investigated VLM personalization to understand user-provided concepts. However, they mainly focus on single concepts, neglecting the existence and interplay of multiple concepts, which limits real-world applicability. This paper proposes MC-LLaVA, a multi-concept personalization paradigm. Specifically, MC-LLaVA employs a multi-concept instruction tuning strategy, effectively integrating multiple concepts in a single training step. To reduce the training costs, we propose a personalized textual prompt that uses visual token information to initialize concept tokens. Additionally, we introduce a personalized visual prompt during inference, aggregating location maps for enhanced recognition and grounding capabilities. To further push the performance upper bound, we incorporate an optional auxiliary loss, better enhancing the proposed personalized prompts. To decorate the VLM personalization research, we contribute a high-quality dataset. We carefully collect images with multiple characters and objects from movies and manually create question-answer samples for multi-concept scenarios, featuring superior diversity. Comprehensive experiments demonstrate that MC-LLaVA achieves impressive multi-concept personalized responses, paving the way for VLMs to become better user assistants. The code and dataset will be released at \href{https://github.com/arctanxarc/MC-LLaVA}{https://github.com/arctanxarc/MC-LLaVA}.

cs.CV↗

M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions

This work addresses the challenge of personalized question answering in long-term human-machine interactions: when conversational history spans weeks or months and exceeds the context window, existing personalization mechanisms struggle to continuously absorb and leverage users' incremental concepts, aliases, and preferences. Current personalized multimodal models are predominantly static-concepts are fixed at initialization and cannot evolve during interactions. We propose M2A, an agentic dual-layer hybrid memory system that maintains personalized multimodal information through online updates. The system employs two collaborative agents: ChatAgent manages user interactions and autonomously decides when to query or update memory, while MemoryManager breaks down memory requests from ChatAgent into detailed operations on the dual-layer memory bank, which couples a RawMessageStore (immutable conversation log) with a SemanticMemoryStore (high-level observations), providing memories at different granularities. In addition, we develop a reusable data synthesis pipeline that injects concept-grounded sessions from Yo'LLaVA and MC-LLaVA into LoCoMo long conversations while preserving temporal coherence. Experiments show that M2A significantly outperforms baselines, demonstrating that transforming personalization from one-shot configuration to a co-evolving memory mechanism provides a viable path for high-quality individualized responses in long-term multimodal interactions. The code is available at https://github.com/Little-Fridge/M2A.

cs.AI↗