SearcharxivSearch

arXiv subjects

Xinyuan Hu

Publications and source records attributed to Xinyuan Hu.

15 recordsLinked to original sources

DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views

Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as embodied AI and autonomous driving continue to emerge, reconstructing clean 3D scenes from sparse rainy views in a feed-forward manner becomes increasingly important. Existing feed-forward 3D Gaussian Splatting (3DGS) methods often assume clean inputs and collapse under rainy conditions. To this end, we present \textbf{\textit{DerainSplat}}, a feed-forward framework that reconstructs clean 3D scenes from only a few rainy views. To support this task, we build a large-scale multi-view derain dataset through a four-stage synthesis pipeline that sequentially models overcast illumination, depth-dependent haze, rain streaks, and lens raindrops, producing privileged weather factors. We introduce a weather net that predicts the weather factors from rainy context and yields two support maps. Scene support modulates cross-view cost-volume matching, while radiance support drives depth-aligned appearance fusion to fill corrupted pixels. The derived geometry evidence further attenuates Gaussian opacity to reduce spurious structures. A rainy cycle consistency re-renders clean views using the predicted factors and aligns them with rainy inputs. Extensive experiments show that \textbf{\textit{DerainSplat}} outperforms existing methods on various datasets, including RealEstate10K, ACID, Mip-NeRF360, and real-world rainy scenes, with strong cross-dataset generalization.

cs.CV

Effective potentials for polar molecules under non-orthogonal dual microwave fields

Dual-microwave shielding has emerged as a powerful tool for stabilizing ultracold polar molecules while tuning their intermolecular interactions. However, the two microwave fields are generally not perfectly orthogonal in experiments. Such misalignment introduces an in-plane component of the linearly polarized microwave, whose frequency differs from that of the elliptically polarized field. This component prevents complete cancellation of the dipole-dipole interaction and, more critically, renders the single-molecule dressed state intrinsically time-dependent, so that the conventional time-independent scattering framework is no longer available. Here we develop a Floquet theory that yields an analytic effective potential and enables accurate scattering calculations for polar molecules in non-orthogonal dual microwave fields. We find that, though misalignment weakens the shielding moderately, inelastic losses remain strongly suppressed under experimentally relevant conditions. Meanwhile, misalignment provides additional tunability of the interaction anisotropy and strength, which has been directly applied to recent experimental observations on the gas-to-droplet transition~[Z. Shi \textit{et al}, arXiv:2508.20518 (2025)] and Fermi-surface deformation in microwave-shielded molecular gases~[S. Biswas \textit{et al}, arXiv:2602.22447]. The framework is not restricted to dual-microwave shielding and can be generalized straightforwardly to arbitrary multi-frequency driving, providing a versatile tool for manipulating ultracold polar molecules under complex microwave configurations.

cond-mat.quant-gas

Holographic Airy Beamforming: Curved Trajectory Optimization for Blockage-Resilient Terahertz Communications

Terahertz communication offers vast bandwidth for high-speed transmission in the 6G networks but faces severe blockage challenges in the near-field region due to large antenna arrays. To overcome the limitation that near-field focused beams are susceptible to obstacles, wavefront engineering is leveraged to generate an Airy beam that propagates along a parabolic trajectory to circumvent blockages. In this paper, we consider the reconfigurable holographic surface (RHS) as a potential solution for such precise wavefront engineering owing to its compact radiation element spacing being much smaller than half-wavelength. We reveal that the adjustable effective aperture of the RHS allows the parabolic offset to be located within the antenna aperture, which enhances the freedom in designing Airy beam trajectories. An analog beamforming method, named the holographic Airy beamforming scheme based on amplitude control, is then proposed to generate the curved beam that propagates along the desired trajectory. To maximize the received power of a blocked user, we develop a geometry-based trajectory optimization algorithm. Simulation results validate that, compared to traditional phase-controlled arrays with analog beamforming, the RHS can leverage its adjustable effective aperture to improve the received power of the blocked user by over 10 dB.

eess.SP

Holographic Surface Enabled Integrated Sensing and Communications

Integrated sensing and communications (ISAC) is an essential 6G capability for joint data transmission and environmental sensing. To support 6G scenarios with stringent ISAC performance requirements, existing massive-MIMO-based systems are expected to scale toward ultra-massive MIMO. However, this scaling incurs prohibitive cost and power consumption when realized using widely adopted phased arrays with complex phase shifters and feeding networks. Recently, holographic integrated sensing and communications (HISAC) has emerged as a promising paradigm to address this issue. It employs reconfigurable holographic surfaces (RHSs), a type of leaky-wave antenna, as a cost- and energy-efficient implementation of ultra-massive MIMO-based ISAC, and offers enhanced flexibility for ISAC beam synthesis through holographic beamforming. In this paper, we provide a comprehensive tutorial on HISAC, focusing on how RHS-enabled holographic beamforming can be exploited to jointly support communication and sensing under practical hardware constraints. We first introduce the fundamentals of RHSs and discuss the unique leakage power constraint of holographic beamforming. We then present a general optimization framework for HISAC and show how HISAC enhances joint communication and sensing, sensing-assisted communication, and communication-assisted sensing. We further present HISAC system implementations and experimental results. Finally, we outline promising research directions for HISAC, highlighting the potential of HISAC in advancing efficient, flexible, and high-performance ISAC networks.

eess.SP

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and results. The challenge seeks to identify robust reconstruction pipelines that are robust under real-world adverse conditions, specifically extreme low-light and smoke-degraded environments, as captured by our RealX3D benchmark. A total of 279 participants registered for the competition, of whom 33 teams submitted valid results. We thoroughly evaluate the submitted approaches against state-of-the-art baselines, revealing significant progress in 3D reconstruction under adverse conditions. Our analysis highlights shared design principles among top-performing methods and provides insights into effective strategies for handling 3D scene degradation.

cs.CV

GenSmoke-GS: A Multi-Stage Method for Novel View Synthesis from Smoke-Degraded Images Using a Generative Model

This paper describes our method for Track 2 of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge on smoke-degraded images. In this task, smoke reduces image visibility and weakens the cross-view consistency required by scene optimization and rendering. We address this problem with a multi-stage pipeline consisting of image restoration, dehazing, MLLM-based enhancement, 3DGS-MCMC optimization, and averaging over repeated runs. The main purpose of the pipeline is to improve visibility before rendering while limiting scene-content changes across input views. Experimental results on the challenge benchmark show improved quantitative performance and better visual quality than the provided baselines. The code is available at https://github.com/plbbl/GenSmoke-GS. Our method achieved a ranking of 1 out of 14 participants in Track 2 of the NTIRE 3DRR Challenge, as reported on the official competition website: https://www.codabench.org/competitions/13993/#/results-tab.

cs.CV

Robust quantized transport from topological quasienergy winding in long-range-coupling synthetic quantum walks

Quantized transport is a prominent feature in topological physics, with canonical examples being the quantum Hall effect and adiabatic Thouless pump, which are based on the Chern number, a topological invariant of 2D systems. Going beyond the Chern-number-based paradigms, quantized transports can also arise from k-direction quasienergy winding unique to periodically driven (Floquet) systems, which are free of dimensionality and adiabaticity limitations. However, lattices displaying winding of their quasienergy bands require asymmetric long-range couplings that are difficult to achieve in lattices of real-space coupled sites. Here, by leveraging photonic synthetic dimensions we construct asymmetric long-range-couplings in a one-dimensional temporal quantum walk based on three coupled fiber loops. We demonstrate quantized transport arising from the winding of quasienergy bands in k direction. We show that the average group velocity of an initial wave packet is proportional to the winding number, which leads to a quantized transport displacement. To better visualize this quantized displacement, we cascade two regions with flipped nearest/long-range couplings and observe a focusing effect with a quantized spatial shift in the focusing point. We also probe the robust properties of quantized transport against obstacles and disorders. The study initiates quasienergy-winding-based topological transports, which can feature applications in precise and robust imaging and information processing.

physics.optics

SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images

Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency information in LR inputs. To address this, we propose \textbf{SRSplat}, a feed-forward framework that reconstructs high-resolution 3D scenes from only a few LR views. Our main insight is to compensate for the deficiency of texture information by jointly leveraging external high-quality reference images and internal texture cues. We first construct a scene-specific reference gallery, generated for each scene using Multimodal Large Language Models (MLLMs) and diffusion models. To integrate this external information, we introduce the \textit{Reference-Guided Feature Enhancement (RGFE)} module, which aligns and fuses features from the LR input images and their reference twin image. Subsequently, we train a decoder to predict the Gaussian primitives using the multi-view fused feature obtained from \textit{RGFE}. To further refine predicted Gaussian primitives, we introduce \textit{Texture-Aware Density Control (TADC)}, which adaptively adjusts Gaussian density based on the internal texture richness of the LR inputs. Extensive experiments demonstrate that our SRSplat outperforms existing methods on various datasets, including RealEstate10K, ACID, and DTU, and exhibits strong cross-dataset and cross-resolution generalization capabilities.

cs.CV

Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction

Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible. In this paper, we propose Sparse4DGS, the first method for sparse-frame dynamic scene reconstruction. We observe that dynamic reconstruction methods fail in both canonical and deformed spaces under sparse-frame settings, especially in areas with high texture richness. Sparse4DGS tackles this challenge by focusing on texture-rich areas. For the deformation network, we propose Texture-Aware Deformation Regularization, which introduces a texture-based depth alignment loss to regulate Gaussian deformation. For the canonical Gaussian field, we introduce Texture-Aware Canonical Optimization, which incorporates texture-based noise into the gradient descent process of canonical Gaussians. Extensive experiments show that when taking sparse frames as inputs, our method outperforms existing dynamic or few-shot techniques on NeRF-Synthetic, HyperNeRF, NeRF-DS, and our iPhone-4D datasets.

cs.CV

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often struggle to interpret ambiguous, reasoning-based instructions, while 2D vision-language models that excel at such reasoning lack intrinsic 3D spatial understanding. In this paper, we introduce REALM, an innovative MLLM-agent framework that enables open-world reasoning-based segmentation without requiring extensive 3D-specific post-training. We perform segmentation directly on 3D Gaussian Splatting representations, capitalizing on their ability to render photorealistic novel views that are highly suitable for MLLM comprehension. As directly feeding one or more rendered views to the MLLM can lead to high sensitivity to viewpoint selection, we propose a novel Global-to-Local Spatial Grounding strategy. Specifically, multiple global views are first fed into the MLLM agent in parallel for coarse-level localization, aggregating responses to robustly identify the target object. Then, several close-up novel views of the object are synthesized to perform fine-grained local segmentation, yielding accurate and consistent 3D masks. Extensive experiments show that REALM achieves remarkable performance in interpreting both explicit and implicit instructions across LERF, 3D-OVS, and our newly introduced REALM3D benchmarks. Furthermore, our agent framework seamlessly supports a range of 3D interaction tasks, including object removal, replacement, and style transfer, demonstrating its practical utility and versatility. Project page: https://ChangyueShi.github.io/REALM.

cs.CV

Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents

The rapid expansion of web content has made on-device AI assistants indispensable for helping users manage the increasing complexity of online tasks. The emergent reasoning ability in large language models offer a promising path for next-generation on-device AI agents. However, deploying full-scale Large Language Models (LLMs) on resource-limited local devices is challenging. In this paper, we propose Division-of-Thoughts (DoT), a collaborative reasoning framework leveraging the synergy between locally deployed Smaller-scale Language Models (SLMs) and cloud-based LLMs. DoT leverages a Task Decomposer to elicit the inherent planning abilities in language models to decompose user queries into smaller sub-tasks, which allows hybrid language models to fully exploit their respective strengths. Besides, DoT employs a Task Scheduler to analyze the pair-wise dependency of sub-tasks and create a dependency graph, facilitating parallel reasoning of sub-tasks and the identification of key steps. To allocate the appropriate model based on the difficulty of sub-tasks, DoT leverages a Plug-and-Play Adapter, which is an additional task head attached to the SLM that does not alter the SLM's parameters. To boost adapter's task allocation capability, we propose a self-reinforced training method that relies solely on task execution feedback. Extensive experiments on various benchmarks demonstrate that our DoT significantly reduces LLM costs while maintaining competitive reasoning accuracy. Specifically, DoT reduces the average reasoning time and API costs by 66.12% and 83.57%, while achieving comparable reasoning accuracy with the best baseline methods.

cs.CL

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Language has long been conceived as an essential tool for human reasoning. The breakthrough of Large Language Models (LLMs) has sparked significant research interest in leveraging these models to tackle complex reasoning tasks. Researchers have moved beyond simple autoregressive token generation by introducing the concept of "thought" -- a sequence of tokens representing intermediate steps in the reasoning process. This innovative paradigm enables LLMs' to mimic complex human reasoning processes, such as tree search and reflective thinking. Recently, an emerging trend of learning to reason has applied reinforcement learning (RL) to train LLMs to master reasoning processes. This approach enables the automatic generation of high-quality reasoning trajectories through trial-and-error search algorithms, significantly expanding LLMs' reasoning capacity by providing substantially more training data. Furthermore, recent studies demonstrate that encouraging LLMs to "think" with more tokens during test-time inference can further significantly boost reasoning accuracy. Therefore, the train-time and test-time scaling combined to show a new research frontier -- a path toward Large Reasoning Model. The introduction of OpenAI's o1 series marks a significant milestone in this research direction. In this survey, we present a comprehensive review of recent progress in LLM reasoning. We begin by introducing the foundational background of LLMs and then explore the key technical components driving the development of large reasoning models, with a focus on automated data construction, learning-to-reason techniques, and test-time scaling. We also analyze popular open-source projects at building large reasoning models, and conclude with open challenges and future research directions.

cs.AI

Two- and many-body physics of ultracold molecules dressed by dual microwave fields

We investigate the two- and many-body physics of the ultracold polar molecules dressed by dual microwaves with distinct polarizations. Using Floquet theory and multichannel scattering calculations, we identify a regime with the largest elastic-to-inelastic scattering ratio which is favorable for performing evaporative cooling. Furthermore, we derive and, subsequently, validate an effective interaction potential that accurately captures the dynamics of microwave-shielded polar molecules (MSPMs). We also explore the ground-state properties of the ultracold gases of MSPMs by computing physical quantities such as gas density, condensate fraction, momentum distribution, and second-order correlation. It is shown that the system supports a weakly correlated expanding gas state and a strongly correlated self-bound gas state. Since the dual-microwave scheme introduces addition control knob and is essential for creating ultracold Bose gases of polar molecules, our work pave the way for studying two- and many-body physics of the ultracold polar molecules dressed by dual microwaves.

cond-mat.quant-gas

LIMP: Large Language Model Enhanced Intent-aware Mobility Prediction

Human mobility prediction is essential for applications like urban planning and transportation management, yet it remains challenging due to the complex, often implicit, intentions behind human behavior. Existing models predominantly focus on spatiotemporal patterns, paying less attention to the underlying intentions that govern movements. Recent advancements in large language models (LLMs) offer a promising alternative research angle for integrating commonsense reasoning into mobility prediction. However, it is a non-trivial problem because LLMs are not natively built for mobility intention inference, and they also face scalability issues and integration difficulties with spatiotemporal models. To address these challenges, we propose a novel LIMP (LLMs for Intent-ware Mobility Prediction) framework. Specifically, LIMP introduces an "Analyze-Abstract-Infer" (A2I) agentic workflow to unleash LLM's commonsense reasoning power for mobility intention inference. Besides, we design an efficient fine-tuning scheme to transfer reasoning power from commercial LLM to smaller-scale, open-source language model, ensuring LIMP's scalability to millions of mobility records. Moreover, we propose a transformer-based intention-aware mobility prediction model to effectively harness the intention inference ability of LLM. Evaluated on two real-world datasets, LIMP significantly outperforms baseline models, demonstrating improved accuracy in next-location prediction and effective intention inference. The interpretability of intention-aware mobility prediction highlights our LIMP framework's potential for real-world applications. Codes and data can be found in https://github.com/tsinghua-fib-lab/LIMP .

cs.CL

RIS-based IMT-2030 Testbed for MmWave Multi-stream Ultra-massive MIMO Communications

As one enabling technique of the future sixth generation (6G) network, ultra-massive multiple-input-multiple-output (MIMO) can support high-speed data transmissions and cell coverage extension. However, it is hard to realize the ultra-massive MIMO via traditional phased arrays due to unacceptable power consumption. To address this issue, reconfigurable intelligent surface-based (RIS-based) antennas are an energy-efficient enabler of the ultra-massive MIMO, since they are free of energy-hungry phase shifters. In this article, we report the performances of the RIS-enabled ultra-massive MIMO via a project called Verification of MmWave Multi-stream Transmissions Enabled by RIS-based Ultra-massive MIMO for 6G (V4M), which was proposed to promote the evolution towards IMT-2030. In the V4M project, we manufacture RIS-based antennas with 1024 one-bit elements working at 26 GHz, based on which an mmWave dual-stream ultra-massive MIMO prototype is implemented for the first time. To approach practical settings, the Tx and Rx of the prototype are implemented by one commercial new radio base station and one off-the-shelf user equipment, respectively. The measured data rate of the dual-stream prototype approaches the theoretical peak rate. Our contributions to the V4M project are also discussed by presenting technological challenges and corresponding solutions.

cs.IT