SearcharxivSearch

arXiv subjects

Xin Zeng

Publications and source records attributed to Xin Zeng.

At least 19 recordsLinked to original sources

GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirrors or transmitted content behind glass rather than the physical surface itself. Adapting these models with real-world data is challenging because conventional depth sensors are also unreliable in such regions. We observe that while the appearance of a non-Lambertian surface varies with its reflected or transmitted environment, its underlying geometry remains unchanged. Based on this observation, we propose GIFT (Geometry-Invariant Fine-Tuning), a parameter-efficient post-training framework that requires no measured depth labels. We collect groups of RGB images under controlled appearance changes while keeping the camera and target geometry fixed. GIFT exploits geometric invariance across these observations to suppress non-Lambertian depth hallucinations while retaining general depth estimation capability. We further construct a controlled benchmark that evaluates non-Lambertian depth recovery, robustness to appearance changes, and performance retention in other regions. Experiments on our benchmark and an independent real-world dataset demonstrate that GIFT improves depth prediction for mirrors and transparent objects while largely preserving the base model's performance, providing a practical and low-cost approach for adapting monocular depth foundation models to non-Lambertian scenes.

cs.CV

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared tier, with a margin approaching 20 points at full scale. BM25 also anchors the low-cost end of the Pareto frontier without LLM-based construction. Dense retrieval remains efficient but less accurate, whereas graph-based RAG encounters construction walls before deployment scale and its scalable variants remain below BM25 at shared tiers. Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable default, while agentic reasoning works best after ranked discovery rather than in place of it.

cs.CL

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC), explicitly connecting high-level task understanding, behavior planning, and low-level control. The datasets are collected from both real-world and simulated outdoor environments for training and evaluation. We further develop ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, and deploy it on a physical UAV platform. Experiments with representative VLMs and VLA models show that current UAV agents still struggle with behavior planning, viewpoint adjustment, and robust task completion in active perception. These results establish ActiveFly-Bench as a new testbed for embodied aerial intelligence.

cs.RO

Reconfigurable ultrafast perovskite polariton logic gates via nonlinear dynamics

Exciton-polaritons provide a great platform for developing ultrafast all-optical logic gates for quantum and optical chips. However, progress toward practical polariton logic remains limited due to incomplete logical functionality on a single device. Herein, we present a single-device perovskite polariton platform enabling reconfigurable, ultrafast logic gates with functional completeness. The device consists of an optically trapped perovskite microwire, generating well-controlled non-equilibrium polariton condensation states for multiple logic operation channels. By tailoring the power of signal and gate beams, the same device is programmed to execute three basic Boolean functions (AND,OR,and NOT) and a high-order XOR function with a high on/off ratio of 21 dB, and a fast response time 6.7 ps. The reconfigurability arises from the selective activation of different nonlinear responses of polariton condensates, including amplification, seeding state transitions, and nonlinear interaction. These results provide valuable insights for advancing exciton-polariton logic gates.

physics.optics

Phase conjugated master oscillator fiber power amplifier

High-power narrow-linewidth fiber lasers are fundamentally limited by stimulated Brillouin scattering (SBS), which constrains further power scaling while maintaining spectral linewidth. Traditional mitigation techniques, such as active phase modulation, often introduce trade-offs among complexity, cost, and spectral brightness. In this study, we propose and experimentally demonstrate a novel all-optical approach for spectral linewidth manipulation and SBS suppression using optical phase conjugation (OPC). By leveraging nonlinear spectral broadening followed by phase conjugation, this method enables sophisticated linewidth narrowing in fiber amplifier, resulting in narrow linewidth output and enhanced SBS threshold. Using a low-cost fiber oscillator as the seed source, we achieve a spectral compression ratio exceeding 3 times. This method not only eliminates the need for complex electro-optic modulation systems but also provides a pathway toward simpler, high-brightness fiber laser systems. Our findings underscore the viability of OPC as a transformative tool for nonlinearity management and power scaling in high-performance fiber laser architectures.

physics.optics

Deep Submillimeter and Radio Observations in the SSA22 Field. IV. Spectral Energy Distributions, Star Formation Histories, and the Infrared-Radio Correlation of the 850 $\mu$m-selected SMGs

We analyze the spectral energy distributions (SEDs), star formation histories (SFHs), and infrared-radio correlation (IRRC) of 221 850 $\mu$m-selected submillimeter galaxies (SMGs) in the SSA22 deep field. The median mass-weighted age is 567 Myr. Most galaxies in our sample began forming $\sim$ 1.68 Gyr after the Big Bang, entered the `SMG phase' after $\sim$ 1 Gyr of evolution -- when they are predominantly observed -- and largely transitioned out of the `SMG phase' to become quiescent within an additional $\sim$ 0.2 Gyr. A subset of massive galaxies shows rapid early assembly with high star formation efficiencies ($\sim$0.2-0.8). The majority of SMGs reside at the high-mass end of the star-forming main sequence, with a characteristic stellar mass of $M_{star} \sim 10^{11}$ M$_\odot$, above which galaxies are predominantly either on the main sequence or already quenched. We observe a downsizing trend: more massive galaxies tend to ``mature" earlier, completing their major episodes of star formation at higher redshifts compared to lower-mass systems. Our sample contributes $\sim$ 21% (28%) to the cosmic star formation rate density (stellar mass density), including the overdensity, with its relative contribution peaking at 50-60% in the redshift range $z=2.5-3.5$. The median infrared-radio correlation parameter $q_{IR}$ is 2.37, evolving as $(1+z)^{-0.11}$, likely due to AGN contributions at high redshift and intrinsic differences between low- and high-redshift populations.

astro-ph.GA

Deep Submillimeter and Radio Observations in the SSA22 Field. III. Multiwavelength Identifications and Properties of the 850 $\mu$m-selected Submillimeter Galaxies

We present a multiwavelength analysis of 850 $\mu$m-selected SMGs (deblended S$_{\rm 850}\gtrsim$ 1mJy) in the SSA22 field, where our deepest JCMT/SCUBA-2 observations reach a sensitivity of $\sigma_{850}\sim$ 0.80mJy beam$^{-1}$. Using multiple identification methods, we have identified 248 deblended SMG candidates for 192 SCUBA-2 sources. The average multiplicity of SCUBA-2 sources is $\sim$26%, with brighter sources exhibiting higher multiplicity. After applying quality cuts based on SED fitting reliability, our final sample comprises 221 SMGs associated with 186 SCUBA-2 sources. The SSA22 SMGs have a median infrared luminosity of (2.25$\pm$0.25) $\times$10$^{12}$ L$_{\odot}$, with $\sim$ 63% ($\sim$ 8%) of the sample classified as ULIRGs (HLIRGs). The median redshift of the sample is $z = 2.00 \pm 0.08$, while optically faint galaxies exhibit higher median redshift ($\sim 2.20 \pm 0.17$). The comoving volume density of SMGs increases by a factor of $\sim 6$ at $z \lesssim 4$, plateauing at $\sim$ 1.78-3.16 $\times$ 10$^{-5}$ cMpc$^{-3}$ over $z \sim$ 1-3 (including the overdensity). The significant overdensity of SMGs within large-scale structures demonstrates their reliability as tracers of cosmic structure formation at high redshift. The median stellar mass and SFR of our SMG sample are $(1.55 \pm 0.22) \times 10^{11}$ M$_\odot$ and $166 \pm 25$ M$_\odot$ yr$^{-1}$, respectively. We observe a clear ``downsizing" signature: after cosmic noon ($z \lesssim 2$), massive SMGs exhaust their gas reservoirs and transition to quiescence, while lower-mass SMGs continue forming stars and dominate the cosmic SFR density. The sample has a median dust mass of (1.95 $\pm$ 0.14) $\times$ 10$^{9}$ M$_{\odot}$. The dust fraction ($ M_{\text{dust}}/M_{\text{star}}$) has a median value of (1.4 $\pm$ 0.2) $\times$ 10$^{-2}$. The median $A_V$ of SMGs is 3.09$\pm$0.07mag.

astro-ph.GA

Ultrafast Exciton-Polariton Transport and Relaxation in Halide Perovskite

Halide perovskites offer a great platform for room-temperature exciton-polaritons (EPs) due to their strong oscillator strength and large exciton binding energy, promising applications in next-generation photonic and polaritonic devices. Efficient manipulation of EP transport and relaxation is critical for device performance, yet their spatiotemporal dynamics across different in-plane momenta (k//) remain poorly understood due to limitations in experimental access. In this work, we employ energy-resolved transient reflectance microscopy (TRM) combined with the dispersion relation of EPs to achieve high-resolution imaging of EP transport at specific k//. This approach directly reveals the quasi-ballistic transport and ultrafast relaxation of EPs in different k// regions, showcasing diffusion as fast as ~490 cm2/s and a relaxation time of ~95.1 fs. Furthermore, by tuning the detuning parameter, we manipulate the ballistic transport group velocity and relaxation time of EPs across varying k//. Our results reveal key insights into the dynamics of EP transport and relaxation, providing valuable guidance for the design and optimization of polaritonic devices.

cond-mat.mtrl-sci

Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space

Unmanned aerial vehicles (UAVs) have emerged as powerful embodied agents. One of the core abilities is autonomous navigation in large-scale three-dimensional environments. Existing navigation policies, however, are typically optimized for low-level objectives such as obstacle avoidance and trajectory smoothness, lacking the ability to incorporate high-level semantics into planning. To bridge this gap, we propose ANWM, an aerial navigation world model that predicts future visual observations conditioned on past frames and actions, thereby enabling agents to rank candidate trajectories by their semantic plausibility and navigational utility. ANWM is trained on 4-DoF UAV trajectories and introduces a physics-inspired module: Future Frame Projection (FFP), which projects past frames into future viewpoints to provide coarse geometric priors. This module mitigates representational uncertainty in long-distance visual generation and captures the mapping between 3D trajectories and egocentric observations. Empirical results demonstrate that ANWM significantly outperforms existing world models in long-distance visual forecasting and improves UAV navigation success rates in large-scale environments.

cs.RO

OmniTFT: Omni Target Forecasting for Vital Signs and Laboratory Result Trajectories in Multi Center ICU Data

Accurate multivariate time-series prediction of vital signs and laboratory results is crucial for early intervention and precision medicine in intensive care units (ICUs). However, vital signs are often noisy and exhibit rapid fluctuations, while laboratory tests suffer from missing values, measurement lags, and device-specific bias, making integrative forecasting highly challenging. To address these issues, we propose OmniTFT, a deep learning framework that jointly learns and forecasts high-frequency vital signs and sparsely sampled laboratory results based on the Temporal Fusion Transformer (TFT). Specifically, OmniTFT implements four novel strategies to enhance performance: sliding window equalized sampling to balance physiological states, frequency-aware embedding shrinkage to stabilize rare-class representations, hierarchical variable selection to guide model attention toward informative feature clusters, and influence-aligned attention calibration to enhance robustness during abrupt physiological changes. By reducing the reliance on target-specific architectures and extensive feature engineering, OmniTFT enables unified modeling of multiple heterogeneous clinical targets while preserving cross-institutional generalizability. Across forecasting tasks, OmniTFT achieves substantial performance improvement for both vital signs and laboratory results on the MIMIC-III, MIMIC-IV, and eICU datasets. Its attention patterns are interpretable and consistent with known pathophysiology, underscoring its potential utility for quantitative decision support in clinical care.

cs.LG

Magnonic entanglement in a chiral cavity-magnon coupling system

The generation of magnon entanglement and squeezing plays a crucial role in quantum information processing. In this study, we propose a scheme based on a chiral cavity-magnon system, which consists of a torus-shaped cavity and two yttrium iron garnet spheres. The magnon mode of each yttrium iron garnet sphere is selectively coupled to one of the two degenerate rotating microwave modes of the toroidal cavity. The system aims to achieve entangled and squeezed magnon states through the mediation of the cavity. We further show that bipartite entanglement can be achieved by tuning external driving parameters. Additionally, our scheme does not rely on the magnon Kerr nonlinearity, which is usually extremely weak in yttrium iron garnet spheres. This work provides insights and methods for the research of quantum states in cavity-magnon systems.

quant-ph

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of point clouds over other modalities remain unclear. Moreover, existing 3D benchmarks are insufficient for fairly evaluating the ability of multimodal LLMs to comprehend spatial concepts. To address these challenges, we introduce ScanReQA, a 3D spatial reasoning benchmark encompassing text, vision, and point cloud modalities. We then evaluate the performance of text, 2D, and 3D LLMs on the benchmark to compare the effectiveness of different modalities in understanding spatial concepts. Furthermore, we analyze the reasoning mechanisms behind 3D LLMs using point clouds. Our findings reveal that: 1) binary spatial reasoning remains challenging for current 3D LLMs, 2) MLLMs based on point cloud and visual modalities demonstrate stronger spatial reasoning capabilities than LLMs, and 3) 3D LLMs exhibit the attention sink phenomenon similar to that in 2D LLMs, impairing spatial reasoning. We think these conclusions can help the next step of 3D LLMs and also offer insights for foundation models in other modalities. We release datasets and codes in the project page: https://github.com/EmbodiedCity/ScanReQA.code.

cs.CV

Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space

Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs spanning 7 general spatial reasoning tasks, including multiple-choice, true/false, and short-answer formats, and supports both visual and point cloud modalities. The questions are automatically generated from spatial relations extracted from both real-world and simulated aerial scenes. Evaluation on 13 popular MLLMs reveals that: 1) Models are generally better at answering questions about relative spatial relations than absolute distances, 2) 3D LLMs fail to demonstrate significant advantages over 2D LLMs, and 3) Fine-tuning solely on the simulated dataset can significantly improve the model's spatial reasoning performance in real-world scenarios. We release our benchmark, data generation pipeline, and evaluation toolkit to support further research: https://github.com/EmbodiedCity/Open3D-VQA.code.

cs.CV

All-optical and ultrafast control of high-order exciton-polariton orbital modes

Exciton-polaritons flows within closed quantum circuits can spontaneously form phase-locked modes that carry orbital angular momentum (OAM). With its infinite set of angular momentum quantum numbers, high-order OAM represents a transformative solution to the bandwidth bottleneck in multiplexed optical communication. However, its practical application is hindered by the limited choice of materials which in general requires cryogenic temperatures and the reliance on mechanical switching. In this work, we achieve stable and high-order (up to order of 33) OAM modes by constructing a closed quantum circuit using the halide perovskite microcavities at room temperature. By controlling the spatial and temporal symmetry of the closed quantum circuits using another laser pulse, we achieve significant tuning OAM of EP flows from 8 to 12. Our work demonstrate all-optical and ultrafast control of high-order OAM using exciton-polariton condensates in perovskite microcavities that would have important applications in high-throughput optical communications.

physics.optics

Generation of high-fidelity Greenberger-Horne-Zeilinger states in a driven hybrid quantum system

In this study, we propose a theoretical scheme for achieving long-distance Greenberger-Horne-Zeilinger states in a driven hybrid quantum system. By applying a microwave field to the YIG sphere, we utilize the Kerr effect to induce the squeezing of the magnon, thereby achieving an exponential enhancement of the coupling strength between the magnonic mode and spins, and we also discuss in detail the relationship between the squeezing parameter and the external microwave field. By means of the Schrieffer-Wolff transformation, the magnonic mode can be adiabatically eliminated under the large detuning condition, thereby establishing a robust effective interaction between spins essential for realizing the desired entangled state. Numerical simulations indicate that the squeezing parameter can be effectively increased by adjusting the driving field, and our proposal can generate high-fidelity Greenberger-Horne-Zeilinger states even in dissipative systems. Additionally, we extensively discuss the influence of inhomogeneous broadening on the entangled states, and the experimental feasibility shows that our results provide possibilities in the realms of quantum networking and quantum computing.

quant-ph

AUGlasses: Continuous Action Unit based Facial Reconstruction with Low-power IMUs on Smart Glasses

Recent advancements in augmented reality (AR) have enabled the use of various sensors on smart glasses for applications like facial reconstruction, which is vital to improve AR experiences for virtual social activities. However, the size and power constraints of smart glasses demand a miniature and low-power sensing solution. AUGlasses achieves unobtrusive low-power facial reconstruction by placing inertial measurement units (IMU) against the temporal area on the face to capture the skin deformations, which are caused by facial muscle movements. These IMU signals, along with historical data on facial action units (AUs), are processed by a transformer-based deep learning model to estimate AU intensities in real-time, which are then used for facial reconstruction. Our results show that AUGlasses accurately predicts the strength (0-5 scale) of 14 key AUs with a cross-user mean absolute error (MAE) of 0.187 (STD = 0.025) and achieves facial reconstruction with a cross-user MAE of 1.93 mm (STD = 0.353). We also integrated various preprocessing and training techniques to ensure robust performance for continuous sensing. Micro-benchmark tests indicate that our system consistently performs accurate continuous facial reconstruction with a fine-tuned cross-user model, achieving an AU MAE of 0.35.

cs.HC

xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data

Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M records for humans with each record measuring 20,000 genes. This highlights the need for unsupervised representation learning to fully ingest these data, yet classical transformer architectures are prohibitive to train on such data in terms of both computation and memory. To address this challenge, we propose a novel asymmetric encoder-decoder transformer for scRNA-seq data, called xTrimoGene$^α$ (or xTrimoGene for short), which leverages the sparse characteristic of the data to scale up the pre-training. This scalable design of xTrimoGene reduces FLOPs by one to two orders of magnitude compared to classical transformers while maintaining high accuracy, enabling us to train the largest transformer models over the largest scRNA-seq dataset today. Our experiments also show that the performance of xTrimoGene improves as we scale up the model sizes, and it also leads to SOTA performance over various downstream tasks, such as cell type annotation, perturb-seq effect prediction, and drug combination prediction. xTrimoGene model is now available for use as a service via the following link: https://api.biomap.com/xTrimoGene/apply.

cs.LG

xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein

Protein language models have shown remarkable success in learning biological information from protein sequences. However, most existing models are limited by either autoencoding or autoregressive pre-training objectives, which makes them struggle to handle protein understanding and generation tasks concurrently. We propose a unified protein language model, xTrimoPGLM, to address these two types of tasks simultaneously through an innovative pre-training framework. Our key technical contribution is an exploration of the compatibility and the potential for joint optimization of the two types of objectives, which has led to a strategy for training xTrimoPGLM at an unprecedented scale of 100 billion parameters and 1 trillion training tokens. Our extensive experiments reveal that 1) xTrimoPGLM significantly outperforms other advanced baselines in 18 protein understanding benchmarks across four categories. The model also facilitates an atomic-resolution view of protein structures, leading to an advanced 3D structural prediction model that surpasses existing language model-based tools. 2) xTrimoPGLM not only can generate de novo protein sequences following the principles of natural ones, but also can perform programmable generation after supervised fine-tuning (SFT) on curated sequences. These results highlight the substantial capability and versatility of xTrimoPGLM in understanding and generating protein sequences, contributing to the evolving landscape of foundation models in protein science.

q-bio.QM