SearcharxivSearch

arXiv subjects

Henry Zheng

Publications and source records attributed to Henry Zheng.

18 recordsLinked to original sources

ZAYA1-8B Technical Report

We present ZAYA1-8B, a reasoning-focused mixture-of-experts (MoE) model with 700M active and 8B total parameters, built on Zyphra's MoE++ architecture. ZAYA1-8B's core pretraining, midtraining, and supervised fine-tuning (SFT) were performed on a full-stack AMD compute, networking, and software platform. With under 1B active parameters, ZAYA1-8B matches or exceeds DeepSeek-R1-0528 on several challenging mathematics and coding benchmarks, and remains competitive with substantially larger open-weight reasoning models. ZAYA1-8B was trained from scratch for reasoning, with reasoning data included from pretraining onward using an answer-preserving trimming scheme. Post-training uses a four-stage RL cascade: reasoning warmup on math and puzzles; a 400-task RLVE-Gym curriculum; math and code RL with test-time compute traces and synthetic code environments built from competitive-programming references; and behavioral RL for chat and instruction following. We also introduce Markovian RSA, a test-time compute method that recursively aggregates parallel reasoning traces while carrying forward only bounded-length reasoning tails between rounds. In TTC evaluation, Markovian RSA raises ZAYA1-8B to 91.9\% on AIME'25 and 89.6\% on HMMT'25 while carrying forward only a 4K-token tail, narrowing the gap to much larger reasoning models including Gemini-2.5 Pro, DeepSeek-V3.2, and GPT-5-High.

cs.AI

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding

Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains underexplored. Effective 3D reasoning hinges on accurate grounding: to answer open-ended queries, a model must first identify query-relevant objects and regions in a complex scene, and then reason about their spatial and geometric relationships. Recent approaches have demonstrated strong potential for grounded 3D reasoning. However, they often rely on in-domain tuning or hand-crafted reasoning pipelines, which limit their flexibility and zero-shot generalization to novel environments. In this work, we present MAG-3D, a training-free multi-agent framework for grounded 3D reasoning with off-the-shelf VLMs. Instead of relying on task-specific training or fixed reasoning procedures, MAG-3D dynamically coordinates expert agents to address the key challenges of 3D reasoning. Specifically, we propose a planning agent that decomposes the task and orchestrates the overall reasoning process, a grounding agent that performs free-form 3D grounding and relevant frame retrieval from extensive 3D scene observations, and a coding agent that conducts flexible geometric reasoning and explicit verification through executable programs. This multi-agent collaborative design enables flexible training-free 3D grounded reasoning across diverse scenes and achieves state-of-the-art performance on challenging benchmarks.

cs.CV

CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration

High-quality 3D scene generation from a single image is crucial for AR/VR and embodied AI applications. Early approaches struggle to generalize due to reliance on specialized models trained on curated small datasets. While recent advancements in large-scale 3D foundation models have significantly enhanced instance-level generation, coherent scene generation remains a challenge, where performance is limited by inaccurate per-object pose estimations and spatial inconsistency. To this end, this paper introduces CC-FMO, a zero-shot, camera-conditioned pipeline for single-image to 3D scene generation that jointly conforms to the object layout in input image and preserves instance fidelity. CC-FMO employs a hybrid instance generator that combines semantics-aware vector-set representation with detail-rich structured latent representation, yielding object geometries that are both semantically plausible and high-quality. Furthermore, CC-FMO enables the application of foundational pose estimation models in the scene generation task via a simple yet effective camera-conditioned scale-solving algorithm, to enforce scene-level coherence. Extensive experiments demonstrate that CC-FMO consistently generates high-fidelity camera-aligned compositional scenes, outperforming all state-of-the-art methods.

cs.CV

Efficient evaluation of the dark-matter two-loop power spectrum in the EFT of LSS

Rapid progress in cosmological Large Scale Structure (LSS) surveys motivates precise theoretical predictions. The Effective Field Theory of Large-Scale Structure (EFTofLSS) is routinely applied to data, and requires fast computation of its predictions when sampling the large space of cosmological parameters. Going beyond existing one-loop techniques, we present a method to rapidly evaluate the two-loop power spectrum. Our method decomposes the typically small difference between a given linear power spectrum and a reference power spectrum into a cosmology-independent basis of functions resembling massive scalar propagators in Quantum Field Theory. By taking the leading terms in such a small difference, we numerically evaluate the cosmology-independent loop integrals where in the integrand only the relevant combinations of basis functions appear. We achieve an efficient numerical evaluation via physically motivated local ultraviolet subtractions and by arranging the cancellation of infrared singularities locally in the integrands. Final predictions are obtained by contracting these precomputed integrals with the cosmology-dependent coordinates of the expansion in the fixed basis. We present and publicly release the precomputed integrals for the renormalized two-loop dark-matter power spectrum in the EFTofLSS. These require eight EFT counterterms, which include the effect of generated vorticity, and are sufficient to analyze the lensing galaxy signal in LSS surveys at this order.

astro-ph.CO

PyBird-JAX: Accelerated inference in large-scale structure with model-independent emulation of one-loop galaxy power spectra

We present $\texttt{PyBird-JAX}$, a differentiable, $\texttt{JAX}$-based implementation of $\texttt{PyBird}$, using internal neural network emulators to accelerate computationally costly operations for rapid large-scale structure (LSS) analysis. $\texttt{PyBird-JAX}$ computes one-loop EFTofLSS predictions for redshift-space galaxy power spectrum multipoles in 1.2 ms on a CPU and 0.2 ms on a GPU, achieving 3-4 orders of magnitude speed-up over $\texttt{PyBird}$. The emulators take a compact spline-based representation of the input linear power spectrum $P(k)$ as feature vectors, making the approach applicable to a wide range of cosmological models. We rigorously validate its accuracy against large-volume simulations and on BOSS data, including cosmologies not explicitly represented in the training set. Leveraging automatic differentiation, $\texttt{PyBird-JAX}$ supports Fisher forecasting, Taylor expansion of model predictions, gradient-based searches, and vectorised ensemble sampling. Interfaced with a variety of samplers and Boltzmann solvers, $\texttt{PyBird-JAX}$ provides a high-performance, end-to-end inference pipeline. Combined with a symbolic-$P(k)$ generator, a typical Stage-4 LSS MCMC converges in minutes on a GPU. Our results demonstrate that $\texttt{PyBird-JAX}$ delivers the precision and speed required for upcoming LSS surveys, opening the door to accelerated cosmological inference with minimal accuracy loss and no pretraining. In a companion paper [1], we put $\texttt{PyBird-JAX}$ to use in achieving LSS marginalised constraints free from volume projection effects through non-flat measures.

astro-ph.CO

Debiasing inference in large-scale structure with non-flat volume measures

Increasingly large parameter spaces, used to more accurately model precision observables in physics, can paradoxically lead to large deviations in the inferred parameters of interest -- a bias known as volume projection effects -- when marginalising over many nuisance parameters. For posterior distributions that admit a Laplace expansion, we show that this artefact of Bayesian inference can be mitigated by defining expectation values with respect to a non-flat volume measure, such that the posterior mean becomes unbiased on average. We begin by finding a measure that ensures the mean is an unbiased estimator of the mode. Although the mode itself, as we rediscover, is biased under sample averaging, this choice yields the least biased estimator due to a cancellation we clarify. We further explain why bias in marginal posteriors can appear relatively large, yet remains correctable, when the number of nuisances is large. To demonstrate our approach, we present mock analyses in large-scale structure (LSS) wherein cosmological parameters are subject to large projection effects (at the 1-2$\sigma$ level) under a flat measure, that are however recovered at high fidelity ($<0.1\sigma$) when estimated using non-flat counterparts. Our cosmological analyses are enabled by $\texttt{PyBird-JAX}$, a fast, differentiable pipeline for LSS developed in our companion paper [1].

astro-ph.CO

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based on verbal descriptions. However, this task faces two significant challenges: (1) loss of fine-grained visual semantics due to sparse fusion of point clouds with ego-centric multi-view images, (2) limited textual semantic context due to arbitrary language descriptions. We propose DenseGrounding, a novel approach designed to address these issues by enhancing both visual and textual semantics. For visual features, we introduce the Hierarchical Scene Semantic Enhancer, which retains dense semantics by capturing fine-grained global scene features and facilitating cross-modal alignment. For text descriptions, we propose a Language Semantic Enhancer that leverages large language models to provide rich context and diverse language descriptions with additional context during model training. Extensive experiments show that DenseGrounding significantly outperforms existing methods in overall accuracy, with improvements of 5.81% and 7.56% when trained on the comprehensive full dataset and smaller mini subset, respectively, further advancing the SOTA in egocentric 3D visual grounding. Our method also achieves 1st place and receives the Innovation Award in the CVPR 2024 Autonomous Grand Challenge Multi-view 3D Visual Grounding Track, validating its effectiveness and robustness.

cs.CV

Optimizers for Stabilizing Likelihood-free Inference

A growing number of applications in particle physics and beyond use neural networks as unbinned likelihood ratio estimators applied to real or simulated data. Precision requirements on the inference tasks demand a high-level of stability from these networks, which are affected by the stochastic nature of training. We show how physics concepts can be used to stabilize network training through a physics-inspired optimizer. In particular, the Energy Conserving Descent (ECD) optimization framework uses classical Hamiltonian dynamics on the space of network parameters to reduce the dependence on the initial conditions while also stabilizing the result near the minimum of the loss function. We develop a version of this optimizer known as $ECD_{q=1}$, which has few free hyperparameters with limited ranges guided by physical reasoning. We apply $ECD_{q=1}$ to representative likelihood-ratio estimation tasks in particle physics and find that it out-performs the widely-used Adam optimizer. We expect that ECD will be a useful tool for wide array of data-limited problems, where it is computationally expensive to exhaustively optimize hyperparameters and mitigate fluctuations with ensembling.

hep-ph

ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding

Embodied intelligence requires agents to interact with 3D environments in real time based on language instructions. A foundational task in this domain is ego-centric 3D visual grounding. However, the point clouds rendered from RGB-D images retain a large amount of redundant background data and inherent noise, both of which can interfere with the manifold structure of the target regions. Existing point cloud enhancement methods often require a tedious process to improve the manifold, which is not suitable for real-time tasks. We propose Proxy Transformation suitable for multimodal task to efficiently improve the point cloud manifold. Our method first leverages Deformable Point Clustering to identify the point cloud sub-manifolds in target regions. Then, we propose a Proxy Attention module that utilizes multimodal proxies to guide point cloud transformation. Built upon Proxy Attention, we design a submanifold transformation generation module where textual information globally guides translation vectors for different submanifolds, optimizing relative spatial relationships of target regions. Simultaneously, image information guides linear transformations within each submanifold, refining the local point cloud manifold of target regions. Extensive experiments demonstrate that Proxy Transformation significantly outperforms all existing methods, achieving an impressive improvement of 7.49% on easy targets and 4.60% on hard targets, while reducing the computational overhead of attention blocks by 40.6%. These results establish a new SOTA in ego-centric 3D visual grounding, showcasing the effectiveness and robustness of our approach.

cs.CV

Neutrino masses from large-scale structures: future sensitivity and theory dependence

In the incoming years, cosmological surveys aim at measuring the sum of neutrino masses $Σm_ν$, complementing the determination of their mass ordering from laboratory experiments. In order to assess the full potential of large-scale structures (LSS), we employ state-of-the-art predictions from the effective field theory of LSS (EFTofLSS) at one loop to perform Fisher forecasts on the sensitivity (combining power spectrum and bispectrum) of ongoing and future surveys (DESI, MegaMapper) in combination with CMB measurements (Planck, Litebird and Stage-4). We find that the 1$σ$ sensitivity on $Σm_ν$ is expected to be 15 meV with Planck+DESI, and 7 meV with S4+MegaMapper, where $\sim 10\%$ and $30\%$ of the constraints are brought by the one-loop bispectrum respectively. To understand how robust are these bounds, we explore how they are relaxed when considering extensions to the standard model, dubbed `new physics'. We find that the shift induced on $Σm_ν$ by a $1σ$ shift on new physics parameters (we consider extra relativistic species, neutrino self-interactions, curvature or a time-evolving electron mass) could be $\mathcal O(10)$ meV for Planck+DESI, but it will be suppressed down to $\mathcal O(1)$ meV in S4+MegaMapper. Our study highlights the quantitative impact of including the bispectrum at one loop in the EFTofLSS, and the robustness of the sensitivity to $Σm_ν$ against potential new physics thanks to the synergy of cosmological probes.

astro-ph.CO

Training an Open-Vocabulary Monocular 3D Object Detection Model without 3D Data

Open-vocabulary 3D object detection has recently attracted considerable attention due to its broad applications in autonomous driving and robotics, which aims to effectively recognize novel classes in previously unseen domains. However, existing point cloud-based open-vocabulary 3D detection models are limited by their high deployment costs. In this work, we propose a novel open-vocabulary monocular 3D object detection framework, dubbed OVM3D-Det, which trains detectors using only RGB images, making it both cost-effective and scalable to publicly available data. Unlike traditional methods, OVM3D-Det does not require high-precision LiDAR or 3D sensor data for either input or generating 3D bounding boxes. Instead, it employs open-vocabulary 2D models and pseudo-LiDAR to automatically label 3D objects in RGB images, fostering the learning of open-vocabulary monocular 3D detectors. However, training 3D models with labels directly derived from pseudo-LiDAR is inadequate due to imprecise boxes estimated from noisy point clouds and severely occluded objects. To address these issues, we introduce two innovative designs: adaptive pseudo-LiDAR erosion and bounding box refinement with prior knowledge from large language models. These techniques effectively calibrate the 3D labels and enable RGB-only training for 3D detectors. Extensive experiments demonstrate the superiority of OVM3D-Det over baselines in both indoor and outdoor scenarios. The code will be released.

cs.CV

Reducing Overtreatment of Indeterminate Thyroid Nodules Using a Multimodal Deep Learning Model

Objective: Molecular testing (MT) classifies cytologically indeterminate thyroid nodules as benign or malignant with high sensitivity but low positive predictive value (PPV), only using molecular profiles, ignoring ultrasound (US) imaging and biopsy. We address this limitation by applying attention multiple instance learning (AMIL) to US images. Methods: We retrospectively reviewed 333 patients with indeterminate thyroid nodules at UCLA medical center (259 benign, 74 malignant). A multi-modal deep learning AMIL model was developed, combining US images and MT to classify the nodules as benign or malignant and enhance the malignancy risk stratification of MT. Results: The final AMIL model matched MT sensitivity (0.946) while significantly improving PPV (0.477 vs 0.448 for MT alone), indicating fewer false positives while maintaining high sensitivity. Conclusion: Our approach reduces false positives compared to MT while maintaining the same ability to identify positive cases, potentially reducing unnecessary benign thyroid resections in patients with indeterminate nodules.

q-bio.QM

Mask Grounding for Referring Image Segmentation

Referring Image Segmentation (RIS) is a challenging task that requires an algorithm to segment objects referred by free-form language expressions. Despite significant progress in recent years, most state-of-the-art (SOTA) methods still suffer from considerable language-image modality gap at the pixel and word level. These methods generally 1) rely on sentence-level language features for language-image alignment and 2) lack explicit training supervision for fine-grained visual grounding. Consequently, they exhibit weak object-level correspondence between visual and language features. Without well-grounded features, prior methods struggle to understand complex expressions that require strong reasoning over relationships among multiple objects, especially when dealing with rarely used or ambiguous clauses. To tackle this challenge, we introduce a novel Mask Grounding auxiliary task that significantly improves visual grounding within language features, by explicitly teaching the model to learn fine-grained correspondence between masked textual tokens and their matching visual objects. Mask Grounding can be directly used on prior RIS methods and consistently bring improvements. Furthermore, to holistically address the modality gap, we also design a cross-modal alignment loss and an accompanying alignment module. These additions work synergistically with Mask Grounding. With all these techniques, our comprehensive approach culminates in MagNet (Mask-grounded Network), an architecture that significantly outperforms prior arts on three key benchmarks (RefCOCO, RefCOCO+ and G-Ref), demonstrating our method's effectiveness in addressing current limitations of RIS algorithms. Our code and pre-trained weights will be released.

cs.CV

Peeking into the next decade in Large-Scale Structure Cosmology with its Effective Field Theory

After the successful full-shape analyses of BOSS data using the Effective Field Theory of Large-Scale Structure, we investigate what upcoming galaxy surveys might achieve. We introduce a ``perturbativity prior" that ensures that loop terms are as large as theoretically expected, which is effective in the case of a large number of EFT parameters. After validating our technique by comparison with already-performed analyses of BOSS data, we provide Fisher forecasts using the one-loop prediction for power spectrum and bispectrum for two benchmark surveys: DESI and MegaMapper. We find overall great improvements on the cosmological parameters. In particular, we find that MegaMapper (DESI) should obtain at least a 12$\sigma$ ($2\sigma$) evidence for non-vanishing neutrino masses, bound the curvature $\Omega_k$ to 0.0012 (0.012), and primordial inflationary non-Gaussianities as follows: $f_{\text{NL}}^{\text{loc.}}$ to $\pm 0.26$ (3.3), $f_{\text{NL}}^{\text{eq.}}$ to $\pm16$ (92), $f_{\text{NL}}^{\text{orth.}}$ to $\pm 4.2$ (27). Such measurements would provide much insight on the theory of Inflation. We investigate the limiting factor of shot noise and ignorance of the EFT parameters.

astro-ph.CO

Efficiently evaluating loop integrals in the EFTofLSS using QFT integrals with massive propagators

We develop a new way to analytically calculate loop integrals in the Effective Field Theory of Large Scale-Structure. Previous available methods show severe limitations beyond the one-loop power spectrum due to analytical challenges and computational and memory costs. Our new method is based on fitting the linear power spectrum with cosmology-independent functions that resemble integer powers of quantum field theory massive propagators with complex masses. A remarkable small number of them is sufficient to reach enough accuracy. Similarly to former approaches, the cosmology dependence is encoded in the coordinate vector of the expansion of the linear power spectrum in our basis. We first produce cosmology-independent tensors where each entry is the loop integral evaluated on a given combination of basis vectors. For each cosmology, the evaluation of a loop integral amounts to contracting this tensor with the coordinate vector of the linear power spectrum. The 3-dimensional loop integrals for our basis functions can be evaluated using techniques familiar to particle physics, such as recursion relations and Feynman parametrization. We apply our formalism to evaluate the one-loop bispectrum of galaxies in redshift space. The final analytical expressions are quite simple and can be evaluated with little computational and memory cost. We show that the same expressions resolve the integration of all one-loop $N$-point function in the EFTofLSS. This method, which is originally presented here, has already been applied in the first one-loop bispectrum analysis of the BOSS data to constraint $Λ$CDM parameters and primordial non-Gaussianities, see arXiv:2206.08327 and arXiv:2201.11518.

astro-ph.CO

Joint Representation Learning for Text and 3D Point Cloud

Recent advancements in vision-language pre-training (e.g. CLIP) have shown that vision models can benefit from language supervision. While many models using language modality have achieved great success on 2D vision tasks, the joint representation learning of 3D point cloud with text remains under-explored due to the difficulty of 3D-Text data pair acquisition and the irregularity of 3D data structure. In this paper, we propose a novel Text4Point framework to construct language-guided 3D point cloud models. The key idea is utilizing 2D images as a bridge to connect the point cloud and the language modalities. The proposed Text4Point follows the pre-training and fine-tuning paradigm. During the pre-training stage, we establish the correspondence of images and point clouds based on the readily available RGB-D data and use contrastive learning to align the image and point cloud representations. Together with the well-aligned image and text features achieved by CLIP, the point cloud features are implicitly aligned with the text embeddings. Further, we propose a Text Querying Module to integrate language information into 3D representation learning by querying text embeddings with point cloud features. For fine-tuning, the model learns task-specific 3D representations under informative language guidance from the label set without 2D images. Extensive experiments demonstrate that our model shows consistent improvement on various downstream tasks, such as point cloud semantic segmentation, instance segmentation, and object detection. The code will be available here: https://github.com/LeapLabTHU/Text4Point

cs.CV

Robotic Dough Shaping

Robotic manipulation of deformable objects gains great attention due to its wide applications including medical surgery, home assistance, and automatic food preparation. The ability to deform soft objects remains a great challenge for robots due to difficulties in defining the problem mathematically. In this paper, we address the problem of shaping a piece of dough-like deformable material into a 2D target shape presented upfront. We use a 6 degree-of-freedom WidowX-250 Robot Arm equipped with a rolling pin and information collected from an RGB-D camera and a tactile sensor. We present and compare several control policies, including a dough shrinking action, in extensive experiments across three kinds of deformable materials and across three target dough shape sizes, achieving the intersection over union (IoU) of 0.90. Our results show that: i) rolling dough from the highest dough point is more efficient than from the 2D/3D dough centroid; ii) it might be better to stop the roll movement at the current dough boundary as opposed to the target shape outline; iii) the shrink action might be beneficial only if properly tuned with respect to the expand action; and iv) the Play-Doh material is easier to shape to a target shape as compared to Plasticine or Kinetic sand. Video demonstrations of our work are available at https://youtu.be/ZzLMxuITdt4

cs.RO

The Hubble Tension in Light of the Full-Shape Analysis of Large-Scale Structure Data

The disagreement between direct late-time measurements of the Hubble constant from the SH0ES collaboration, and early-universe measurements based on the $Λ$CDM model from the Planck collaboration might, at least in principle, be explained by new physics in the early universe. Recently, the application of the Effective Field Theory of Large-Scale Structure to the full shape of the power spectrum of the SDSS/BOSS data has revealed a new, rather powerful, way to measure the Hubble constant and the other cosmological parameters from Large-Scale Structure surveys. In light of this, we analyze two models for early universe physics, Early Dark Energy and Rock 'n' Roll, that were designed to significantly ameliorate the Hubble tension. Upon including the information from the full shape to the Planck, BAO, and Supernovae measurements, we find that the degeneracies in the cosmological parameters that were introduced by these models are well broken by the data, so that these two models do not significantly ameliorate the tension.

astro-ph.CO