SearcharxivSearch

arXiv subjects

Cheng Zheng

Publications and source records attributed to Cheng Zheng.

At least 19 recordsLinked to original sources

Bridging Probabilistic LLMs and Deterministic Statistical Validation: The PROVE Multi-Agent Framework for Clinical Trial Reporting

Ensuring the accuracy and consistency of clinical trial Tables, Figures, and Listings (TFLs) remains a major challenge in regulatory reporting. Independent programming and manual review are essential quality-control practices, but cross-output verification still depends heavily on reviewer inspection and may miss structural, logical, or arithmetic discrepancies. Large language models (LLMs) can help interpret varied table language and navigate lengthy study documents, but they are not reliable substitutes for programmed statistical checks. We introduce PROVE (Programmatic Reporting and Output Verification Engine), an auditable framework that uses optional LLM and retrieval support for table interpretation while reserving numerical and logical decisions for programmed validators. PROVE links findings to source evidence, supports cross-output consistency checks, and allows LLM use to be enabled or disabled based on study requirements. We evaluated PROVE using ten replicated synthetic oncology reporting packages generated from raw data through SDTM, ADaM, and TFL outputs, with paired clean and discrepancy-injected packages; each replicate included 15 randomly injected discrepancies. We examined two table-label settings: exact labels matching the validator vocabulary and labels with similar clinical meaning but different wording. Within the implemented rule classes, all automated PROVE variants achieved perfect classification in the exact-label setting. In the label-variation setting, LLM-assisted semantic matching improved overall recall from 0.588 to 0.993 and overall F1 from 0.735 to 0.996 compared with exact-match, fuzzy lexical, and embedding-similarity variants. These findings suggest that LLMs are most useful for interpreting real-world variation in TFL wording and formatting, while executable checks should remain responsible for final numerical validation.

stat.AP

A Criterion for Equidistribution along the $\Omega$ Function over Polynomial Sequences with Applications

Let $P (Y_1, ..., Y_d)$ be a certain fixed homogeneous polynomial of integral coefficients. In this paper, we establish a quantitative equidistribution criterion for the ergodic averages along $\Omega (|P (n_1, ..., n_d)|)$. Consequently, by an estimate of Lachand, we prove the following variant of a theorem of Bergelson and Richter: if $P$ is an irreducible binary cubic form and $ (X, T)$ is a uniquely ergodic system with unique invariant measure $\mu$, then for any $x \in X$ and $f \in C(X)$, \begin{equation*} \lim_{N \rightarrow \infty} \frac 1 {N^2} {\mathop{\sum\sum}_{n_1, n_2 \leqslant N}} f \big( T^{ \Omega (|P (n_1, n_2)| ) } x \big) = \int_{X} f \ \mathrm{d} \mu . \end{equation*} Moreover, we prove in the appendix a related conjecture of C\'espedes and Donoso over number fields.

math.DS

Joint Model for Mediation Analysis with Causally Related Longitudinal and Recurrent Event Mediators for Survival Outcome

Recurrent events and repeated measures are commonly encountered in clinical longitudinal studies, often holding strong associations with patient outcomes. Although joint models for repeated measures, recurrent events, and a terminal event have been developed to account for their correlation, limited methodologies exist to examine causal mediation mechanisms involving multiple types of mediators, especially when mediators are causally related. This study addresses this gap by proposing a novel causal mediation analysis framework to quantify natural direct and indirect effects when both recurrent events and repeated measures act as mediators with causal dependencies. We extend joint modeling approaches by incorporating shared random effects (frailties) structures, relaxing the commonly used ``sequential ignorability" assumption, and accounting for unmeasured time-independent confounders through shared random effects. We apply our method to the Terry Beirn Community Programs for Clinical Research on AIDS (CPCRA) study and demonstrate that both recurrent opportunistic infections (OIs) and repeated CD4 measurements mediate the effects of prior AIDS-defining conditions on survival outcomes. Additionally, the shared random effects between repeated CD4 and survival models highlight the presence of unmeasured confounding between CD4 counts and mortality. Simulation studies demonstrate the robustness and finite sample performance of our estimators for natural direct and indirect effects. The proposed methodology enables a more comprehensive investigation of causal pathways in longitudinal studies with multiple mediators, providing insights into treatment mechanisms and informing clinical decision-making.

stat.ME

Field Demonstration of a Multi-User Continuous-Variable Quantum Access Network for Quantum-to-the-Home

Realizing scalable Quantum-to-the-Home (QTTH) faces a bottleneck: link asymmetry in broadcast continuous-variable quantum access networks (CV-QANs) hinders the selection of a globally optimal modulation variance. We demonstrate a downstream broadcast CV-QAN connecting a Quantum Line Terminal (QLT) to multiple Quantum Network Units (QNUs) over commercial fiber. Operating within a trusted local network domain, we establish a multi-user utility model to select the optimal shared variance, balancing network efficiency and user fairness. Supported by robust digital signal processing, our 1:16 field trial achieves Mbit/s-level asymptotic secure key rates, bridging theoretical protocols with Fiber-to-the-Home reality and guiding future scalable access architectures.

quant-ph

Joint analysis for multivariate longitudinal and event time data with a change point anchored at interval-censored event time

Huntington's disease (HD) is an autosomal dominant neurodegenerative disorder characterized by motor dysfunction, psychiatric disturbances, and cognitive decline. The onset of HD is marked by severe motor impairment, which may be predicted by prior cognitive decline and, in turn, exacerbate cognitive deficits. Clinical data, however, are often collected at discrete time points, so the timing of disease onset is subject to interval censoring. To address the challenges posed by such data, we develop a joint model for multivariate longitudinal biomarkers with a change point anchored at an interval-censored event time. The model simultaneously assesses the effects of longitudinal biomarkers on the event time and the changes in biomarker trajectories following the event. We conduct a comprehensive simulation study to demonstrate the finite-sample performance of the proposed method for causal inference. Finally, we apply the method to PREDICT-HD, a multisite observational cohort study of prodromal HD individuals, to ascertain how cognitive impairment and motor dysfunction interact during disease progression.

stat.ME

HEIR: Learning Graph-Based Motion Hierarchies

Hierarchical structures of motion exist across research fields, including computer vision, graphics, and robotics, where complex dynamics typically arise from coordinated interactions among simpler motion components. Existing methods to model such dynamics typically rely on manually-defined or heuristic hierarchies with fixed motion primitives, limiting their generalizability across different tasks. In this work, we propose a general hierarchical motion modeling method that learns structured, interpretable motion relationships directly from data. Our method represents observed motions using graph-based hierarchies, explicitly decomposing global absolute motions into parent-inherited patterns and local motion residuals. We formulate hierarchy inference as a differentiable graph learning problem, where vertices represent elemental motions and directed edges capture learned parent-child dependencies through graph neural networks. We evaluate our hierarchical reconstruction approach on three examples: 1D translational motion, 2D rotational motion, and dynamic 3D scene deformation via Gaussian splatting. Experimental results show that our method reconstructs the intrinsic motion hierarchy in 1D and 2D cases, and produces more realistic and interpretable deformations compared to the baseline on dynamic 3D Gaussian splatting scenes. By providing an adaptable, data-driven hierarchical modeling paradigm, our method offers a formulation applicable to a broad range of motion-centric tasks. Project Page: https://light.princeton.edu/HEIR/

cs.CV

Collaborative On-Sensor Array Cameras

Modern nanofabrication techniques have enabled us to manipulate the wavefront of light with sub-wavelength-scale structures, offering the potential to replace bulky refractive surfaces in conventional optics with ultrathin metasurfaces. In theory, arrays of nanoposts provide unprecedented control over manipulating the wavefront in terms of phase, polarization, and amplitude at the nanometer resolution. A line of recent work successfully investigates flat computational cameras that replace compound lenses with a single metalens or an array of metasurfaces a few millimeters from the sensor. However, due to the inherent wavelength dependence of metalenses, in practice, these cameras do not match their refractive counterparts in image quality for broadband imaging, and may even suffer from hallucinations when relying on generative reconstruction methods. In this work, we investigate a collaborative array of metasurface elements that are jointly learned to perform broadband imaging. To this end, we learn a nanophotonics array with 100-million nanoposts that is end-to-end jointly optimized over the full visible spectrum--a design task that existing inverse design methods or learning approaches cannot support due to memory and compute limitations. We introduce a distributed meta-optics learning method to tackle this challenge. This allows us to optimize a large parameter array along with a learned meta-atom proxy and a non-generative reconstruction method that is parallax-aware and noise-aware. The proposed camera performs favorably in simulation and in all experimental tests irrespective of the scene illumination spectrum.

physics.optics

Model-free High Dimensional Mediator Selection with False Discovery Rate Control

There is a challenge in selecting high-dimensional mediators when the mediators have complex correlation structures and interactions. In this work, we frame the high-dimensional mediator selection problem into a series of hypothesis tests with composite nulls, and develop a method to control the false discovery rate (FDR) which has mild assumptions on the mediation model. We show the theoretical guarantee that the proposed method and algorithm achieve FDR control. We present extensive simulation results to demonstrate the power and finite sample performance compared with existing methods. Lastly, we demonstrate the method for analyzing the Alzheimer's Disease Neuroimaging Initiative (ADNI) data, in which the proposed method selects the volume of the hippocampus and amygdala, as well as some other important MRI-derived measures as mediators for the relationship between gender and dementia progression.

stat.ME

Can Video Diffusion Model Reconstruct 4D Geometry?

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods either require specialized 4D representation or sophisticated optimization. In this paper, we present Sora3R, a novel framework that taps into the rich spatiotemporal priors of large-scale video diffusion models to directly infer 4D pointmaps from casual videos. Sora3R follows a two-stage pipeline: (1) we adapt a pointmap VAE from a pretrained video VAE, ensuring compatibility between the geometry and video latent spaces; (2) we finetune a diffusion backbone in combined video and pointmap latent space to generate coherent 4D pointmaps for every frame. Sora3R operates in a fully feedforward manner, requiring no external modules (e.g., depth, optical flow, or segmentation) or iterative global alignment. Extensive experiments demonstrate that Sora3R reliably recovers both camera poses and detailed scene geometry, achieving performance on par with state-of-the-art methods for dynamic 4D reconstruction across diverse scenarios.

cs.CV

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities. However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects (3D objects with temporal evolution over time). In this paper, we introduce 4D-Bench, the first benchmark to evaluate the capabilities of MLLMs in 4D object understanding, featuring tasks in 4D object Question Answering (4D object QA) and 4D object captioning. 4D-Bench provides 4D objects with diverse categories, high-quality annotations, and tasks necessitating multi-view spatial-temporal understanding, different from existing 2D image/video-based benchmarks. With 4D-Bench, we evaluate a wide range of open-source and closed-source MLLMs. The results from the 4D object captioning experiment indicate that MLLMs generally exhibit weaker temporal understanding compared to their appearance understanding, notably, while open-source models approach closed-source performance in appearance understanding, they show larger performance gaps in temporal understanding. 4D object QA yields surprising findings: even with simple single-object videos, MLLMs perform poorly, with state-of-the-art GPT-4o achieving only 63\% accuracy compared to the human baseline of 91\%. These findings highlight a substantial gap in 4D object understanding and the need for further advancements in MLLMs.

cs.CV

Landau-Level Quantization and Band Splitting of FeSe Monolayers Revealed by Scanning Tunneling Spectroscopy

Two-dimensional (2D) superconductors that reside on substrates must be influenced by Rashba spin-orbit coupling (SOC). The intriguing effect of Rashba-type SOCs on iron-based superconductors (IBSs) has remained largely a mystery. In this work, we unveil modified Landau-level spectroscopy and the intricate band splitting of FeSe monolayers through the precision of scanning tunneling spectroscopy, which unequivocally demonstrates the presence of Rashba SOC. The discovery sheds light on a nonparabolic electron band at the X/Y point, displaying a distinctive Landau quantization behavior characterized by $E_n\propto(nB)^{4/3}$. The theoretical model aligns with our experimental insights, positing that the k$^4$-term of the electron band becomes predominant and profoundly reshapes the band structure. Our research underscores the pivotal role of the Rashba SOC effect on 2D superconductors and sets the stage to probe new quantum states in systems with remarkably low carrier concentrations.

cond-mat.supr-con

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeling such multi-dimensional distribution. To this end, we propose a novel diffusion-based pipeline that generates high-quality multi-view videos centered around a dynamic 3D object from text. Specifically, we factor the T2MVid problem into viewpoint-space and time components. Such factorization allows us to combine and reuse layers of advanced pre-trained multi-view image and 2D video diffusion models to ensure multi-view consistency as well as temporal coherence for the generated multi-view videos, largely reducing the training cost. We further introduce alignment modules to align the latent spaces of layers from the pre-trained multi-view and the 2D video diffusion models, addressing the reused layers' incompatibility that arises from the domain gap between 2D and multi-view data. In support of this and future research, we further contribute a captioned multi-view video dataset. Experimental results demonstrate that our method generates high-quality multi-view videos, exhibiting vivid motions, temporal coherence, and multi-view consistency, given a variety of text prompts.

cs.CV

A shrinking target problem in homogeneous spaces of semisimple algebraic groups

In this paper, we study a shrinking target problem with target at infinity in a homogeneous space of a semisimple algebraic group from the representation-theoretic point of view. Let $\rho:\mathbf G\to\mathbf{GL}(V)$ be an irreducible $\mathbb Q$-rational representation of a connected semisimple $\mathbb Q$-algebraic group $\mathbf G$ on a complex vector space $V$, $\{a_t\}_{t\in\mathbb R}$ a one-parameter subgroup in a $\mathbb Q$-split torus in $\mathbf G$ and $\psi:\mathbb R_+\to\mathbb R_+$ a positive function on $\mathbb R_+$. We define a subset $S_\rho(\psi)$ of $\psi$-Diophantine elements in $\mathbf G(\mathbb R)$ in terms of the representation $\rho$ and $\{a_t\}_{t\in\mathbb R}$, and prove formulas for the Hausdorff dimension of the complement of $S_{\rho}(\psi)$. We also discuss the connections of our results to Diophantine approximation on flag varieties and rational approximation to linear subspaces in Grassmann varieties.

math.DS

Variable selection with FDR control for noisy data -- an application to screening metabolites that are associated with breast and colorectal cancer

The rapidly expanding field of metabolomics presents an invaluable resource for understanding the associations between metabolites and various diseases. However, the high dimensionality, presence of missing values, and measurement errors associated with metabolomics data can present challenges in developing reliable and reproducible methodologies for disease association studies. Therefore, there is a compelling need to develop robust statistical methods that can navigate these complexities to achieve reliable and reproducible disease association studies. In this paper, we focus on developing such a methodology with an emphasis on controlling the False Discovery Rate during the screening of mutual metabolomic signals for multiple disease outcomes. We illustrate the versatility and performance of this procedure in a variety of scenarios, dealing with missing data and measurement errors. As a specific application of this novel methodology, we target two of the most prevalent cancers among US women: breast cancer and colorectal cancer. By applying our method to the Wome's Health Initiative data, we successfully identify metabolites that are associated with either or both of these cancers, demonstrating the practical utility and potential of our method in identifying consistent risk factors and understanding shared mechanisms between diseases.

stat.ME

Neural Lithography: Close the Design-to-Manufacturing Gap in Computational Optics with a 'Real2Sim' Learned Photolithography Simulator

We introduce neural lithography to address the 'design-to-manufacturing' gap in computational optics. Computational optics with large design degrees of freedom enable advanced functionalities and performance beyond traditional optics. However, the existing design approaches often overlook the numerical modeling of the manufacturing process, which can result in significant performance deviation between the design and the fabricated optics. To bridge this gap, we, for the first time, propose a fully differentiable design framework that integrates a pre-trained photolithography simulator into the model-based optical design loop. Leveraging a blend of physics-informed modeling and data-driven training using experimentally collected datasets, our photolithography simulator serves as a regularizer on fabrication feasibility during design, compensating for structure discrepancies introduced in the lithography process. We demonstrate the effectiveness of our approach through two typical tasks in computational optics, where we design and fabricate a holographic optical element (HOE) and a multi-level diffractive lens (MDL) using a two-photon lithography system, showcasing improved optical performance on the task-specific metrics.

physics.optics

Learning to Read Analog Gauges from Synthetic Data

Manually reading and logging gauge data is time inefficient, and the effort increases according to the number of gauges available. We present a computer vision pipeline that automates the reading of analog gauges. We propose a two-stage CNN pipeline that identifies the key structural components of an analog gauge and outputs an angular reading. To facilitate the training of our approach, a synthetic dataset is generated thus obtaining a set of realistic analog gauges with their corresponding annotation. To validate our proposal, an additional real-world dataset was collected with 4.813 manually curated images. When compared against state-of-the-art methodologies, our method shows a significant improvement of 4.55 in the average error, which is a 52% relative improvement. The resources for this project will be made available at: https://github.com/fuankarion/automatic-gauge-reading.

cs.CV

Controlling FDR in selecting group-level simultaneous signals from multiple data sources with application to the National Covid Collaborative Cohort data

One challenge in exploratory association studies using observational data is that the associations between the predictors and the outcome are potentially weak and rare, and the candidate predictors have complex correlation structures. False discovery rate (FDR) controlling procedures can provide important statistical guarantees for replicability in predictor identification in exploratory research. In the recently established National COVID Collaborative Cohort (N3C), electronic health record (EHR) data on the same set of candidate predictors are independently collected in multiple different sites, offering opportunities to identify true associations by combining information from different sources. This paper presents a general knockoff-based variable selection algorithm to identify associations from unions of group-level conditional independence tests (simultaneous signals) with exact FDR control guarantees under finite sample settings. This algorithm can work with general regression settings, allowing heterogeneity of both the predictors and the outcomes across multiple data sources. We demonstrate the performance of this method with extensive numerical studies and an application to the N3C data.

stat.ME