SearcharxivSearch

arXiv subjects

Yi Hua

Publications and source records attributed to Yi Hua.

14 recordsLinked to original sources

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate complete shapes that are often misaligned with the input. We introduce World Tracing, a generative pixel-aligned geometry representation that predicts 3D points aligned with observed pixels while completing geometry beyond the visible surface. For each input pixel, World Tracing predicts an ordered stack of camera-space 3D points, where the first layer represents the visible surface and subsequent layers represent front-to-back intersections with occluded surfaces. We instantiate this representation with a world-tracing diffusion transformer, WT-DiT, which treats multiple geometry layers as separate denoising tokens coupled through factorized and global attention. WT-DiT is trained with pixel-space flow matching and a mixed noise schedule that balances visible-surface reconstruction with occluded-geometry generation. World Tracing achieves strong performance on visible-surface reconstruction and complete geometry generation across object, scene, and dynamic benchmarks, outperforming both depth predictors and image-to-3D generators. It also preserves 2D-to-3D correspondence, enabling text-driven 3D scene editing, geometry-conditioned novel-view video synthesis, and training-free integration with textured-mesh generators.

cs.CV

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

The evolution of Multimodal Large Language Models (MLLMs) has shifted the focus from text generation to active behavioral execution, particularly via OS agents navigating complex GUIs. However, the transition of these agents into trustworthy daily partners is hindered by a lack of rigorous evaluation regarding safety, efficiency, and multi-modal robustness. Current benchmarks suffer from narrow safety scenarios, noisy trajectory labeling, and limited robustness metrics. To bridge this gap, we propose OS-SPEAR, a comprehensive toolkit for the systematic analysis of OS agents across four dimensions: Safety, Performance, Efficiency, and Robustness. OS-SPEAR introduces four specialized subsets: (1) a S(afety)-subset encompassing diverse environment- and human-induced hazards; (2) a P(erformance)-subset curated via trajectory value estimation and stratified sampling; (3) an E(fficiency)-subset quantifying performance through the dual lenses of temporal latency and token consumption; and (4) a R(obustness)-subset that applies cross-modal disturbances to both visual and textual inputs. Additionally, we provide an automated analysis tool to generate human-readable diagnostic reports. We conduct an extensive evaluation of 22 popular OS agents using OS-SPEAR. Our empirical results reveal critical insights into the current landscape: notably, a prevalent trade-off between efficiency and safety or robustness, the performance superiority of specialized agents over general-purpose models, and varying robustness vulnerabilities across different modalities. By providing a multidimensional ranking and a standardized evaluation framework, OS-SPEAR offers a foundational resource for developing the next generation of reliable and efficient OS agents. The dataset and codes are available at https://github.com/Wuzheng02/OS-SPEAR.

cs.CL

Apple Intelligence Foundation Language Models: Tech Report 2025

We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architectural innovations such as KV-cache sharing and 2-bit quantization-aware training; and ii a scalable server model built on a novel Parallel-Track Mixture-of-Experts PT-MoE transformer that combines track parallelism, mixture-of-experts sparse computation, and interleaved global-local attention to deliver high quality with competitive cost on Apple's Private Cloud Compute platform. Both models are trained on large-scale multilingual and multimodal datasets sourced via responsible web crawling, licensed corpora, and high-quality synthetic data, then further refined with supervised fine-tuning and reinforcement learning on a new asynchronous platform. The resulting models support several additional languages while understanding images and executing tool calls. In public benchmarks and human evaluations, both the server model and the on-device model match or surpass comparably sized open baselines. A new Swift-centric Foundation Models framework exposes guided generation, constrained tool calling, and LoRA adapter fine-tuning, allowing developers to integrate these capabilities with a few lines of code. The latest advancements in Apple Intelligence models are grounded in our Responsible AI approach with safeguards like content filtering and locale-specific evaluation, as well as our commitment to protecting our users' privacy with innovations like Private Cloud Compute.

cs.LG

Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models

Multi-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently memorize private content that is entirely irrelevant to the training tasks. In this paper, we investigate how randomly generated task-irrelevant private content can become spuriously correlated with downstream objectives due to partial mini-batch training dynamics, thus causing inadvertent memorization. Concretely, we randomly generate task-irrelevant watermarks into VQA fine-tuning images at varying probabilities and propose a novel probing framework to determine whether MLLMs have inadvertently encoded such content. Our experiments reveal that MLLMs exhibit notably different training behaviors in partial mini-batch settings with task-irrelevant watermarks embedded. Furthermore, through layer-wise probing, we demonstrate that MLLMs trigger distinct representational patterns when encountering previously seen task-irrelevant knowledge, even if this knowledge does not influence their output during prompting. Our code is available at https://github.com/illusionhi/ProbingPrivacy.

cs.CV

fNeRF: High Quality Radiance Fields from Practical Cameras

In recent years, the development of Neural Radiance Fields has enabled a previously unseen level of photo-realistic 3D reconstruction of scenes and objects from multi-view camera data. However, previous methods use an oversimplified pinhole camera model resulting in defocus blur being `baked' into the reconstructed radiance field. We propose a modification to the ray casting that leverages the optics of lenses to enhance scene reconstruction in the presence of defocus blur. This allows us to improve the quality of radiance field reconstructions from the measurements of a practical camera with finite aperture. We show that the proposed model matches the defocus blur behavior of practical cameras more closely than pinhole models and other approximations of defocus blur models, particularly in the presence of partial occlusions. This allows us to achieve sharper reconstructions, improving the PSNR on validation of all-in-focus images, on both synthetic and real datasets, by up to 3 dB.

cs.CV

Performance Enhancement via XPM Suppression in a Linear all-PM NPE Mode-locked Fiber Oscillator

We demonstrate strong performance enhancement of an all polarization-maintaining fiber oscillator mode-locked using NPE in a linear self-stabilized fiber interferometer via suppression of cross-phase modulation (XPM). Numerical simulations reveal that XPM significantly affects the saturable absorber dynamics resulting in distortions of mode-locked steady-states. In the experiment, we construct an oscillator with XPM suppression, employing an intra-cavity YVO4 crystal, and compare its characteristics with a reference oscillator in standard configuration. It is shown, that XPM suppression not only lowers the mode-locking threshold by more than 40%, but further results in improved spectral pulse quality at the output ports and reduced nonlinear loss of the artificial saturable absorber.

physics.optics

Angle Sensitive Pixels for Lensless Imaging on Spherical Sensors

We propose OrbCam, a lensless architecture for imaging with spherical sensors. Prior work in lensless imager techniques have focused largely on using planar sensors; for such designs, it is important to use a modulation element, e.g. amplitude or phase masks, to construct a invertible imaging system. In contrast, we show that the diversity of pixel orientations on a curved surface is sufficient to improve the conditioning of the mapping between the scene and the sensor. Hence, when imaging on a spherical sensor, all pixels can have the same angular response function such that the lensless imager is comprised of pixels that are identical to each other and differ only in their orientations. We provide the computational tools for the design of the angular response of the pixels in a spherical sensor that leads to well-conditioned and noise-robust measurements. We validate our design in both simulation and a lab prototype. The implications of our design is that the lensless imaging can be enabled easily for curved and flexible surfaces thereby opening up a new set of application domains.

cs.CV

Large-mode-area Soliton Fiber Oscillator Mode-locked with Linear Self-stabilized Interferometer

In this work, we investigate an approach to scale up the output pulse energy in an all polarization-maintaining 17 MHz Yb-doped fiber oscillator via implementation of 25 um core-diameter large-mode-area fibers. The artificial saturable absorber in form of a Kerr-type self-stabilized fiber-interferometer enables highly stable mode-locked steady-states in the soliton-like operation regime with 170 mW average output power and a total output pulse energy of ~10 nJ distributed between two output ports. An experimental parameter comparison with a reference oscillator made of 5.5 um core-sized standard fiber-components reveals an increase of pulse energy by a factor of 36 with simultaneously reduced intensity-noise in the high frequency range > 100 kHz.

physics.optics

All-PM Divided Pulse Fiber Oscillator Mode-locked with the Optical Kerr-effect

In this letter, we investigate a Yb-doped mode-locked fiber oscillator that uses coherent pulse division and recombination to avoid excessive nonlinear phase shifts. The mode-locking mechanism of the laser is based on the accumulation of a differential nonlinear phase between orthogonal polarization modes in the polarization-maintaining fiber segment. The inserted coherent pulse divider, based on YVO4-crystals rotated successively by 45°, enables stable and undistorted mode-locked steady-states. The output pulse energy is increased from 89 pJ in the non-divided operation by ~6.5 dB to more than 400 pJ with three divisions. Measurements of the amplitude-fluctuations reveal a simultaneous broadband reduction of up to ~9 dB in the frequency range from 10 kHz to 2MHz.

physics.optics

A Simple Framework for 3D Lensless Imaging with Programmable Masks

Lensless cameras provide a framework to build thin imaging systems by replacing the lens in a conventional camera with an amplitude or phase mask near the sensor. Existing methods for lensless imaging can recover the depth and intensity of the scene, but they require solving computationally-expensive inverse problems. Furthermore, existing methods struggle to recover dense scenes with large depth variations. In this paper, we propose a lensless imaging system that captures a small number of measurements using different patterns on a programmable mask. In this context, we make three contributions. First, we present a fast recovery algorithm to recover textures on a fixed number of depth planes in the scene. Second, we consider the mask design problem, for programmable lensless cameras, and provide a design template for optimizing the mask patterns with the goal of improving depth estimation. Third, we use a refinement network as a post-processing step to identify and remove artifacts in the reconstruction. These modifications are evaluated extensively with experimental results on a lensless camera prototype to showcase the performance benefits of the optimized masks and recovery algorithms over the state of the art.

eess.IV

Nonlinear Fiber System for Shot-noise Limited Intensity Noise Suppression and Amplification

We propose a nonlinear fiber system for shot-noise limited, all-optical intensity-noise reduction and signal amplification. The mechanism is based on the accumulation of different nonlinear phase shifts between orthogonal polarization modes in a polarization-maintaining fiber amplifier in combination with an implemented sinusoidal transmission-function. The resulting correlation between the input intensity-fluctuations and the system transmission enables tunable intensity noise reduction of the input pulse train. In the experiment, the noise spectral density of a mode-locked oscillator is suppressed by up to ~20 dB to the theoretical shot-noise limit of the measurement at -151.3 dBc/Hz with simultaneous pulse amplification of 13.5dB.

physics.optics

Compact, all-PM fiber integrated and alignment-free ultrafast Yb:fiber NALM laser with sub-femtosecond timing jitter

We report a simple and compact design of a dispersion compensated mode-locked Yb:fiber oscillator based on a nonlinear amplifying loop mirror (NALM). The fully polarization maintaining (PM) fiber integrated laser features a chirped fiber Bragg grating (CFBG) for dispersion compensation and a fiber integrated compact non-reciprocal phase bias device, which is alignment-free. The main design parameters were determined by numerically simulating the pulse evolution in the oscillator and by analyzing their impact on the laser performance. Experimentally, we achieved an 88 fs compressed pulse duration with sub-fs timing jitter at 54 MHz repetition rate and 51 mW of output power with 5.5 * 10-5 [20 Hz, 1 MHz] integrated relative intensity noise (RIN). Furthermore, we demonstrate tight phase-locking of the laser's carrier-envelope offset frequency (fceo) to a stable radio frequency (RF) reference and of one frequency comb tooth to a stable optical reference at 291 THz.

physics.optics

Intrinsic Amplitude-Noise Suppression in Fiber Lasers Mode-locked with Nonlinear Amplifying Loop Mirrors

In this work, we investigate the steady-states of a fiber lasers mode-locked with a nonlinear amplifying loop-mirror that has an inherent amplitude noise-suppression mechanism. Due to the interaction of the sinusoidal transmission function with the fluctuating intracavity pulse amplitude we show that this mechanism may lead to a detectable difference in relative intensity noise at the reflected and transmitted output port under specific preconditions. We present systematic intensity noise measurements with a nonlinear fiber-based system that replicates a single roundtrip in the laser cavity. Experimental results and simulations clearly show a reduction of the intracavity amplitude fluctuations up to 4 dB for certain steady-states.

physics.optics

Segmented Terahertz Electron Accelerator and Manipulator (STEAM)

Acceleration and manipulation of ultrashort electron bunches are the basis behind electron and X-ray devices used for ultrafast, atomic-scale imaging and spectroscopy. Using laser-generated THz drivers enables intrinsic synchronization as well as dramatic gains in field strengths, field gradients and component compactness, leading to shorter electron bunches, higher spatio-temporal resolution and smaller infrastructures. We present a segmented THz electron accelerator and manipulator (STEAM) with extended interaction lengths capable of performing multiple high-field operations on the energy and phase-space of ultrashort bunches with moderate charge. With this single device, powered by few-microjoule, single-cycle, 0.3 THz pulses, we demonstrate record THz-device acceleration of >30 keV, streaking with <10 fs resolution, focusing with >2 kT/m strengths, compression to ~100 fs as well as real-time switching between these modes of operation. The STEAM device demonstrates the feasibility of future THz-based compact electron guns, accelerators, ultrafast electron diffractometers and Free-Electron Lasers with transformative impact.

physics.acc-ph