SearcharxivSearch

arXiv subjects

Weiwei Cai

Publications and source records attributed to Weiwei Cai.

At least 19 recordsLinked to original sources

Multidimensional Light Detection with Symmetry-engineered Heterojunctions

Miniaturised multidimensional light detection, encompassing full-Stokes polarimetry and spectroscopy in ultracompact footprints, is attracting growing interest for its potential to capture a comprehensive set of light properties in portable platforms. Although significant progress has been made in miniaturised schemes for independent polarisation and spectral detection, achieving simultaneous high-dimensional light detection remains an outstanding challenge that limits their development toward full integration. We overcome this limitation by breaking the rotational and inversion symmetries in a symmetry-engineered van der Waals heterojunction to realise an ultracompact multidimensional photodetector. The dual symmetry breaking gives rise to non-trivial quantum geometric and topological features, enabling simultaneous broadband polarisation- and spectrum-resolved light detection, in contrast to previous van der Waals material-based devices, which could detect only one of these modalities. Our device, with an effective area of only 10 micron x 10 micron, reconstructs full-Stokes polarisation with overall root-mean-square errors below 0.05 and resolves spectral peaks separated by 0.4 nm, capabilities not previously achieved in single-pixel detectors. By unifying high-fidelity polarimetry and sub-nanometre spectroscopy in a single electrically tunable junction, our work eliminates the need for cascaded detection architectures and establishes a foundation for multidimensional detector arrays for integrated photonics, quantum information processing, and precision imaging.

physics.optics

Differentiable eigendecomposition-free RCWA for full-tensor anisotropic photonics

Full-tensor anisotropy transforms rigorous coupled-wave analysis (RCWA) into a large, fully coupled non-Hermitian eigenproblem, making eigendecomposition expensive and difficult to differentiate. We introduce a differentiable, eigendecomposition-free RCWA framework for spatially patterned media with fully coupled permittivity tensors, using boundary fields rather than internal eigenmodes as the layer representation. A boundary-value cascade constructs scattering operators directly from these fields, enabling automatic differentiation and efficient GPU execution. Benchmarks against finite-element and transfer-matrix solutions show close agreement in scattering responses, while automatic-differentiation gradients agree with finite differences and enable topology optimization. At 529 Fourier harmonics, layer construction is 22.8 times faster than conventional eigendecomposition on the same GPU. Our framework offers a general computational route toward scalable forward modeling and inverse design in anisotropic photonic systems.

physics.optics

InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search

Training-free neural architecture search promises efficient discovery of high-performance networks without costly training. However, existing zero-cost proxies rely on fragmented heuristics that fail to capture the fundamental question: what makes an architecture trainable? This paper introduces Intrinsic Trainability (InTrain), a unified theoretical proxy that formalizes trainability as an architectural invariant emerging from two synergistic components: geometric capacity and optimization resilience. We operationalize intrinsic trainability through analysis of neural information processing. Geometric capacity is quantified via the participation ratio of activation covariance eigenspectrum, capturing the effective dimensionality of representation manifolds. Optimization resilience is measured through cumulative gradient health, assessing the robustness of backpropagation across network depth. InTrain synthesizes these dimensions through a scale-invariant multiplicative coupling, which we hypothesize is essential for capturing their synergistic, non-additive relationship. Extensive experiments on standard NAS benchmarks and search spaces demonstrate that InTrain achieves ranking correlations on par with state-of-the-art ensemble-based proxies and outperforms other single-metric methods.

cs.LG

MiCU: End-to-End Smart Home Command Understanding with Large Language Model

Command understanding systems in smart home ecosystems can automate device control and substantially improve user experience. However, while they perform well on precise utterances (e.g., "turn on the bedroom light"), they struggle with ambiguous or misaligned commands (e.g., "make the bedroom cozy"). Large language models (LLMs) generalize well across various domains and can outperform traditional rule-based systems on such tasks, but their effectiveness is often constrained by scarce domain-specific data, insufficient task-specific adaptation, and high computational costs. In this paper, we propose an automated training data synthesis workflow using user logs and LLMs; then we build MiCU, a domain-specific LLM that excels at command understanding. Specifically, we employ curriculum learning to inject domain knowledge into the base LLM, then we enhance its reasoning ability via cold-start training combined with reinforcement learning (RL) guided by domain-specific thinking rules. Additionally, we introduce a token compression technique that condenses device description into a single special token, substantially reducing inference overhead and enabling \model-fast, an efficient variant optimized for long inputs. Extensive experiments show that MiCU significantly outperforms baselines, with an average accuracy gain of 20.01% across all device categories. We have deployed MiCU in the Xiaomi Home app, receiving approximately 1.7 million page views per day. Production evaluations show that MiCU reduces user correction rate by 1.57% and increases human audited accuracy by 32.05%. Our data and code are available at https://github.com/xiaomi-research/iot_spec_llm

cs.CL

Volumetric Optical Scattering Neural Networks

Optical neural networks offer a route to low-latency and energy-efficient inference by encoding computation in light propagation. However, most existing implementations rely on planar photonic circuits or discretely spaced diffractive layers, restricting volumetric integration and imposing stringent alignment requirements. Here we demonstrate a volumetric optical scattering neural network (OSNN) in which densely packed weak scatterers form a three-dimensional, locally connected optical computing medium. In contrast to fully connected diffractive architectures, the OSNN uses near-field scattering interactions, described under the first-Born approximation, to compress optical interconnections into a monolithic volume. We implement this concept using resilient inverse design and two-photon nanolithography, yielding OSNN devices with a volume of ~$3.8*10^{-4}mm^{3}$ and a record-breaking neuron density of $1.0*10^{9}/mm^{3}$. Experimentally, the fabricated classifier achieves $94.8\%$ blind-test accuracy on MNIST, while the imager performs optical compressed imaging with a $1-{\mu}m$ effective resolution and average FSIM values of $0.93$ on Fashion-MNIST and $0.91$ on VesselMNIST3D. OSNN paves the way for ultra-dense, ultra-compact, and efficient optical computing, creating a universal platform for embedded optical intelligence and promising widespread application in AI fields ranging from autonomous driving to medical diagnosis.

physics.optics

Photocurrent at oblique illumination and reconstruction of wavefront direction with 2d photodetectors

Many contemporary photodetectors operate beyond the readout of light intensity and enable the reconstruction of spectrum and polarization at the single-pixel level. However, the determination of light incidence direction with reconstructive detectors has not been realized so far. We show that photodetectors based on symmetric junctions of metals and 2d electron systems (2DES) enable (1) zero-bias photocurrent at oblique light incidence (2) reconstruction of incidence direction based on photocurrent measurements at variable carrier density. The former effect is based on peculiar electrodynamics of metal-contacted 2DES, where spatial variations of incident field phase translate into strong variations of local field amplitude. The local absorbances at two opposite metal-2DES junctions at oblique incidence are dissimilar, which results in finite photocurrent independent of microscopic rectification mechanism at these junctions. The direction of photocurrent uniquely determines the quadrant of light incidence. Quantitative determination of incidence angle becomes possible under conditions of 2d plasmon resonance at variable carrier density. In such a case, obliquely incident radiation excites the asymmetric plasmon modes, which amplitude carries unique information about angle of incidence.

cond-mat.mes-hall

Native 3D Editing with Full Attention

Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimization-based approaches are prohibitively slow, while feed-forward approaches relying on multi-view 2D editing often suffer from inconsistent geometry and degraded visual quality. To address these issues, we propose a novel native 3D editing framework that directly manipulates 3D representations in a single, efficient feed-forward pass. Specifically, we create a large-scale, multi-modal dataset for instruction-guided 3D editing, covering diverse addition, deletion, and modification tasks. This dataset is meticulously curated to ensure that edited objects faithfully adhere to the instructional changes while preserving the consistency of unedited regions with the source object. Building upon this dataset, we explore two distinct conditioning strategies for our model: a conventional cross-attention mechanism and a novel 3D token concatenation approach. Our results demonstrate that token concatenation is more parameter-efficient and achieves superior performance. Extensive evaluations show that our method outperforms existing 2D-lifting approaches, setting a new benchmark in generation quality, 3D consistency, and instruction fidelity.

cs.CV

Reconstructive comb spectroscopy: A single-pixel detection paradigm beyond dual-comb limitations

Frequency comb spectroscopy has revolutionized broadband molecular fingerprinting with mode-defined resolution. While dual-comb spectroscopy stands as a dominant paradigm for high-resolution measurements, it relies on mutually coherent dual combs, and its applicability to non-cooperative sensing is limited by the requirement for phase-sensitive detection and controlled optical returns. Here, we introduce reconstructive comb spectroscopy, a fundamentally different paradigm that eliminates these constraints. By integrating a mode-programmable optical comb with a computational sensing scheme based on single-pixel detection, our method achieves picometer-level spectral resolution over a 10-nm (1.27-THz) instantaneous bandwidth, with single-photon sensitivity down to 10^-4 photons per pulse, and compressed spectral acquisition at 2.5% sampling while maintaining reconstruction errors below 10%. We demonstrate robust performance through scattering media and from non-cooperative targets. These capabilities establish reconstructive comb spectroscopy as a new platform for gas sensing, with broad applicability in remote atmospheric monitoring, industrial leak detection, and standoff chemical-threat identification.

physics.optics

Ultracompact Wide-FOV Near-infrared Camera with Wafer-level Manufactured Meta-Aspheric Lens

Overcoming the trade-off between wide field of view (FOV) and compactness remains a central challenge for integrating near-infrared (NIR) imaging into smartphones and AR glasses. Existing refractive NIR optics cannot simultaneously achieve ultra-wide angles above 100{\deg} and ultrathin total track length (TTL) below 5 mm, limiting their use in portable devices. Here, we present a wafer-level-manufactured meta-aspheric lens (MAL) that achieves a 101.5{\deg} FOV, 3.39 mm TTL, and F/1.64 aperture within a compact volume of 0.02 cubic centimeters. Unlike previous hybrid lenses with separate refractive and diffractive components, our MAL features a fully integrated structure, which enables a compact form factor. This integration also simplifies fabrication, allowing high-throughput production via micrometer-level precision alignment and bonding on a single wafer, with only one dicing step and no need for additional mechanical fixtures. Furthermore, the design process explicitly considers manufacturability and accurately models metalens dispersion, ensuring that experimental performance matches simulated results. We validate our MAL through both direct and computational imaging experiments. Despite its small form factor, our scalable MAL demonstrates strong NIR imaging performance in blood vessel imaging, eye tracking, and computational pixel super-resolution tasks. This scalable MAL technology establishes a new benchmark for high-performance, miniaturized NIR imaging and opens the door to next-generation smartphone and AR optical systems.

physics.optics

Prediction, Generation of WWTPs microbiome community structures and Clustering of WWTPs various feature attributes using DE-BP model, SiTime-GAN model and DPNG-EPMC ensemble clustering algorithm with modulation of microbial ecosystem health

Microbiomes not only underpin Earth's biogeochemical cycles but also play crucial roles in both engineered and natural ecosystems, such as the soil, wastewater treatment, and the human gut. However, microbiome engineering faces significant obstacles to surmount to deliver the desired improvements in microbiome control. Here, we use the backpropagation neural network (BPNN), optimized through differential evolution (DE-BP), to predict the microbial composition of activated sludge (AS) systems collected from wastewater treatment plants (WWTPs) located worldwide. Furthermore, we introduce a novel clustering algorithm termed Directional Position Nonlinear Emotional Preference Migration Behavior Clustering (DPNG-EPMC). This method is applied to conduct a clustering analysis of WWTPs across various feature attributes. Finally, we employ the Similar Time Generative Adversarial Networks (SiTime-GAN), to synthesize novel microbial compositions and feature attributes data. As a result, we demonstrate that the DE-BP model can provide superior predictions of the microbial composition. Additionally, we show that the DPNG-EPMC can be applied to the analysis of WWTPs under various feature attributes. Finally, we demonstrate that the SiTime-GAN model can generate valuable incremental synthetic data. Our results, obtained through predicting the microbial community and conducting analysis of WWTPs under various feature attributes, develop an understanding of the factors influencing AS communities.

cs.LG

ViStoryBench: Comprehensive Benchmark Suite for Story Visualization

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to short prompts, lacking character references, or single-image cases, failing to reflect real-world narrative complexity and obscuring true model performance.We introduce ViStoryBench, a comprehensive benchmark designed to evaluate story visualization models across varied narrative structures, visual styles, and character settings. It features richly annotated multi-shot scripts derived from curated stories spanning literature, film, and folklore. Large language models assist in story summarization and script generation, with all outputs verified by humans for coherence and fidelity. Character references are carefully curated to maintain consistency across different artistic styles. ViStoryBench proposes a suite of multi-dimensional automated metrics to evaluate character consistency, style similarity, prompt alignment, aesthetic quality, and artifacts like copy-paste behavior. These metrics are validated through human studies and used to assess a broad range of open-source and commercial models, enabling systematic analysis and encouraging advances in visual storytelling.

cs.CV

Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets

While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamental challenges such as data scarcity, algorithmic limitations, and ecosystem fragmentation. To this end, we present Step1X-3D, an open framework addressing these challenges through: (1) a rigorous data curation pipeline processing >5M assets to create a 2M high-quality dataset with standardized geometric and textural properties; (2) a two-stage 3D-native architecture combining a hybrid VAE-DiT geometry generator with an diffusion-based texture synthesis module; and (3) the full open-source release of models, training code, and adaptation modules. For geometry generation, the hybrid VAE-DiT component produces TSDF representations by employing perceiver-based latent encoding with sharp edge sampling for detail preservation. The diffusion-based texture synthesis module then ensures cross-view consistency through geometric conditioning and latent-space synchronization. Benchmark results demonstrate state-of-the-art performance that exceeds existing open-source methods, while also achieving competitive quality with proprietary solutions. Notably, the framework uniquely bridges the 2D and 3D generation paradigms by supporting direct transfer of 2D control techniques~(e.g., LoRA) to 3D synthesis. By simultaneously advancing data quality, algorithmic fidelity, and reproducibility, Step1X-3D aims to establish new standards for open research in controllable 3D asset generation.

cs.CV

Achromatic single-layer hologram

Phase retrieval is a fundamental technique of advanced optical technologies, enabling precise control over wavefront properties. A persistent challenge in diffractive optical element (DOE) design is that a single hologram typically operates within a single wavelength or color channel, limiting it to monochromatic image generation. This limitation in channel capacity significantly restricts the applicability of DOE in optical applications. In this study, we propose a design strategy for full-color, single-layer hologram based on a variable-scale diffraction model. By imposing strict constraints in Fourier domain and reducing depth of focus (DOF), we achieve the simultaneous encryption and storage of red, green, and blue channel information within a single achromatic hologram. This strategy facilitates color separation in large-depth 3D holography and enables achromatic full-color image displays. We demonstrated full-color holographic video playback at a full refresh rate of 60 Hz, achieving a temporal resolution three times greater than that of existing methods. Furthermore, we successfully fabricated achromatic, twin-image-free, full-color binary pure-phase DOEs at low cost. This achromatic strategy addresses the demands across various fields in optics, including high-refresh-rate full-color displays, high-density optical information storage, advanced optical security, high-reusability holographic metasurface optical element, and high-performance achromatic metalenses.

physics.optics

Neural refractive index field: Unlocking the Potential of Background-oriented Schlieren Tomography in Volumetric Flow Visualization

Background-oriented Schlieren tomography (BOST) is a prevalent method for visualizing intricate turbulent flows, valued for its ease of implementation and capacity to capture three-dimensional distributions of a multitude of flow parameters. However, the voxel-based meshing scheme leads to significant challenges, such as inadequate spatial resolution, substantial discretization errors, poor noise immunity, and excessive computational costs. This work presents an innovative reconstruction approach termed neural refractive index field (NeRIF) which implicitly represents the flow field with a neural network, which is trained with tailored strategies. Both numerical simulations and experimental demonstrations on turbulent Bunsen flames suggest that our approach can significantly improve the reconstruction accuracy and spatial resolution while concurrently reducing computational expenses. Although showcased in the context of background-oriented schlieren tomography here, the key idea embedded in the NeRIF can be readily adapted to various other tomographic modalities including tomographic absorption spectroscopy and tomographic particle imaging velocimetry, broadening its potential impact across different domains of flow visualization and analysis.

physics.flu-dyn

DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting

Dynamic scene reconstruction has garnered significant attention in recent years due to its capabilities in high-quality and real-time rendering. Among various methodologies, constructing a 4D spatial-temporal representation, such as 4D-GS, has gained popularity for its high-quality rendered images. However, these methods often produce suboptimal surfaces, as the discrete 3D Gaussian point clouds fail to align with the object's surface precisely. To address this problem, we propose DynaSurfGS to achieve both photorealistic rendering and high-fidelity surface reconstruction of dynamic scenarios. Specifically, the DynaSurfGS framework first incorporates Gaussian features from 4D neural voxels with the planar-based Gaussian Splatting to facilitate precise surface reconstruction. It leverages normal regularization to enforce the smoothness of the surface of dynamic objects. It also incorporates the as-rigid-as-possible (ARAP) constraint to maintain the approximate rigidity of local neighborhoods of 3D Gaussians between timesteps and ensure that adjacent 3D Gaussians remain closely aligned throughout. Extensive experiments demonstrate that DynaSurfGS surpasses state-of-the-art methods in both high-fidelity surface reconstruction and photorealistic rendering.

cs.CV

Broadband miniaturized spectrometers with a van der Waals tunnel diode

Miniaturized spectrometers are of immense interest for various on-chip and implantable photonic and optoelectronic applications. State-of-the-art conventional spectrometer designs rely heavily on bulky dispersive components (such as gratings, photodetector arrays, and interferometric optics) to capture different input spectral components that increase their integration complexity. Here, we report a high-performance broadband spectrometer based on a simple and compact van der Waals heterostructure diode, leveraging a careful selection of active van der Waals materials -- molybdenum disulfide and black phosphorus, their electrically tunable photoresponse, and advanced computational algorithms for spectral reconstruction. We achieve remarkably high peak wavelength accuracy of ~2 nanometers, and broad operation bandwidth spanning from ~500 to 1600 nanometers in a device with a ~30x20 μm2 footprint. This diode-based spectrometer scheme with broadband operation offers an attractive pathway for various applications, such as sensing, surveillance and spectral imaging.

physics.optics

Real Differences between OT and CRDT in Building Co-Editing Systems and Real World Applications

OT (Operational Transformation) was invented for supporting real-time co-editors in the late 1980s and has evolved to become a core technique used in today's working co-editors and adopted in major industrial products. CRDT (Commutative Replicated Data Type) for co-editors was first proposed around 2006, under the name of WOOT (WithOut Operational Transformation). Follow-up CRDT variations are commonly labeled as "post-OT" techniques and have made broad claims of superiority over OT solutions, in terms of correctness, time and space complexity, simplicity, etc. Over one decade later, however, OT remains the choice for building the vast majority of co-editors, whereas CRDT is rarely found in working co-editors. Why? To seek truth from facts, we set out to conduct a comprehensive and critical review of representative OT and CRDT solutions and working co-editors based on them. From this work, we have made important discoveries about OT and CRDT, and revealed facts and evidences that refute CRDT claims over OT on all accounts. We present our discoveries in three related and complementary articles. In prior two articles, we have revealed the similarities of OT and CRDT in following the same general transformation approach in co-editors, and their real differences in correctness and complexity. In this article, we examine the role of building working co-editors in shaping OT and CRDT research and solutions, and consequential differences in the choice between OT and CRDT in real world co-editors and industry products. In particular, we review the evolution of co-editors from research vehicles to real world applications, and discuss representative OT-based co-editors and alternative approaches in industry products and open source projects. Moreover, we evaluate CRDT-based co-editors in relation to published CRDT solutions, and clarify some myths surrounding "peer-to-peer" co-editing.

cs.DC

Real Differences between OT and CRDT under a General Transformation Framework for Consistency Maintenance in Co-Editors

OT (Operational Transformation) was invented for supporting real-time co-editors in the late 1980s and has evolved to become a core technique used in today's working co-editors and adopted in major industrial products. CRDT (Commutative Replicated Data Type) for co-editors was first proposed around 2006, under the name of WOOT (WithOut Operational Transformation). Follow-up CRDT variations are commonly labeled as "post-OT" techniques capable of making concurrent operations natively commutative in co-editors. On top of that, CRDT solutions have made broad claims of superiority over OT solutions, and routinely portrayed OT as an incorrect, complex and inefficient technique. Over one decade later, however, OT remains the choice for building the vast majority of co-editors, whereas CRDT is rarely found in working co-editors. Contradictions between the reality and CRDT's purported advantages have been the source of much confusion and debate in co-editing communities. Have the vast majority of co-editors been unfortunate in choosing the faulty and inferior OT, or those CRDT claims are false? What are the real differences between OT and CRDT for co-editors? What are the key factors and underlying reasons behind the choices between OT and CRDT in the real world? To seek truth from facts, we set out to conduct a comprehensive and critical review on representative OT and CRDT solutions and working co-editors based on them. From this work, we have made important discoveries about OT and CRDT, and revealed facts and evidences that refute CRDT claims over OT on all accounts. We report our discoveries in a series of three articles and the current article is the first one in this series. We hope the discoveries from this work help clear up common misconceptions and confusions surrounding OT and CRDT, and accelerate progress in co-editing technology for real world applications.

cs.DC