SearcharxivSearch

arXiv subjects

Zhengyang Sun

Publications and source records attributed to Zhengyang Sun.

12 recordsLinked to original sources

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test-time future imagination. We present SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without explicit future-frame generation at inference. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.

cs.CV

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA , a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsistencies by selecting discriminative self- and cross-attention heads to derive a countable latent layout. It then refines this layout conservatively and modulates cross-attention to guide regeneration. On the introduced CountBench, NUMINA improves counting accuracy by up to 7.4% on Wan2.1-1.3B, and by 4.9% and 5.5% on 5B and 14B models, respectively. Furthermore, CLIP alignment is improved while maintaining temporal consistency. These results demonstrate that structural guidance complements seed search and prompt enhancement, offering a practical path toward count-accurate text-to-video diffusion. The code is available at https://github.com/H-EmbodVis/NUMINA.

cs.CV

Mitigating Hallucination on Hallucination in RAG via Ensemble Voting

Retrieval-Augmented Generation (RAG) aims to reduce hallucinations in Large Language Models (LLMs) by integrating external knowledge. However, RAG introduces a critical challenge: hallucination on hallucination," where flawed retrieval results mislead the generation model, leading to compounded hallucinations. To address this issue, we propose VOTE-RAG, a novel, training-free framework with a two-stage structure and efficient, parallelizable voting mechanisms. VOTE-RAG includes: (1) Retrieval Voting, where multiple agents generate diverse queries in parallel and aggregate all retrieved documents; (2) Response Voting, where multiple agents independently generate answers based on the aggregated documents, with the final output determined by majority vote. We conduct comparative experiments on six benchmark datasets. Our results show that VOTE-RAG achieves performance comparable to or surpassing more complex frameworks. Additionally, VOTE-RAG features a simpler architecture, is fully parallelizable, and avoids the problem drift" risk. Our work demonstrates that simple, reliable ensemble voting is a superior and more efficient method for mitigating RAG hallucinations.

cs.CL

A camera system for real-time optical calibration of water-based neutrino telescopes

Calibrating the optical properties within the detection medium of a neutrino telescope is crucial for determining its angular resolution and energy scale. For the next generation of neutrino telescopes planned to be constructed in deep water, such as the TRopIcal DEep-sea Neutrino Telescope (TRIDENT), there are additional challenges due to the dynamic nature and potential non-uniformity of the water medium. This necessitates a real-time optical calibration system distributed throughout the large detector array. This study introduces a custom-designed CMOS camera system equipped with rapid image processing algorithms, providing a real-time optical calibration method for TRIDENT and other similar projects worldwide. In September 2021, the TRIDENT Pathfinder experiment (TRIDENT Explorer, T-REX for short) successfully deployed this camera system in the West Pacific Ocean at a depth of 3420 meters. Within 30 minutes, about 3000 images of the T-REX light source were captured, allowing for the in-situ measurement of seawater attenuation and absorption lengths under three wavelengths. This deep-sea experiment for the first time showcased a technical demonstration of a functioning camera calibration system in a dynamic neutrino telescope site, solidifying a substantial part of the calibration strategies for the future TRIDENT project.

physics.ins-det

The PMT System of the TRIDENT Pathfinder Experiment

Next generation neutrino telescopes are highly anticipated to boost the development of neutrino astronomy. A multi-cubic-kilometer neutrino telescope, TRopIcal DEep-sea Neutrino Telescope (TRIDENT), was proposed to be built in the South China Sea. The detector aims to achieve ~ 0.1 degree angular resolution for track-like events at energy above 100 TeV by using hybrid digital optical modules, opening new opportunities for neutrino astronomy. In order to measure the water optical properties and marine environment of the proposed TRIDENT site, a pathfinder experiment was conducted, in which a 100-meter-long string consisting of three optical modules was deployed at a depth of 3420 m to perform in-situ measurements. The central module emits light by housing LEDs, whereas the other two modules detect light with two independent and complementary systems: the PMT and the camera systems. By counting the number of detected photons and analyzing the photon arrival time distribution, the PMT system can measure the absorption and scattering lengths of sea water, which serve as the basic inputs for designing the neutrino telescope. In this paper, we present the design concept, calibration and performance of the PMT system in the pathfinder experiment.

hep-ex

The trigger system for the CSR external-target experiment

A trigger system has been designed and implemented for the HIRFL-CSR external target experiment (CEE), the spectrometer for studying nuclear matter properties with heavy ion collisions in the GeV energy region. The system adopts master-slave structure and serial data transmission mode using optical fiber to deal with different types of detectors and long-distance signal transmission. The trigger logic can be accessed based on command register and controlled by a remote computer. The overall field programmable gate array (FPGA) logic can be flexibly reconfigured online to match the physical requirements of the experiment. The trigger system has been tested in beam experiment. It is demonstrated that the trigger system functions correctly and meets the physical requirements of CEE.

physics.ins-det

A fast tunable driver of light source for the TRIDENT Pathfinder experiment

TRIDENT (The tRopIcal DEep-sea Neutrino Telescope) is a proposed next-generation neutrino telescope to be constructed in the South China Sea. In September 2021, the TRIDENT Pathfinder experiment (TRIDENT EXplorer, T-REX for short) was conducted to evaluate the in-situ optical properties of seawater. The T-REX experiment deployed three digital optical modules at a depth of 3420 meters, including a light emitter module (LEM) and two light receiver modules (LRMs) equipped with photomultiplier tubes (PMTs) and cameras to detect light signals. The LEM emits light in pulsing and steady modes. It features a fast tunable driver to activate light-emitting diodes (LEDs) that emit nanosecond-width light pulses with tunable intensity. The PMTs in the LRM receive single photo-electron (SPE) signals with an average photon number of approximately 0.3 per 1-microsecond time window, which is used to measure the arrival time distribution of the SPE signals. The fast tunable driver can be remotely controlled in real-time by the data acquisition system onboard the research vessel, allowing for convenient adjustments to the driver's parameters and facilitating the acquisition of high-quality experimental data. This paper describes the requirements, design scheme, and test results of the fast tunable driver, highlighting its successful implementation in the T-REX experiment and its potential for future deep-sea experiments.

physics.ins-det

The Light Source of the TRIDENT Pathfinder Experiment

In September 2021, a site scouting mission known as the TRIDENT pathfinder experiment (TRIDENT EXplorer, T-REX for short) was conducted in the South China Sea with the goal of envisaging a next-generation multi-cubic-kilometer neutrino telescope. One of the main tasks is to measure the in-situ optical properties of seawater at depths between $2800~\mathrm{m}$ and $3500~\mathrm{m}$, where the neutrino telescope will be instrumented. To achieve this, we have developed a light emitter module equipped with a clock synchronization system to serve as the light source, which could be operated in pulsing and steady modes. Two light receiver modules housing both photomultiplier tubes (PMTs) and cameras are employed to detect the photons emitted by the light source. This paper presents the instrumentation of the light source in T-REX, including its design, calibration, and performance.

astro-ph.IM

INT: Towards Infinite-frames 3D Detection with An Efficient Framework

It is natural to construct a multi-frame instead of a single-frame 3D detector for a continuous-time stream. Although increasing the number of frames might improve performance, previous multi-frame studies only used very limited frames to build their systems due to the dramatically increased computational and memory cost. To address these issues, we propose a novel on-stream training and prediction framework that, in theory, can employ an infinite number of frames while keeping the same amount of computation as a single-frame detector. This infinite framework (INT), which can be used with most existing detectors, is utilized, for example, on the popular CenterPoint, with significant latency reductions and performance improvements. We've also conducted extensive experiments on two large-scale datasets, nuScenes and Waymo Open Dataset, to demonstrate the scheme's effectiveness and efficiency. By employing INT on CenterPoint, we can get around 7% (Waymo) and 15% (nuScenes) performance boost with only 2~4ms latency overhead, and currently SOTA on the Waymo 3D Detection leaderboard.

cs.CV

Simulation study on the optical processes at deep-sea neutrino telescope sites

The performance of a large-scale water Cherenkov neutrino telescope relies heavily on the transparency of the surrounding water, quantified by its level of light absorption and scattering. A pathfinder experiment was carried out to measure the optical properties of deep seawater in South China Sea with light-emitting diodes (LEDs) as light sources, photon multiplier tubes (PMTs) and cameras as photon sensors. Here, we present an optical simulation program employing the Geant4 toolkit to understand the absorption and scattering processes in the deep seawater, which helps to extract the underlying optical properties from the experimental data. The simulation results are compared with the experimental data and show good agreements. We also verify the analysis methods that utilize various observables of the PMTs and the cameras with this simulation program, which can be easily adapted by other neutrino telescope pathfinder experiments and future large-scale detectors.

astro-ph.IM

Target observation of complex networks

How to observe the state of a network from a limited number of measurements has become an important issue in complex networks, engineering, communication, epidemiology, etc. Under some scenarios, it is neither unfeasible nor unnecessary to observe the entire network. Therefore, we investigate the target observation of a network in this paper. We propose a target minimal dominating set problem corresponding to target observation, which is a natural generalization of classical minimal dominating set problem. Three algorithms are proposed to approximate the minimum set of occupied nodes sufficient for target observation. Extensive numerical results on computer-generated random networks and real-world networks demonstrate that the proposed algorithms offer superior performance in identification of a target minimal dominating set.

physics.soc-ph

An integrative sparse boosting analysis of cancer genomic commonality and difference

In cancer research, high-throughput profiling has been extensively conducted. In recent studies, the integrative analysis of data on multiple cancer patient groups/subgroups has been conducted. Such analysis has the potential to reveal the genomic commonality as well as difference across groups/subgroups. However, in the existing literature, methods with a special attention to the genomic commonality and difference are very limited. In this study, a novel estimation and marker selection method based on the sparse boosting technique is developed to address the commonality/difference problem. In terms of technical innovation, a new penalty and computation of increments are introduced. The proposed method can also effectively accommodate the grouping structure of covariates. Simulation shows that it can outperform direct competitors under a wide spectrum of settings. The analysis of two TCGA (The Cancer Genome Atlas) datasets is conducted, showing that the proposed analysis can identify markers with important biological implications and have satisfactory prediction and stability.

stat.ME