SearcharxivSearch

arXiv subjects

Junyu Zhu

Publications and source records attributed to Junyu Zhu.

10 recordsLinked to original sources

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

Dynamic 3D scene reconstruction has made significant progress with multi-camera systems, often relying on temporally aligned observations across views. However, in real-world scenarios, temporal asynchrony among capturing devices remains a common limitation, leading to severe motion blur and geometric artifacts. Existing asynchronous reconstruction methods typically estimate temporal offsets through photometric supervision, but appearance matching provides weak temporal cues under large offsets and complex motions. We attribute this limitation to two critical issues: texture-induced collapse, where low-textured regions provide nearly vanishing alignment signals, and deformation-induced entanglement, where temporal errors are absorbed into distorted geometry or motion rather than being explicitly corrected. To address these issues, we propose ASTRA (Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment), a framework that introduces 2D motion trajectories as explicit, texture-robust supervision for asynchronous dynamic reconstruction. Instead of synchronizing cameras solely through rendered color residuals, ASTRA jointly optimizes temporal offsets and dynamic 3D representations by aligning the projected motion of reconstructed 3D points with observed 2D trajectories, while using dynamic and certainty masking to suppress unreliable trajectory constraints. Extensive experiments on different dynamic Gaussian Splatting backbones show that ASTRA preserves high-frequency spatial details and sustains strong robustness even under severe asynchrony with up to 25-frame offsets, achieving approximately 1.4 dB PSNR improvement, reducing temporal-offset MAE by 54.0%, and nearly quadrupling the synchronization success rate.

cs.CV

High-frequency gravitational waves from axion inflation in the weak-backreaction regime

Axion inflation, characterized by a Chern-Simons interaction between the inflaton and a gauge field, provides a powerful mechanism for generating primordial gravitational waves (GWs) through tachyonic enhancement of the gauge field. While recent literature has predominantly focused on the Strong Backreaction (SB) regime to maximize GW signals for future interferometers, this regime suffers from computational complexities as well as the risk of overproducing scalar perturbations. In this work, we investigate gauge field amplification and GW production strictly within the theoretically safer Weak Backreaction (WB) regime, with a particular focus on the largely unexplored non-instantaneous reheating phase. Because the tachyonic enhancement during slow-roll typically increases as inflation approaches its end, it is crucial to investigate how the production of GWs behaves at the very end of inflation and thereafter. By continuously tracking the evolution from slow-roll through reheating to radiation domination, we present a complete picture of inflationary and post-inflationary GW production in this framework. A particularly interesting feature of the post-inflationary phase is that the oscillatory behavior of the inflaton during reheating leads to frequent sign-flips of the instability parameter $\xi$, exciting both helical modes of the gauge field. Our analysis reveals that axion inflation can naturally generate one of the strongest known primordial GW signals at high-frequency bands. The yield is relevant for future precision measurements of the effective number of neutrino species, $N_{{\rm eff}}$, and also strongly motivates the development of novel high-frequency GW detectors.

hep-ph

Controllable Sim Agents with Behavior Latents

Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test autonomous systems without real-world risk. We introduce Controllable Neural Variational Agents (CNeVA), a controllable simulated-agent framework that learns to infer a per-agent Gaussian behavior latent from per-channel discounted returns via a closed-form conjugate variational update, conditioning a rectified-flow trajectory generator trained on a mixed channel-mask curriculum for classifier-free guidance. To tackle scarcity in reward signals, we propose soft eligibility gates that replace hard binary thresholds with smooth exponential decay, preserving the gradient signal for near-threshold agents. On the Waymo Open Motion Dataset, CNeVA attains competitive realism on the benchmark while exposing per-channel controllability that the higher-ranked imitation models lack. Speed- and acceleration-based steering produces monotone responses without stall-induced reward hacking. Safety controllability is monotone and substantial with the introduction of soft eligibility. We manage to achieve steerable map compliance under a context-residual return measure. Furthermore, our experiment demonstrates that steering metrics must be read alongside physical-plausibility guardrails to avoid reward-hacking confounds.

cs.RO

Full-Spectrum Analysis of Gravitational Wave Production from Inflation to Reheating

In this work, we systematically study gravitational wave (GW) production during both the inflationary and post-inflationary epochs. While inflationary GWs can be readily derived from tensor perturbations during inflation, post-inflationary GWs arise from a variety of processes during reheating and require detailed treatment for quantitative analysis. We consider four distinct production channels: $(i)$ pure inflaton annihilation, $(ii)$ graviton bremsstrahlung from inflaton decay, $(iii)$ radiation-catalyzed inflaton-graviton conversion, and $(iv)$ scattering among fully thermalized radiation particles. For each channel, we solve the corresponding Boltzmann equation to obtain the GW spectrum and derive a simple yet accurate analytical expression for it. By employing a consistent treatment of all production channels, our analysis yields for the first time the full spectrum of GWs produced during the inflationary and post-inflationary epochs. We find that, while inflationary GWs dominate at low frequencies, post-inflationary processes generally produce high-frequency GWs with considerably high energy densities that may significantly exceed that of inflationary GWs.

hep-ph

Self-supervised Event-based Monocular Depth Estimation using Cross-modal Consistency

An event camera is a novel vision sensor that can capture per-pixel brightness changes and output a stream of asynchronous ``events''. It has advantages over conventional cameras in those scenes with high-speed motions and challenging lighting conditions because of the high temporal resolution, high dynamic range, low bandwidth, low power consumption, and no motion blur. Therefore, several supervised monocular depth estimation from events is proposed to address scenes difficult for conventional cameras. However, depth annotation is costly and time-consuming. In this paper, to lower the annotation cost, we propose a self-supervised event-based monocular depth estimation framework named EMoDepth. EMoDepth constrains the training process using the cross-modal consistency from intensity frames that are aligned with events in the pixel coordinate. Moreover, in inference, only events are used for monocular depth prediction. Additionally, we design a multi-scale skip-connection architecture to effectively fuse features for depth estimation while maintaining high inference speed. Experiments on MVSEC and DSEC datasets demonstrate that our contributions are effective and that the accuracy can outperform existing supervised event-based and unsupervised frame-based methods.

cs.CV

Camera-based 3D Semantic Scene Completion with Sparse Guidance Network

Semantic scene completion (SSC) aims to predict the semantic occupancy of each voxel in the entire 3D scene from limited observations, which is an emerging and critical task for autonomous driving. Recently, many studies have turned to camera-based SSC solutions due to the richer visual cues and cost-effectiveness of cameras. However, existing methods usually rely on sophisticated and heavy 3D models to process the lifted 3D features directly, which are not discriminative enough for clear segmentation boundaries. In this paper, we adopt the dense-sparse-dense design and propose a one-stage camera-based SSC framework, termed SGN, to propagate semantics from the semantic-aware seed voxels to the whole scene based on spatial geometry cues. Firstly, to exploit depth-aware context and dynamically select sparse seed voxels, we redesign the sparse voxel proposal network to process points generated by depth prediction directly with the coarse-to-fine paradigm. Furthermore, by designing hybrid guidance (sparse semantic and geometry guidance) and effective voxel aggregation for spatial geometry cues, we enhance the feature separation between different categories and expedite the convergence of semantic propagation. Finally, we devise the multi-scale semantic propagation module for flexible receptive fields while reducing the computation resources. Extensive experimental results on the SemanticKITTI and SSCBench-KITTI-360 datasets demonstrate the superiority of our SGN over existing state-of-the-art methods. And even our lightweight version SGN-L achieves notable scores of 14.80\% mIoU and 45.45\% IoU on SeamnticKITTI validation with only 12.5 M parameters and 7.16 G training memory. Code is available at https://github.com/Jieqianyu/SGN.

cs.CV

The $\nu_{R}$-philic scalar dark matter

Right-handed neutrinos ($\nu_{R}$) offer an intriguing portal to new physics in hidden sectors where dark matter (DM) may reside. In this work, we delve into the simplest hidden sector involving only a real scalar exclusively coupled to $\nu_{R}$, referred to as the $\nu_{R}$-philic scalar. We investigate the viability of the $\nu_{R}$-philic scalar to serve as a DM candidate, under the constraint that the coupling of $\nu_{R}$ to the standard model is determined by the seesaw relation and is responsible for the observed DM abundance. By analyzing the DM decay channels and solving Boltzmann equations, we identify the viable parameter space. In particular, our study reveals a lower bound ($2.6\times10^{5}$ GeV) on the mass of $\nu_{R}$ for the $\nu_{R}$-philic scalar to be DM. The DM mass may vary from sub-MeV to sub-GeV. Within the viable parameter space, monochromatic neutrino lines from DM decay can be an important signal for DM indirect detection.

hep-ph

Semi-Supervised Learning for Visual Bird's Eye View Semantic Segmentation

Visual bird's eye view (BEV) semantic segmentation helps autonomous vehicles understand the surrounding environment only from images, including static elements (e.g., roads) and dynamic elements (e.g., vehicles, pedestrians). However, the high cost of annotation procedures of full-supervised methods limits the capability of the visual BEV semantic segmentation, which usually needs HD maps, 3D object bounding boxes, and camera extrinsic matrixes. In this paper, we present a novel semi-supervised framework for visual BEV semantic segmentation to boost performance by exploiting unlabeled images during the training. A consistency loss that makes full use of unlabeled data is then proposed to constrain the model on not only semantic prediction but also the BEV feature. Furthermore, we propose a novel and effective data augmentation method named conjoint rotation which reasonably augments the dataset while maintaining the geometric relationship between the front-view images and the BEV semantic segmentation. Extensive experiments on the nuScenes and Argoverse datasets show that our semi-supervised framework can effectively improve prediction accuracy. To the best of our knowledge, this is the first work that explores improving visual BEV semantic segmentation performance using unlabeled data. The code is available at https://github.com/Junyu-Z/Semi-BEVseg

cs.CV

FG-Depth: Flow-Guided Unsupervised Monocular Depth Estimation

The great potential of unsupervised monocular depth estimation has been demonstrated by many works due to low annotation cost and impressive accuracy comparable to supervised methods. To further improve the performance, recent works mainly focus on designing more complex network structures and exploiting extra supervised information, e.g., semantic segmentation. These methods optimize the models by exploiting the reconstructed relationship between the target and reference images in varying degrees. However, previous methods prove that this image reconstruction optimization is prone to get trapped in local minima. In this paper, our core idea is to guide the optimization with prior knowledge from pretrained Flow-Net. And we show that the bottleneck of unsupervised monocular depth estimation can be broken with our simple but effective framework named FG-Depth. In particular, we propose (i) a flow distillation loss to replace the typical photometric loss that limits the capacity of the model and (ii) a prior flow based mask to remove invalid pixels that bring the noise in training loss. Extensive experiments demonstrate the effectiveness of each component, and our approach achieves state-of-the-art results on both KITTI and NYU-Depth-v2 datasets.

cs.CV

Giant-diamagnetic and magnetization-step effects in HgMnTe monocrystal

In Hg$_{1-x}$Mn$_x$Te (x$\geq$0.16) monocrystal, the giant-diamagnetic (GDM) and magnetization-step phenomena have been observed in spin glass (SG) regime. The susceptibility of GDM is about 100-1000 times than that of classic diamagnetic. It can be interpreted that: due to the long-range antiferromagnetic (AF) exchange interactions and the non-uniform random distribution of Mn$^{2+}$ ions in Hg$_{1-x}$Mn$_x$Te, a quasi-static spin wave forms and produces the GDM phenomenon below the critical temperature and magnetic field. Meanwhile, this theory is proved by Monte Carlo simulations in a two-dimensional AF cluster based on XY model. Hence, it is possible to emerge long-range magnetic order structure in SG state.

cond-mat.mtrl-sci