SearcharxivSearch

arXiv subjects

Yuhao Yan

Publications and source records attributed to Yuhao Yan.

11 recordsLinked to original sources

Frequency-Domain Mixing Data Augmentation for Malicious Traffic Detection

The strong dynamics of network traffic often force malicious traffic detection models to handle out-of-distribution data. Typically, deep learning-based malicious traffic detection models require a large amount of high-quality training data. However, owing to challenges such as high labeling difficulty and resource consumption, existing datasets often suffer from insufficient diversity and fail to capture evolving traffic patterns, leading to poor out-of-distribution generalization ability of the trained models. Data augmentation has been widely adopted to improve data diversity and model generalization. Recently, frequency-domain mixing augmentation has shown promising performance because it effectively perturbs data while preserving key structural information. This approach shows potential for enhancing malicious traffic detection models. However, existing studies lack theoretical interpretation of the mixing mechanism, and do not adapt to the characteristics of network traffic. In this paper, we first conduct a theoretical analysis of the current frequency-domain mixing method, revealing its underlying principles and limitations. We further propose an improved frequency-domain mixing-based data augmentation method for network traffic data, which enhances the diversity of sequence features in network traffic and improves the out-of-distribution generalization of malicious traffic detection models. Extensive experiments on multiple artificial and real-world datasets demonstrate that our method substantially improves detection performance across diverse network environments and outperforms other data augmentation approaches.

cs.CR

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.

cs.LG

MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation

Medical vision-language models (VLMs) are increasingly expected to support clinical workflows through diagnostic text and relevant medical images. However, current medical visual benchmarks have three recurring limitations: query-image misalignment from queries weakly grounded in specific image instances, closed-ended formats that narrow answer space and encourage shortcut-based prediction, and text-centric output paradigms that limit evaluation of image-generation and image-editing capabilities. We introduce MedGEN-Bench, a benchmark for open-ended multimodal medical generation. The evaluation snapshot reported in this manuscript comprises 6,422 image-text pairs reviewed by clinical experts and models, spanning 6 canonical imaging modalities, 15 clinical tasks, and 27 named subtasks. It includes 1,100 Visual Question Answering (VQA) pairs, 3,872 Image Editing pairs, and 1,450 Contextual Multimodal Generation pairs. MedGEN-Bench centers on contextual entanglement: dependence of an instruction's intended output on the particular image instance rather than on task wording alone. The benchmark operationalizes this concept through image-grounded instructions and extends evaluation to open-ended multimodal outputs. Its tiered evaluation protocol combines reproducible reference-based fidelity and similarity measures with a structured, checklist-guided assessment by a medical VLM judge. We evaluate 10 compositional frameworks, 2 dedicated image-editing models, 3 unified models, and 5 VLMs. The results show image-output tasks remain unsaturated. Contextual augmentation increases mean image-instruction similarity from 0.273 to 0.372, while a 1,000-case medical-expert audit shows moderate agreement between judge scores and clinician ratings. Source code and dataset are available at https://yangjj007.github.io/medgen.

cs.CV

Characterizing the Reliability of a Novel Upright CT for Proton Therapy

Purpose: To evaluate reliability of upright CT for proton dose calculation and feasibility of a simplified phantom configuration for accelerated routine QA. Methods: A calibration phantom was scanned on an upright CT following consensus guidelines for 14 sessions/7 months. CT number repeatability was assessed by standard deviation (SD). Stopping power ratio (SPR) look-up table was derived. Phantom size dependency was assessed. The simplified phantom configuration was scanned for 15 sessions/8 months. Repeatability was assessed. CT numbers and SPR were compared with consensus configuration. Both configurations were scanned on a recumbent CT to validate the findings. An anthropomorphic phantom was scanned on upright and recumbent CT. Targets were drawn mimicking spine and prostate tumor. Proton plans were developed using pencil beam scanning techniques and robust optimization. Equivalence of dose calculation were assessed via controlled comparisons. Results: Simplified configuration measured all CT numbers in 1 scan vs 5 for consensus guidelines. Upright CT demonstrated excellent longitudinal stability (inter- and intrasession SD <4.9 HU and 1.6 HU, respectively). Size dependency was identified with significant (p<.05) differences in CT numbers, propagated to $\Delta$SPR <5.3%. Significant (p<.05) differences were found comparing upright CT numbers measured by 2 configurations ($\Delta$SPR<2.6%). Recumbent CT showed smaller $\Delta$SPR (<0.7%). Both dosimetric comparison showed local differences (<8% of prescription dose) while clinical equivalence was found with target coverage differences <0.2% and gamma pass rates=100% at 3 mm/3% for all controlled comparison of different CT machines and phantom configurations. Conclusions: The upright CT demonstrated reliability to support adaptive proton therapy. The simplified configuration shows feasibility for rapid QA.

physics.med-ph

Technical assessment of a novel vertical CT system for upright radiotherapy simulation and treatment planning

Purpose: To characterize image quality, imaging dose, and dose calculation accuracy for an upright CT scanner with a six-degree-of-freedom patient positioning system. Methods: Imaging dose (CTDIvol) was measured at 120 kVp and 200 mAs. Image quality was evaluated using an ACR-464 phantom. Mean CT number accuracy was assessed within inserts of known material and uniformity as the difference in values at the center and periphery of uniform phantoms. High-contrast resolution was assessed by visible line pairs and modulation transfer function (MTF). Low-contrast performance was quantified by contrast-to-noise-ratio (CNR). Spatial integrity was evaluated between fiducials 100 mm apart. Hounsfield unit to mass density and stopping-power-ratio calibrations were performed. Proton and photon treatment plans were optimized on upright CT scans of a thorax phantom in heterogenous and homogeneous regions. Dose was forward computed on a registered recumbent CT scan and agreement evaluated using 3D gamma analysis. Results: CT imaging dose (CTDIvol) was 23.5 mGy for the 16 cm head phantom and 10.1 mGy for the 32 cm body phantom. Mean CT numbers (HU) were within the expected range for water (1.7) and acrylic (120.8). CT numbers were slightly (5-27 HU) out-of-range for air (-950.4), polyethylene (-78.8), and bone (823.0). Image uniformity was 20.2 HU and 35.0 HU for 20 cm and 48 cm diameter phantoms, respectively. Eight high-contrast line pairs were visualized. The MTF equaled 4.4 cm-1 at 50% and 7.1 cm-1 at 10%. The median CNR was 0.93, below the 1.0 tolerance. Spatial integrity was 0.36 mm. Gamma pass rates were 99.8% for photon and 90.6% for proton plans with 1%/1mm criteria, and greater than or equal to 98.0% for all plans with 3%/2mm criteria. Conclusion: Upright CT image quality and dose calculation accuracy are acceptable for photon and proton radiotherapy.

physics.med-ph

SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop

We present an overview of the Spatio-temporal Instance Segmentation (SIS) challenge held in conjunction with the CVPR 2025 Event-based Vision Workshop. The task is to predict accurate pixel-level segmentation masks of defined object classes from spatio-temporally aligned event camera and grayscale camera data. We provide an overview of the task, dataset, challenge details and results. Furthermore, we describe the methods used by the top-5 ranking teams in the challenge. More resources and code of the participants' methods are available here: https://github.com/tub-rip/MouseSIS/blob/main/docs/challenge_results.md

cs.CV

Evaluation of a Novel Quantitative Multiparametric MR Sequence for Radiation Therapy Treatment Response Assessment

Purpose: To evaluate a Deep-Learning-enhanced MUlti-PArametric MR sequence (DL-MUPA) for treatment response assessment for brain metastases patients undergoing stereotactic radiosurgery (SRS) and head-and-neck (HnN) cancer patients undergoing conventionally fractionation adaptive radiation therapy. Methods: DL-MUPA derives quantitative T1 and T2 maps from a single 4-6-minute scan denoised via DL method using dictionary fitting. Phantom benchmarking was performed on a NIST-ISMRM phantom. Longitudinal patient data were acquired on a 1.5T MR-simulator, including pre-treatment (PreTx) and every 3 months after SRS (PostTx) in brain, and PreTx, mid-treatment and 3 months PostTx in HnN. Changes of mean T1 and T2 values were calculated within gross tumor volumes (GTVs), residual disease (RD, HnN), parotids, and submandibular glands (HnN) for treatment response assessment. Uninvolved normal tissues (normal appearing white matter in brain, masseter in HnN) were evaluated to as control. Results: Phantom benchmarking showed excellent inter-session repeatability (coefficient of variance <1% for T1, <7% for T2). Uninvolved normal tissue suggested acceptable in-vivo repeatability (brain |$\Delta$|<6%, HnN |$\Delta$T1|<7%, |$\Delta$T2|<18% (4ms)). Remarkable changes were noted in resolved brain metastasis ($\Delta$T1=14%) and necrotic settings ($\Delta$T1=18-40%, $\Delta$T2=9-41%). In HnN, two primary tumors showed T2 increase (PostTx GTV $\Delta$T2>13%, RD $\Delta$T2>18%). A nodal disease resolved PostTx (GTV $\Delta$T1=-40%, $\Delta$T2=-33%, RD $\Delta$T1=-29%, $\Delta$T2=-35%). Enhancement was found in involved parotids (PostTx $\Delta$T1>12%, $\Delta$T2>13%) and submandibular glands (PostTx $\Delta$T1>15%, $\Delta$T2>35%) while the uninvolved organs remained stable. Conclusions: DL-MUPA shows promise for treatment response assessment and identifying potential endpoints for functional sparing.

physics.med-ph

Revisit the intrinsic features of flip-flopping flow behind side-by-side circular cylinders

As one of the most intriguing wake patterns of two side-by-side circular cylinders at an intermediate gap spacing, the flip-flopping (FF) flow has attracted great attention of fundamental research interest. This FF flow is featured by the intermittently and randomly switching gap flow with correspondingly changing forces of the two cylinders. In this paper, we first present a partition map of the wake patterns behind two side-by-side circular cylinders and briefly introduce intrinsic features of each flow pattern. We focus on the FF flow aiming to explain: (i) the origin of the FF flow between laminar and turbulent regimes, (ii) their connections in different flow regimes, and (iii) mechanisms of the significantly varying flip-over time scale of the FF flows. In the laminar regime, we further divide the FF flow into the sub-classed I (FF1) and II (FF2), based on their different origins from the in-phase and anti-phase synchronized vortex shedding instabilities, respectively. By exploring the vortex interactions, we show that the FF flow in the turbulent regime has the same origin and similar vortex dynamics as the FF2 wake in the laminar regime, despite some minor disparities. Thus, a connection is established between the FF2 pattern in the laminar flow and the FF pattern in the turbulent flow. For the FF flow in the laminar regime (Re < 150-200), the mildly decreasing switching time, is several vortex shedding periods. However, for the FF flow in the weak turbulence regime (150-200 < Re < 1000-1700), the switching time scale increases significantly with Re owing to the increased vortex formation length. The FF in the strong turbulence regime (Re > 1000-1700) has a switching time scale of several orders of magnitude longer than the vortex shedding period, where the switching scale decreases gradually with Re due to the stronger Kelvin-Helmholtz vortices.

physics.flu-dyn

Deep fused flow and topology features for botnet detection basing on pretrained GCN

Nowadays, botnets have become one of the major threats to cyber security. The characteristics of botnets are mainly reflected in bots network behavior and their intercommunication relationships. Existing botnet detection methods use flow features or topology features individually, which overlook the other type of feature. This affects model performance. In this paper, we propose a botnet detection model which uses graph convolutional network (GCN) to deeply fuse flow features and topology features for the first time. We construct communication graphs from network traffic and represent nodes with flow features. Due to the imbalance of existing public traffic flow datasets, it is impossible to train a GCN model on these datasets. Therefore, we use a balanced public communication graph dataset to pretrain a GCN model, thereby guaranteeing its capacity for identify topology features. We then feed the communication graph with flow features into the pretrained GCN. The output from the last hidden layer is treated as the fusion of flow and topology features. Additionally, by adjusting the number of layers in the GCN network, the model can effectively detect botnets under both C2 and P2P structures. Validated on the public ISCX2014 dataset, our approach achieves a remarkable recall rate 92.90% and F1-score 92.76% for C2 botnets, alongside recall rate 94.66% and F1-score of 92.35% for P2P botnets. These results not only demonstrate the effectiveness of our method, but also outperform the performance of the currently leading detection models.

cs.CR

Online Linearized LASSO

Sparse regression has been a popular approach to perform variable selection and enhance the prediction accuracy and interpretability of the resulting statistical model. Existing approaches focus on offline regularized regression, while the online scenario has rarely been studied. In this paper, we propose a novel online sparse linear regression framework for analyzing streaming data when data points arrive sequentially. Our proposed method is memory efficient and requires less stringent restricted strong convexity assumptions. Theoretically, we show that with a properly chosen regularization parameter, the $\ell_2$-norm statistical error of our estimator diminishes to zero in the optimal order of $\tilde{O}({\sqrt{s/t}})$, where $s$ is the sparsity level, $t$ is the streaming sample size, and $\tilde{O}(\cdot)$ hides logarithmic terms. Numerical experiments demonstrate the practical efficiency of our algorithm.

stat.ML

Classification of Positive Radial Solutions to A Weighted Biharmonic Equation

In this paper, we consider the weighted fourth order equation $$\Delta(|x|^{-\alpha}\Delta u)+\lambda \text{div}(|x|^{-\alpha-2}\nabla u)+\mu|x|^{-\alpha-4}u=|x|^\beta u^p\quad \text{in} \quad \mathbb{R}^n \backslash \{0\},$$ where $n\geq 5$, $-n<\alpha 1$ and $(p,\alpha,\beta,n)$ belongs to the critical hyperbola $$\frac{n+\alpha}{2}+\frac{n+\beta}{p+1}=n-2.$$ We prove the existence of radial solutions to the equation for some $\lambda$ and $\mu$. On the other hand, let $v(t):=|x|^{\frac{n-4-\alpha}{2}}u(|x|)$, $t=-\ln |x|$, then for the radial solution $u$ with non-removable singularity at origin, $v(t)$ is a periodic function if $\alpha \in (-2,n-4)$ and $\lambda$, $\mu$ satisfy some conditions; while for $\alpha \in (-n,-2]$, there exists a radial solution with non-removable singularity and the corresponding function $v(t)$ is not periodic. We also get some results about the best constant and symmetry breaking, which is closely related to the Caffarelli-Kohn-Nirenberg type inequality.

math.AP