SearcharxivSearch

arXiv subjects

Shudong Wang

Publications and source records attributed to Shudong Wang.

10 recordsLinked to original sources

Exploring the $t\bar{t}$ threshold at an electron-positron collider

Future electron-positron colliders offer a unique opportunity for high-precision measurements of the top-quark mass, width, strong coupling constant, and top-quark Yukawa coupling via a scan of the $t\bar{t}$ threshold. We present the first prospect study of the simultaneous determination of these parameters, incorporating the latest reference detector design for the Circular Electron-Positron Collider (CEPC). We find that the precision of the top-quark mass measurement can reach a few MeV excluding the theoretical uncertainty on the cross-section, which is nearly two orders of magnitude better than the high-luminosity LHC (HL-LHC) projections. The current theoretical uncertainty of the cross-section calculation is the limiting factor.

hep-ph

Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement Learning

Multi-task reinforcement learning (MTRL) seeks to learn a unified policy for diverse tasks, but often suffers from gradient conflicts across tasks. Existing masking-based methods attempt to mitigate such conflicts by assigning task-specific parameter masks. However, our empirical study shows that coarse-grained binary masks have the problem of over-suppressing key conflicting parameters, hindering knowledge sharing across tasks. Moreover, different tasks exhibit varying conflict levels, yet existing methods use a one-size-fits-all fixed sparsity strategy to keep training stability and performance, which proves inadequate. These limitations hinder the model's generalization and learning efficiency. To address these issues, we propose SoCo-DT, a Soft Conflict-resolution method based by parameter importance. By leveraging Fisher information, mask values are dynamically adjusted to retain important parameters while suppressing conflicting ones. In addition, we introduce a dynamic sparsity adjustment strategy based on the Interquartile Range (IQR), which constructs task-specific thresholding schemes using the distribution of conflict and harmony scores during training. To enable adaptive sparsity evolution throughout training, we further incorporate an asymmetric cosine annealing schedule to continuously update the threshold. Experimental results on the Meta-World benchmark show that SoCo-DT outperforms the state-of-the-art method by 7.6% on MT50 and by 10.5% on the suboptimal dataset, demonstrating its effectiveness in mitigating gradient conflicts and improving overall multi-task performance.

cs.LG

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled attention mechanisms for spatial, temporal, and view dimensions. However, these approaches often struggle to maintain consistency across dimensions, particularly when handling fast-moving objects that appear at different times and viewpoints. In this paper, we present CogDriving, a novel network designed for synthesizing high-quality multi-view driving videos. CogDriving leverages a Diffusion Transformer architecture with holistic-4D attention modules, enabling simultaneous associations across the spatial, temporal, and viewpoint dimensions. We also propose a lightweight controller tailored for CogDriving, i.e., Micro-Controller, which uses only 1.1% of the parameters of the standard ControlNet, enabling precise control over Bird's-Eye-View layouts. To enhance the generation of object instances crucial for autonomous driving, we propose a re-weighted learning objective, dynamically adjusting the learning weights for object instances during training. CogDriving demonstrates strong performance on the nuScenes validation set, achieving an FVD score of 37.8, highlighting its ability to generate realistic driving videos. The project can be found at https://luhannan.github.io/CogDrivingPage/.

cs.CV

Controllable Talking Face Generation by Implicit Facial Keypoints Editing

Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures that are intricately dependent on each other, complicating the process of re-editing image or video inputs. In this work, we present ControlTalk, a talking face generation method to control face expression deformation based on driven audio, which can construct the head pose and facial expression including lip motion for both single image or sequential video inputs in a unified manner. By utilizing a pre-trained video synthesis renderer and proposing the lightweight adaptation, ControlTalk achieves precise and naturalistic lip synchronization while enabling quantitative control over mouth opening shape. Our experiments show that our method is superior to state-of-the-art performance on widely used benchmarks, including HDTF and MEAD. The parameterized adaptation demonstrates remarkable generalization capabilities, effectively handling expression deformation across same-ID and cross-ID scenarios, and extending its utility to out-of-domain portraits, regardless of languages. Code is available at https://github.com/NetEase-Media/ControlTalk.

cs.CV

Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval

The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image frames) representations. However, some problems exist in the representation alignment of video and text, such as a text, and further each word, are of different importance for video frames. Besides, audio usually carries additional or critical information for TVR in the case that frames carry little valid information. Therefore, in TVR task, multi-granularity representation of text, including whole sentence and every word, and the modal of audio are salutary which are underutilized in most existing works. To address this, we propose a novel multi-granularity feature interaction module called MGFI, consisting of text-frame and word-frame, for video-text representations alignment. Moreover, we introduce a cross-modal feature interaction module of audio and text called CMFI to solve the problem of insufficient expression of frames in the video. Experiments on benchmark datasets such as MSR-VTT, MSVD, DiDeMo show that the proposed method outperforms the existing state-of-the-art methods.

cs.CV

Performance studies of jet flavor tagging and measurement of $R_b(R_c)$ using ParticleNet at CEPC

Jet flavor tagging plays a crucial role in the measurement of relative partial decay widths of $Z$ boson, denoted as $R_b$($R_c$), which is considered as a fundamental test of the Standard Model and sensitive probe to new physics. In this study, a Deep Learning algorithm, ParticleNet, is employed to enhance the performance of jet flavor tagging. The combined efficiency and purity of $c$-tagging is improved by more than 50\% compared to the Circular Electron Positron Collider (CEPC) baseline software. In order to measure $R_b$($R_c$) with this new flavor tagging approach, we have adopted the double-tagging method. The precision of $R_b$($R_c$) is improved significantly, in particular to $R_c$, which has seen a reduction in statistical uncertainty by 40\%.

hep-ex

Update of hadronic decays of $J/ψ$ and $ψ(2S)$ though virtual photons

The hadronic decay branching ratios of $J/ψ$ and $ψ(2S)$ through virtual photons, $B(J/ψ, ψ(2S) \rightarrow γ^*\rightarrow \text{hadrons})$, are updated by using the latest published measurements of the $R$ value and the branching ratios of $J/ψ, ψ(2S) \rightarrow l^+l^-$. Their respective precision increases by about 4 and 3 times.

hep-ex

Top quark mass measurements at the $t\bar{t}$ threshold with CEPC

We present a study of top quark mass measurements at the $t\bar{t}$ threshold based on CEPC. A centre-of-mass energy scan near two times of the top mass is performed and the measurement precision of top quark mass, width and $α_S$ are evaluated using the $t\bar{t}$ production rates. Realistic scan strategies at the threshold are discussed to maximise the sensitivity to the measurement of the top quark properties individually and simultaneously in the CEPC scenarios assuming a limited total luminosity of 100 fb$^{-1}$. With the optimal scan for individual property measurements, the top quark mass precision is expected to be 9 MeV, the top quark width precision is expected to be 26 MeV, and $α_S$ can be measured at a precision of 0.00039. Taking into account the uncertainties from theory, background subtraction, beam energy and luminosity spectrum, the top quark mass can be measured at a precision of 14 MeV optimistically and 34 MeV conservatively at CEPC.

hep-ex

Classify the Higgs decays with the PFN and ParticleNet at electron-positron colliders

Various Higgs factories are proposed to study the Higgs boson precisely and systematically in a model-independent way. In this study, the Particle Flow Network and ParticleNet techniques are used to classify the Higgs decays into multi-categories and the ultimate goal is to realize an "end-to-end" analysis. A Monte Carlo simulation study is performed to demonstrate the feasibility, and the performance looks rather promising. This result could be the basis of a "one-shop" analysis to measure all the branching fractions of the Higgs decays simultaneously.

hep-ex

Probing optical excitations in chevron-like armchair graphene nanoribbons

The bottom-up fabrication graphene nanoribbons (GNRs) has opened new opportunities to specifically control their electronic and optical properties by precisely controlling their atomic structure. Here, we address excitations in GNRs with periodic structural wiggles, so-called Chevron GNRs. Based on reflectance difference and high-resolution electron energy loss spectroscopies together with ab-initio simulations, we demonstrate that their excited-state properties are dominated by strongly bound excitons. The spectral fingerprints corresponding to different reaction stages in their bottom-up fabrication are also unequivocally identified, allowing us to follow the exciton build-up from the starting monomer precursor to the final ribbon structure.

cond-mat.mtrl-sci