SearcharxivSearch

arXiv subjects

Yuliang Sun

Publications and source records attributed to Yuliang Sun.

6 recordsLinked to original sources

Tencent Advertising Algorithm Challenge 2025: All-Modality Generative Recommendation

Generative recommender systems are rapidly emerging as a new paradigm for recommendation, where collaborative identifiers and/or multi-modal content are mapped into discrete token spaces and user behavior is modelled with autoregressive sequence models. Despite progress on multi-modal recommendation datasets, there is still a lack of public benchmarks that jointly offer large-scale, realistic and fully all-modality data designed specifically for generative recommendation (GR) in industrial advertising. To foster research in this direction, we organised the Tencent Advertising Algorithm Challenge 2025, a global competition built on top of two all-modality datasets for GR: TencentGR-1M and TencentGR-10M. Both datasets are constructed from real de-identified Tencent Ads logs and contain rich collaborative IDs and multi-modal representations extracted with state-of-the-art embedding models. The preliminary track (TencentGR-1M) provides 1 million user sequences with up to 100 interacted items each, where each interaction is labeled with exposure and click signals, while the final track (TencentGR-10M) scales this to 10 million users and explicitly distinguishes between click and conversion events at both the sequence and target level. This paper presents the task definition, data construction process, feature schema, baseline GR model, evaluation protocol, and key findings from top-ranked and award-winning solutions. Our datasets focus on multi-modal sequence generation in an advertising setting and introduce weighted evaluation for high-value conversion events. We release our datasets at https://huggingface.co/datasets/TAAC2025 and baseline implementations at https://github.com/TencentAdvertisingAlgorithmCompetition/baseline_2025 to enable future research on all-modality generative recommendation at an industrial scale. The official website is https://algo.qq.com/2025.

cs.IR

Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models

Large language models (LLMs) have been increasingly applied to various domains, which triggers increasing concerns about LLMs' safety on specialized domains, e.g. medicine. Despite prior explorations on general jailbreaking attacks, there are two challenges for applying existing attacks on testing the domain-specific safety of LLMs: (1) Lack of professional knowledge-driven attacks, (2) Insufficient coverage of domain knowledge. To bridge this gap, we propose a new task, knowledge-to-jailbreak, which aims to generate jailbreaking attacks from domain knowledge, requiring both attack effectiveness and knowledge relevance. We collect a large-scale dataset with 12,974 knowledge-jailbreak pairs and fine-tune a large language model as jailbreak-generator, to produce domain knowledge-specific jailbreaks. Experiments on 13 domains and 8 target LLMs demonstrate the effectiveness of jailbreak-generator in generating jailbreaks that are both threatening to the target LLMs and relevant to the given knowledge. We also apply our method to an out-of-domain knowledge base, showing that jailbreak-generator can generate jailbreaks that are comparable in harmfulness to those crafted by human experts. Data and code are available at: https://github.com/THU-KEG/Knowledge-to-Jailbreak/.

cs.CL

WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models

To mitigate the potential misuse of large language models (LLMs), recent research has developed watermarking algorithms, which restrict the generation process to leave an invisible trace for watermark detection. Due to the two-stage nature of the task, most studies evaluate the generation and detection separately, thereby presenting a challenge in unbiased, thorough, and applicable evaluations. In this paper, we introduce WaterBench, the first comprehensive benchmark for LLM watermarks, in which we design three crucial factors: (1) For benchmarking procedure, to ensure an apples-to-apples comparison, we first adjust each watermarking method's hyper-parameter to reach the same watermarking strength, then jointly evaluate their generation and detection performance. (2) For task selection, we diversify the input and output length to form a five-category taxonomy, covering $9$ tasks. (3) For evaluation metric, we adopt the GPT4-Judge for automatically evaluating the decline of instruction-following abilities after watermarking. We evaluate $4$ open-source watermarks on $2$ LLMs under $2$ watermarking strengths and observe the common struggles for current methods on maintaining the generation quality. The code and data are available at https://github.com/THU-KEG/WaterBench.

cs.CL

Real-Time Radar-Based Gesture Detection and Recognition Built in an Edge-Computing Platform

In this paper, a real-time signal processing frame-work based on a 60 GHz frequency-modulated continuous wave (FMCW) radar system to recognize gestures is proposed. In order to improve the robustness of the radar-based gesture recognition system, the proposed framework extracts a comprehensive hand profile, including range, Doppler, azimuth and elevation, over multiple measurement-cycles and encodes them into a feature cube. Rather than feeding the range-Doppler spectrum sequence into a deep convolutional neural network (CNN) connected with recurrent neural networks, the proposed framework takes the aforementioned feature cube as input of a shallow CNN for gesture recognition to reduce the computational complexity. In addition, we develop a hand activity detection (HAD) algorithm to automatize the detection of gestures in real-time case. The proposed HAD can capture the time-stamp at which a gesture finishes and feeds the hand profile of all the relevant measurement-cycles before this time-stamp into the CNN with low latency. Since the proposed framework is able to detect and classify gestures at limited computational cost, it could be deployed in an edge-computing platform for real-time applications, whose performance is notedly inferior to a state-of-the-art personal computer. The experimental results show that the proposed framework has the capability of classifying 12 gestures in real-time with a high F1-score.

eess.SP

The effect of internal magnetic field on collective flow in heavy ion collisions at intermediate energies

The properties of nuclear matter under extreme conditions of high temperature, density and isospin-asymmetry have attracted wide attentions in recent years. At present, heavy ion reactions in combination with corresponding model simulations are one of the most important ways to investigate this subject. It is known that a strong magnetic field can be created in heavy ion collisions. However, its effect on the motion of charged particles is usually neglected in previous transport model simulations. In this work, within the Ultra-relativistic Quantum Molecular Dynamics (UrQMD) model, the temporal evolution and spatial distribution of the internal magnetic field are calculated. The magnetic field strength is found to reach about $eB\approx470$ MeV$^{2}$ ($B\approx8\times10^{16}$ G) for Au+Au collisions at $E_{\text{lab}}$=1 GeV/nucleon with impact parameter of 7 fm. The magnetic field in Cu+Au collisions exhibits somewhat different spatial distribution from that in Au+Au collisions. The magnetic field is found to affect the directed flow of pions at forward and backward rapidities to some extent, dependent of the impact parameter and beam energy while the effect on the elliptic flow is small. This suggests that, because $π$ mesons produced in heavy ion collisions at intermediate energies are considered as a sensitive probe for the nuclear symmetry energy, it is necessary to consider the effect of the internal magnetic field.

nucl-th

Elliptic flow from Coulomb interaction and low density elastic scattering

In high energy heavy ion collisions and interacting cold atom systems, large elliptic flow anisotropies have been observed. For the large opacity ($ρσL\sim 10^{3}$) of the latter hydrodynamics is a natural consequence, but for the small opacity ($ρσL\sim 1$) of the former hydrodynamic description is questionable. To shed light onto the situation, we simulate the expansion of a low density Argon ion (or atom) system, initially trapped in an elliptical region, under the Coulomb interaction (or elastic scattering). Significant elliptic anisotropy is found in both cases, and the anisotropy depends on the initial spatial eccentricity and the density of the system. The results may provide insights into the physics of anisotropic flow in high energy heavy ion collisions and its role in the study of quantum chromodynamics.

nucl-th