SearcharxivSearch

arXiv subjects

Kang Tan

Publications and source records attributed to Kang Tan.

4 recordsLinked to original sources

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.

cs.CV

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.

cs.AI

Study of Human Push Recovery

Walking and push recovery controllers for humanoid robots have advanced throughout the years while unsolved gaps leading to undesirable behaviours still exist. Because previous studies are mainly pure engineering methods, while the use of data-driven methods has made impressive achievements in the field of control, we set motivation for exploration of control laws applied by human beings. Successful findings may help fill the gaps in current engineering-based controllers and can potentially improve the performance. In this thesis, we show our complete design and implementation of a set of experiments to collect and process human data, as well as data analysis for model fitting. Using the processed motion data and force data collected, we export the position, velocity and acceleration of our participants' centres of mass to form a one-dimensional point mass model defined by the direction of pushes. As a result, we find that proportional-derivative (PD) control can describe the underlying control law that people used and different PD gains are set for different phases of a push recovery trial within the scope of our study. Our final PD control fittings have an average root mean square error below 0.1m with above 90% of a trial's data points are taken into account on average. We also explore how different error metrics that the PD control model uses influence its performance but neither of the two metrics we proposed can help improve the performance. Finally, we have statistics on how people switch push recovery strategies based on their centre of mass properties at the start of the push. We find that the further the centre of mass is from the steady-state point and the higher the velocity it has, a person is more likely to make a step for push force compensation.

cs.RO

Experimental demonstration of broadband Lorentz non-reciprocity in an integrable photonic architecture based on Mach-Zehnder modulators

We demonstrate the first active optical isolator and circulator implemented in a linear and reciprocal material platform using commercial Mach-Zehnder modulators. In a proof-of-principle experiment based on single-mode polarization-maintaining fibers, we achieve more than 12.5 dB isolation over an unprecedented 8.7 THz bandwidth at telecommunication wavelengths, with only 9.1 dB total insertion loss. Our architecture provides a practical answer to the challenge of non-reciprocal light routing in photonic integrated circuits.

physics.optics