SearcharxivSearch

arXiv subjects

Haoyuan Tang

Publications and source records attributed to Haoyuan Tang.

7 recordsLinked to original sources

The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls

In February 2024 the U.S. Federal Communications Commission (FCC) placed AI-generated voices under the Telephone Consumer Protection Act (TCPA). Yet no peer-reviewed measurement says how much unwanted call traffic is placed by a machine, or how much of that machine speech is synthesized rather than played from a recording. We report both with a disclosed pipeline. An interactive voice honeypot (language-model personas on real U.S. numbers, the caller recorded on its own track) recorded 10,987 calls over 66 days; 11 days on which our stack answered silently are set aside. Three instruments read each opening: an audio fingerprint that finds the same recording played on other calls, a commercial synthetic-speech detector on the caller's first ten seconds, and blinded listeners who check what it flags. Of the 7,233 calls our persona greeted on normal days, 13.8% open with a recording we also heard on another call, and 13.1% with fresh audio the detector labels synthetic. A further 9.9% open with a caller who never spoke after our greeting, 54.2% with fresh audio the detector labels human, and 9.0% could not be scored. Machine-voiced openings are therefore at least 26.9%, a further tenth of calls are silent connections we read as machine-placed, and replays of a recording make up 45% of the detector's own rate (29.3% of 6,192 scored openings). The same waveform played on two calls lands on opposite sides of the detector's threshold 13.6% of the time, and eleven listeners confirm 54.4% of what it flags. Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%); 0.44% disclose automation. Prevalence tracks how long a bait number has circulated (59% against 19% in the same weeks): seeding history, not calendar time, explains the trend. Campaigns outlast their numbers: one recorded compliance notice opens calls in six campaigns, and one synthetic voice serves nine.

cs.CR

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters is not task success but whether the owner would endorse how their proxy handled the call. We evaluate only the text-domain conversational decision layer; speech recognition, audio interaction, end-to-end latency, and handset execution are outside scope. We present CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement. Each is paired, where available, with a counter-metric and an uncertainty estimate; no benchmark-wide Q1-Q5 composite or leaderboard score is defined. Three guardedness diagnostics identify candidate cases for a toolless proxy that holds no credentials and calls no tools. Across three model families represented by paired 4-bit checkpoints (0.6-4B), the primary scoring snapshot gives the larger checkpoint higher point estimates on several service, recall, and plausibility measures, while triage discrimination follows a different ordering. Bare scam-side TPR rewards universal suspicion, and pairwise separation changes when legitimate-side false positives are included and across judge snapshots. Scripted degenerate agents expose further floors, including a hangup-and-echo policy with entity recall 1.000. We report quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.

cs.CR

VDSB-GWSyn: Diffusion Schr\"{o}dinger Bridge for Controllable and Anatomically Feasible Guidewire Synthesis in Coronary Angiography

Coronary guidewire endpoint localization is a fundamental capability for computer-assisted PCI, and its importance increases as robot-assisted PCI is progressively adopted to reduce operator radiation exposure. However, the scarcity of annotated CAG images with guidewires and the limited adaptability of existing guidewire synthesis models remain key bottlenecks for guidewire endpoint localization. To address this issue, we propose VDSB-GWSyn, a Diffusion Schr\"{o}dinger Bridge (DSB) model-based framework, enabling synthesis of controllable, high-fidelity guidewire samples under complex anatomical backgrounds. VDSB-GWSyn first uses our shape prior algorithm to learn the basic guidewire geometry. It then generates guidewire masks under constraints imposed by the vessel segmentation masks and outputs the corresponding endpoint coordinates. Finally, it synthesizes realistic guidewire samples on real CAG images using DSB conditioned with SPADE. Experimental results show that the guidewire samples synthesized by VDSB-GWSyn achieve favorable ROI-FID and ROI-KID, as well as high IPR scores. In addition, incorporating our synthesized data for synthetic pre-training followed by real fine-tuning substantially improves downstream guidewire endpoint localization, reducing MPE from 16.01~px to 7.71~px and increasing PCK at 3~px from 52.63\% to 86.27\%, leading to more clinically reliable deployment of robot-assisted guidewire delivery systems. Moreover, the core design philosophy of controllable device synthesis with strict background preservation and anatomical feasibility constraints has the potential to transfer to other interventional device perception tasks where annotated data are scarce.

cs.CV

FLUID: A Fine-Grained Lightweight Urban Signalized-Intersection Dataset of Dense Conflict Trajectories

The trajectory data of traffic participants (TPs) is a fundamental resource for evaluating traffic conditions and optimizing policies, especially at urban intersections. Although data acquisition using drones is efficient, existing datasets still have limitations in scene representativeness, information richness, and data fidelity. This study introduces FLUID, comprising a fine-grained trajectory dataset that captures dense conflicts at typical urban signalized intersections, and a lightweight, full-pipeline framework for drone-based trajectory processing. FLUID covers three distinct intersection types, with approximately 5 hours of recording time and featuring over 20,000 TPs across 8 categories. Notably, the dataset records an average of 2.8 vehicle conflicts per minute across all scenes, with roughly 15% of all recorded motor vehicles directly involved in these conflicts. FLUID provides comprehensive data, including trajectories, traffic signals, maps, and raw videos. Comparison with the DataFromSky platform and ground-truth measurements validates its high spatio-temporal accuracy. Through a detailed classification of motor vehicle conflicts and violations, FLUID reveals a diversity of interactive behaviors, demonstrating its value for human preference mining, traffic behavior modeling, and autonomous driving research.

cs.RO

Electronic structures and magnetism in van der Waals flat-band material Ni$_{3}$GeTe$_{2}$

The study of magnetism in two-dimensional materials has garnered significant interest, driven by fundamental investigations into low-dimensional magnetic phenomena and their potential for applications in spintronic devices. Through dynamical mean-field theory calculations, we demonstrate that Ni$_{3}$GeTe$_{2}$ exhibits flat-band characteristics resulting from the geometric frustration of its layered triangular lattice. These flat bands are further renormalized due to electronic correlation. Our calculations reveal that the magnetic order of Ni atoms is significantly influenced by both the Coulomb interaction and Hund's coupling, indicating that the physics of Ni atoms is situated in an intermediate region between Hundness and Mottness. Additionally, our results show that Ni atoms experience significant spin fluctuations in their local moments, maintaining paramagnetism at low temperatures. Furthermore, we investigate the effect of vacancies, finding a substantial suppression of the density of states at the Fermi level. The physical mechanisms uncovered by our study provide a comprehensive understanding of the novel properties exhibited in this material.

cond-mat.str-el

Spin disorder state induced by Mg$^{2+}$ doping in a Kitaev material Na$_{3}$Co$_{2}$SbO$_{6}$

Due to the dominant Kitaev exchange interactions, the cobaltate, Na$_{3}$Co$_{2}$SbO$_{6}$, has been considered to be approximate to the Kitaev quantum spin liquid (QSL). Here, we investigate both magnetic dilution and chemical pressure effects of Na$_{3}$Co$_{2}$SbO$_{6}$ by the substitutions of Mg$^{2+}$ for Co$^{2+}$ through the structural, optical, magnetic and thermodynamic measurements. No structural transition has been observed, and the bandgaps remain constant in all doping levels. Combining with the magnetic and thermodynamic measurements, we find that the antiferromagnetic transition temperature is continuously suppressed with increasing Mg doping levels and completely disappears at $x=0.2$. Interestingly, when the doping level $x$ is larger than 0.2, neither long-range magnetic order nor spin glass state has been detected, and the specific heat has a residual linear term at zero field. All features indicate that Na$_{3}$(Co$_{2-x}$Mg$_{x}$)SbO$_{6}$ system enters into a novel spin disorder (NSD) state.

cond-mat.str-el

Planimation

Planimation is a modular and extensible open source framework to visualise sequential solutions of planning problems specified in PDDL. We introduce a preliminary declarative PDDL-like animation profile specification, expressive enough to synthesise animations of arbitrary initial states and goals of a benchmark with just a single profile.

cs.AI