SearcharxivSearch

arXiv subjects

Wenjun Qiu

Publications and source records attributed to Wenjun Qiu.

11 recordsLinked to original sources

Breaking the illusion: Automated Reasoning of GDPR Consent Violations

Recent privacy regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) have established legal requirements for obtaining user consent regarding the collection, use, and sharing of personal data. These regulations emphasize that consent must be informed, freely given, specific, and unambiguous. However, there are still many violations, which highlight a gap between legal expectations and actual implementation. Consent mechanisms embedded in functional web forms across websites play a critical role in ensuring compliance with data protection regulations such as the GDPR and CCPA, as well as in upholding user autonomy and trust. However, current research has primarily focused on cookie banners and mobile app dialogs. These forms are diverse in structure, vary in legal basis, and are often difficult to locate or evaluate, creating a significant challenge for automated consent compliance auditing. In this work, we present Cosmic, a novel automated framework for detecting consent-related privacy violations in web forms. We evaluate our developed tool for auditing consent compliance in web forms, across 5,823 websites and 3,598 forms. Cosmic detects 3,384 violations on 94.1% of consent forms, covering key GDPR principles such as freely given consent, purpose disclosure, and withdrawal options. It achieves 98.6% and 99.1% TPR for consent and violation detection, respectively, demonstrating high accuracy and real-world applicability.

cs.CR

Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning

A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric , which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grain labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly $0.92-$1.71 per labeled text segment. Calpric 's training process also generates a labeled data set of 16K privacy policy text segments across 9 Data categories with balanced positive and negative samples.

cs.CL

A Survey on Poisoning Attacks Against Supervised Machine Learning

With the rise of artificial intelligence and machine learning in modern computing, one of the major concerns regarding such techniques is to provide privacy and security against adversaries. We present this survey paper to cover the most representative papers in poisoning attacks against supervised machine learning models. We first provide a taxonomy to categorize existing studies and then present detailed summaries for selected papers. We summarize and compare the methodology and limitations of existing literature. We conclude this paper with potential improvements and future directions to further exploit and prevent poisoning attacks on supervised models. We propose several unanswered research questions to encourage and inspire researchers for future work.

cs.CR

HistBERT: A Pre-trained Language Model for Diachronic Lexical Semantic Analysis

Contextualized word embeddings have demonstrated state-of-the-art performance in various natural language processing tasks including those that concern historical semantic change. However, language models such as BERT was trained primarily on contemporary corpus data. To investigate whether training on historical corpus data improves diachronic semantic analysis, we present a pre-trained BERT-based language model, HistBERT, trained on the balanced Corpus of Historical American English. We examine the effectiveness of our approach by comparing the performance of the original BERT and that of HistBERT, and we report promising results in word similarity and semantic shift analysis. Our work suggests that the effectiveness of contextual embeddings in diachronic semantic analysis is dependent on the temporal profile of the input text and care should be taken in applying this methodology to study historical semantic change.

cs.CL

Deep Active Learning with Crowdsourcing Data for Privacy Policy Classification

Privacy policies are statements that notify users of the services' data practices. However, few users are willing to read through policy texts due to the length and complexity. While automated tools based on machine learning exist for privacy policy analysis, to achieve high classification accuracy, classifiers need to be trained on a large labeled dataset. Most existing policy corpora are labeled by skilled human annotators, requiring significant amount of labor hours and effort. In this paper, we leverage active learning and crowdsourcing techniques to develop an automated classification tool named Calpric (Crowdsourcing Active Learning PRIvacy Policy Classifier), which is able to perform annotation equivalent to those done by skilled human annotators with high accuracy while minimizing the labeling cost. Specifically, active learning allows classifiers to proactively select the most informative segments to be labeled. On average, our model is able to achieve the same F1 score using only 62% of the original labeling effort. Calpric's use of active learning also addresses naturally occurring class imbalance in unlabeled privacy policy datasets as there are many more statements stating the collection of private information than stating the absence of collection. By selecting samples from the minority class for labeling, Calpric automatically creates a more balanced training set.

cs.CR

Fundamental limits to nanoparticle extinction

We show that there are shape-independent upper bounds to the extinction cross section per unit volume of randomly oriented nanoparticles, given only material permittivity. Underlying the limits are restrictive sum rules that constrain the distribution of quasistatic eigenvalues. Surprisingly, optimally-designed spheroids, with only a single quasistatic degree of freedom, reach the upper bounds for four permittivity values. Away from these permittivities, we demonstrate computationally-optimized structures that surpass spheroids and approach the fundamental limits.

physics.optics

Comment on "A self-assembled three-dimensional cloak in the visible" in Scientific Reports 3, 2328

Mühlig et. al. propose and fabricate a "cloak" comprised of nano-particles on the surface of a sub-wavelength silica sphere. However, the coating only reduces the scattered fields. This is achieved by increased absorption, such that total extinction increases at all wavelengths. An object creating a large shadow is generally not considered to be cloaked; functionally, in contrast to the relatively few structures that can reduce total extinction, there are many that can reduce scattering alone.

physics.optics

Tailorable Stimulated Brillouin Scattering in Nanoscale Silicon Waveguides

While nanoscale modal confinement radically enhances a variety of nonlinear light-matter interactions within silicon waveguides, traveling-wave stimulated Brillouin scattering nonlinearities have never been observed in silicon nanophotonics. Through a new class of hybrid photonic-phononic waveguides, we demonstrate tailorable traveling-wave forward stimulated Brillouin scattering in nanophotonic silicon waveguides for the first time, yielding 3000 times stronger forward SBS responses than any previous waveguide system. Simulations reveal that a coherent combination of electrostrictive forces and radiation pressures are responsible for greatly enhanced photon-phonon coupling at nano-scales. Highly tailorable Brillouin nonlinearities are produced by engineering the structure of a membrane-suspended waveguide to yield Brillouin resonances from 1 to 18 GHz through high quality-factor (>1000) phonon modes. Such wideband and tailorable stimulated Brillouin scattering in silicon photonics could enable practical realization of on-chip slow-light devices, RF-photonic filtering and sensing, and ultra-narrow-band laser sources by using standard semiconductor fabrication and CMOS technologies.

physics.optics

Stimulated brillouin scattering in slow light waveguides

We develop a general method of calculating Stimulated Brillouin Scattering (SBS) gain coefficient in axially periodic waveguides. Applying this method to a silicon periodic waveguide suspended in air, we demonstrate that SBS nonlinearity can be dramatically enhanced at the brillouin zone boundary where the decreased group velocity of light magnifies photon-phonon interaction. In addition, we show that the symmetry plane perpendicular to the propagation axis plays an important role in both forward and backward SBS processes. In forward SBS, only elastic modes which are even about this plane are excitable. In backward SBS, the SBS gain coefficients of elastic modes approach to either infinity or constants, depending on their symmetry about this plane at $q=0$.

physics.optics

Stimulated brillouin scattering in nanoscale silicon step-index waveguides: A general framework of selection rules and calculating SBS gain

We develop a general framework of evaluating the gain coefficient of Stimulated Brillouin Scattering (SBS) in optical waveguides via the overlap integral between optical and elastic eigen-modes. We show that spatial symmetry of the optical force dictates the selection rules of the excitable elastic modes. By applying this method to a rectangular silicon waveguide, we demonstrate the spatial distributions of optical force and elastic eigen-modes jointly determine the magnitude and scaling of SBS gain coefficient in both forward and backward SBS processes. We further apply this method to inter-modal SBS process, and demonstrate that the coupling between distinct optical modes are necessary to excite elastic modes with all possible symmetries.

cond-mat.mes-hall

Dynamic Studies of Scaffold-dependent Mating Pathway in Yeast

The mating pathway in \emph{Saccharomyces cerevisiae} is one of the best understood signal transduction pathways in eukaryotes. It transmits the mating signal from plasma membrane into the nucleus through the G-protein coupled receptor and the mitogen-activated protein kinase (MAPK) cascade. According to the current understandings of the mating pathway, we construct a system of ordinary differential equations to describe the process. Our model is consistent with a wide range of experiments, indicating that it captures some main characteristics of the signal transduction along the pathway. Investigation with the model reveals that the shuttling of the scaffold protein and the dephosphorylation of kinases involved in the MAPK cascade cooperate to regulate the response upon pheromone induction and to help preserving the fidelity of the mating signaling. We explored factors affecting the dose-response curves of this pathway and found that both negative feedback and concentrations of the proteins involved in the MAPK cascade play crucial role. Contrary to some other MAPK systems where signaling sensitivity is being amplified successively along the cascade, here the mating signal is transmitted through the cascade in an almost linear fashion.

q-bio.MN