SearcharxivSearch

arXiv subjects

Han Qin

Publications and source records attributed to Han Qin.

15 recordsLinked to original sources

Detecting Privileged Documents by Ranking Connected Network Entities

This paper presents a link analysis approach for identifying privileged documents by constructing a network of human entities derived from email header metadata. Entities are classified as either counsel or non-counsel based on a predefined list of known legal professionals. The core assumption is that individuals with frequent interactions with lawyers are more likely to participate in privileged communications. To quantify this likelihood, an algorithm assigns a score to each entity within the network. By utilizing both entity scores and the strength of their connections, the method enhances the identification of privileged documents. Experimental results demonstrate the algorithm's effectiveness in ranking legal entities for privileged document detection.

cs.IR

A Comparative Study of Retrieval Methods in Azure AI Search

Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.

cs.IR

Leveraging Machine Learning and Large Language Models for Automated Image Clustering and Description in Legal Discovery

The rapid increase in digital image creation and retention presents substantial challenges during legal discovery, digital archive, and content management. Corporations and legal teams must organize, analyze, and extract meaningful insights from large image collections under strict time pressures, making manual review impractical and costly. These demands have intensified interest in automated methods that can efficiently organize and describe large-scale image datasets. This paper presents a systematic investigation of automated cluster description generation through the integration of image clustering, image captioning, and large language models (LLMs). We apply K-means clustering to group images into 20 visually coherent clusters and generate base captions using the Azure AI Vision API. We then evaluate three critical dimensions of the cluster description process: (1) image sampling strategies, comparing random, centroid-based, stratified, hybrid, and density-based sampling against using all cluster images; (2) prompting techniques, contrasting standard prompting with chain-of-thought prompting; and (3) description generation methods, comparing LLM-based generation with traditional TF-IDF and template-based approaches. We assess description quality using semantic similarity and coverage metrics. Results show that strategic sampling with 20 images per cluster performs comparably to exhaustive inclusion while significantly reducing computational cost, with only stratified sampling showing modest degradation. LLM-based methods consistently outperform TF-IDF baselines, and standard prompts outperform chain-of-thought prompts for this task. These findings provide practical guidance for deploying scalable, accurate cluster description systems that support high-volume workflows in legal discovery and other domains requiring automated organization of large image collections.

cs.IR

Empirical Evaluation of Embedding Models in the Context of Text Classification in Document Review in Construction Delay Disputes

Text embeddings are numerical representations of text data, where words, phrases, or entire documents are converted into vectors of real numbers. These embeddings capture semantic meanings and relationships between text elements in a continuous vector space. The primary goal of text embeddings is to enable the processing of text data by machine learning models, which require numerical input. Numerous embedding models have been developed for various applications. This paper presents our work in evaluating different embeddings through a comprehensive comparative analysis of four distinct models, focusing on their text classification efficacy. We employ both K-Nearest Neighbors (KNN) and Logistic Regression (LR) to perform binary classification tasks, specifically determining whether a text snippet is associated with 'delay' or 'not delay' within a labeled dataset. Our research explores the use of text snippet embeddings for training supervised text classification models to identify delay-related statements during the document review process of construction delay disputes. The results of this study highlight the potential of embedding models to enhance the efficiency and accuracy of document analysis in legal contexts, paving the way for more informed decision-making in complex investigative scenarios.

cs.IR

Probing heavy triplet leptons of the type-III seesaw mechanism at future muon colliders

We study the search potential of heavy triplet leptons in Type III Seesaw mechanism at future high-energy and high-luminosity muon colliders. The impact of up-to-date neutrino oscillation results is taken into account on neutrino parameters and the consequent decay patterns of heavy leptons in Type III Seesaw. The three heavy leptons in the Casas-Ibarra parametrization with diagonal unity matrix are distinguishable in their decay modes. We consider the pair production of charged triplet leptons $E^+ E^-$ through both $\mu^+\mu^-$ annihilation and vector boson fusion (VBF) processes. The leading-order framework of electroweak parton distribution functions are utilized to calculate the VBF cross sections. The charged triplet leptons can be probed for masses above $\sim 1$ TeV. The $E^+E^-\to ZZ\ell^+\ell^-$ channel can be further utilized to fully reconstruct the three triplet leptons and distinguish neutrino mass patterns. The pair production of heavy neutrinos $N$ and the associated production $E^\pm N$ are only induced by VBF processes and lead to lepton-number-violating (LNV) signature. We also study the search potential of LNV processes at future high-energy muon collider with $\sqrt{s}=30$ TeV.

hep-ph

Use Image Clustering to Facilitate Technology Assisted Review

During the past decade breakthroughs in GPU hardware and deep neural networks technologies have revolutionized the field of computer vision, making image analytical potentials accessible to a range of real-world applications. Technology Assisted Review (TAR) in electronic discovery though traditionally has dominantly dealt with textual content, is witnessing a rising need to incorporate multimedia content in the scope. We have developed innovative image analytics applications for TAR in the past years, such as image classification, image clustering, and object detection, etc. In this paper, we discuss the use of image clustering applications to facilitate TAR based on our experiences in serving clients. We describe our general workflow on leveraging image clustering in tasks and use statistics from real projects to showcase the effectiveness of using image clustering in TAR. We also summarize lessons learned and best practices on using image clustering in TAR.

cs.CV

Leptonic Scalars and Collider Signatures in a UV-complete Model

We study the non-standard interactions of neutrinos with light leptonic scalars ($\phi$) in a global $(B-L)$-conserved ultraviolet (UV)-complete model. The model utilizes Type-II seesaw motivated neutrino interactions with an $SU(2)_L$-triplet scalar, along with an additional singlet in the scalar sector. This UV-completion leads to an enriched spectrum and consequently new observable signatures. We examine the low-energy lepton flavor violation constraints, as well as the perturbativity and unitarity constraints on the model parameters. Then we lay out a search strategy for the unique signature of the model resulting from the leptonic scalars at the hadron colliders via the processes $H^{\pm\pm} \to W^\pm W^\pm \phi$ and $H^\pm \to W^\pm \phi$ for both small and large leptonic Yukawa coupling cases. We find that via these associated production processes at the HL-LHC, the prospects of doubly-charged scalar $H^{\pm\pm}$ can reach up to 800 (500) GeV and 1.1 (0.8) TeV at the $2\sigma \ (5\sigma)$ significance for small and large Yukawa couplings, respectively. A future 100 TeV hadron collider will further increase the mass reaches up to 3.8 (2.6) TeV and 4 (2.7) TeV, at the $2\sigma \ (5\sigma)$ significance, respectively. We also demonstrate that the mass of $\phi$ can be determined at about 10% accuracy at the LHC for the large Yukawa coupling case even though it escapes as missing energy from the detectors.

hep-ph

Directly Probing the Higgs-top Coupling at High Scales

We explore the sensitivity to new physics for the coupling of the Higgs boson ($h$) and top quark ($t$) at high energy scales with the process $pp\to t\bar{t}h$ at the high-luminosity LHC. This process probes the coupling in both the space-like and time-like domains at a high scale, complementary to the off-shell Higgs processes in the time-like domain. The effects from physics beyond the Standard Model are parametrized in terms of the effective field theory framework and a non-local Higgs-top form factor. Focusing on the boosted Higgs regime in association with jet substructure techniques, we show that the present search can directly probe the Higgs-top coupling to good precision, providing a strong sensitivity to the new physics scale.

hep-ph

Application of Deep Learning in Recognizing Bates Numbers and Confidentiality Stamping from Images

In eDiscovery, it is critical to ensure that each page produced in legal proceedings conforms with the requirements of court or government agency production requests. Errors in productions could have severe consequences in a case, putting a party in an adverse position. The volume of pages produced continues to increase, and tremendous time and effort has been taken to ensure quality control of document productions. This has historically been a manual and laborious process. This paper demonstrates a novel automated production quality control application which leverages deep learning-based image recognition technology to extract Bates Number and Confidentiality Stamping from legal case production images and validate their correctness. Effectiveness of the method is verified with an experiment using a real-world production data.

cs.IR

Off-shell Higgs Couplings in $H^*\to ZZ\to \ell\ell\nu\nu$

We explore the new physics reach for the off-shell Higgs boson measurement in the ${pp \to H^* \rightarrow Z(\ell^{+}\ell^{-})Z(\nu\bar{\nu})}$ channel at the high-luminosity LHC. The new physics sensitivity is parametrized in terms of the Higgs boson width, effective field theory framework, and a non-local Higgs-top coupling form factor. Adopting Machine-learning techniques, we demonstrate that the combination of a large signal rate and a precise phenomenological probe for the process energy scale, due to the transverse $ZZ$ mass, leads to significant sensitivities beyond the existing results in the literature for the new physics scenarios considered.

hep-ph

Tracing the main elements and electron orbitals that induce superconducting phase transition

The experimental determination of the superconducting transition requires the observation of the emergence of zero-resistance and perfect diamagnetism state. Based on the close relationship between superconducting transition temperature (Tc) and electron density of states (DOS), we take two typical superconducting materials Hg and ZrTe3 as samples and calculate their DOS versus temperature under different pressures by using the first-principle molecular dynamics simulations. According to the analysis of the calculation results, the main contributors that induce superconducting transitions are deduced by tracing the variation of partial density of states near Tc. In particular, the microscopic mechanism of pressure increasing Tc is further analyzed.

cond-mat.supr-con

Empirical Comparisons of CNN with Other Learning Algorithms for Text Classification in Legal Document Review

Research has shown that Convolutional Neural Networks (CNN) can be effectively applied to text classification as part of a predictive coding protocol. That said, most research to date has been conducted on data sets with short documents that do not reflect the variety of documents in real world document reviews. Using data from four actual reviews with documents of varying lengths, we compared CNN with other popular machine learning algorithms for text classification, including Logistic Regression, Support Vector Machine, and Random Forest. For each data set, classification models were trained with different training sample sizes using different learning algorithms. These models were then evaluated using a large randomly sampled test set of documents, and the results were compared using precision and recall curves. Our study demonstrates that CNN performed well, but that there was no single algorithm that performed the best across the combination of data sets and training sample sizes. These results will help advance research into the legal profession's use of machine learning algorithms that maximize performance.

cs.IR

Image Analytics for Legal Document Review: A Transfer Learning Approach

Though technology assisted review in electronic discovery has been focusing on text data, the need of advanced analytics to facilitate reviewing multimedia content is on the rise. In this paper, we present several applications of deep learning in computer vision to Technology Assisted Review of image data in legal industry. These applications include image classification, image clustering, and object detection. We use transfer learning techniques to leverage established pretrained models for feature extraction and fine tuning. These applications are first of their kind in the legal industry for image document review. We demonstrate effectiveness of these applications with solving real world business challenges.

cs.CV

Using Google Analytics to Support Cybersecurity Forensics

Web traffic is a valuable data source, typically used in the marketing space to track brand awareness and advertising effectiveness. However, web traffic is also a rich source of information for cybersecurity monitoring efforts. To better understand the threat of malicious cyber actors, this study develops a methodology to monitor and evaluate web activity using data archived from Google Analytics. Google Analytics collects and aggregates web traffic, including information about web visitors' location, date and time of visit, visited webpages, and searched keywords. This study seeks to streamline analysis of this data and uses rule-based anomaly detection and predictive modeling to identify web traffic that deviates from normal patterns. Rather than evaluating pieces of web traffic individually, the methodology seeks to emulate real user behavior by creating a new unit of analysis: the user session. User sessions group individual pieces of traffic from the same location and date, which transforms the available information from single point-in-time snapshots to dynamic sessions showing users' trajectory and intent. The result is faster and better insight into large volumes of noisy web traffic.

cs.IR

Empirical Study of Deep Learning for Text Classification in Legal Document Review

Predictive coding has been widely used in legal matters to find relevant or privileged documents in large sets of electronically stored information. It saves the time and cost significantly. Logistic Regression (LR) and Support Vector Machines (SVM) are two popular machine learning algorithms used in predictive coding. Recently, deep learning received a lot of attentions in many industries. This paper reports our preliminary studies in using deep learning in legal document review. Specifically, we conducted experiments to compare deep learning results with results obtained using a SVM algorithm on the four datasets of real legal matters. Our results showed that CNN performed better with larger volume of training dataset and should be a fit method in the text classification in legal industry.

cs.IR