SearcharxivSearch

arXiv subjects

Anirban Mukherjee

Publications and source records attributed to Anirban Mukherjee.

At least 19 recordsLinked to original sources

Quantifying translational and bond-orientational order metrics in hyperuniform and nonhyperuniform many-particle systems

Quantifying the degree of order/disorder in many-particle systems remains an outstanding problem in physics, materials science, and mathematics. To this end, we consider the translational order metric $τ_T$, defined as the squared $L^2$ norm of the total correlation function $h(\mathbf{r})$, and introduce its bond-orientational analogue $τ_O$, defined from the weighted total correlation function $h_{\mathbf{f}}(\mathbf{r})$ using local orientational weights (S. Torquato et al., Phys. Rev. X 16, 011042 (2026)). The pair $(τ_T,τ_O)$ places both forms of order on a common two-point statistical footing. We compute both metrics for two nonhyperuniform sphere-packing models in 2D and 3D as functions of packing fraction $ϕ$: 1) equilibrium hard particles and 2) nonequilibrium random sequential addition (RSA) packings. For the nonhyperuniform systems, bond-orientational order remains subdominant along the equilibrium-fluid and RSA configurations; however, its magnitude relative to the translational metric increases near the upper end of the equilibrium-fluid branches, much more strongly in 2D than in 3D, and the two metrics become comparable along the sampled crystal branches. At common packing fractions, equilibrium fluids and RSA packings trace distinct $(τ_T,τ_O)$ trajectories, revealing preparation-dependent differences in structural order. As a representative hyperuniform family, we study 2D disordered stealthy hyperuniform (SHU) ground states for $0<χ<1/2$, where $χ$ is the stealthiness parameter. Within the disordered SHU phase, bond-orientational order remains subdominant, but its magnitude relative to the translational metric increases toward the disorder-to-order threshold. In all three models, $τ_T$ and $τ_O$ are positively correlated beyond the Poisson-reference regime: $τ_O$ increases monotonically with $τ_T$ across the sampled state points.

cond-mat.stat-mech

Copyright Laundering Through the AI Ouroboros: Adapting the 'Fruit of the Poisonous Tree' Doctrine to Recursive AI Training

Copyright enforcement rests on an evidentiary bargain: a plaintiff must show both the defendant's access to the work and substantial similarity in the challenged output. That bargain comes under strain when AI systems are trained through multi-generational pipelines with recursive synthetic data. As successive models are tuned on the outputs of its predecessors, any copyrighted material absorbed by an early model is diffused into deeper statistical abstractions. The result is an evidentiary blind spot where overlaps that emerge look coincidental, while the chain of provenance is too attenuated to trace. These conditions are ripe for "copyright laundering"--the use of multi-generational synthetic pipelines, an "AI Ouroboros," to render traditional proof of infringement impracticable. This Article adapts the "fruit of the poisonous tree" (FOPT) principle to propose a AI-FOPT standard: if a foundational AI model's training is adjudged infringing (either for unlawful sourcing or for non-transformative ingestion that fails fair-use), then subsequent AI models principally derived from the foundational model's outputs or distilled weights carry a rebuttable presumption of taint. The burden shifts to downstream developers--those who control the evidence of provenance--to restore the evidentiary bargain by affirmatively demonstrating a verifiably independent and lawfully sourced lineage or a curative rebuild, without displacing fair-use analysis at the initial ingestion stage. Absent such proof, commercial deployment of tainted models and their outputs is actionable. This Article develops the standard by specifying its trigger, presumption, and concrete rebuttal paths (e.g., independent lineage or verifiable unlearning); addresses counterarguments concerning chilling innovation and fair use; and demonstrates why this lineage-focused approach is both administrable and essential.

cs.CY

Operational Agency: A Permeable Legal Fiction for Tracing Culpability in AI Systems

Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood-a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)-a permeable legal fiction structured as an ex post evidentiary framework-and Operational Agency Graph (OAG), a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI's observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for a standard of care). OAG operationalizes that analysis by embedding these characteristics in a causal graph to trace and apportion culpability among developers, fine-tuners, deployers, and users. Drawing on corporate criminal liability, the innocent-agent doctrine, and secondary and vicarious liability frameworks, the Article shows how OA and OAG strengthen existing doctrines. Across five real-world case studies spanning tort, civil rights, constitutional law, and antitrust, it demonstrates how the framework addresses challenges ranging from autonomous vehicle collisions to algorithmic price-fixing, offering courts a principled evidentiary method-and legislatures and industry a conceptual foundation-to ensure human accountability keeps pace with technological autonomy, without conferring personhood on AI.

cs.CY

Private Again: Artificial Intelligence Agents Restore Anonymity---Foreclosing Discrimination and Its Proof

Artificial intelligence agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without revealing who that principal is. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and *Iqbal*-era pleading needs specific factual allegations---doctrinal predicates that anonymous transactions never generate. The effects fall asymmetrically: those most vulnerable to discrimination are least able to afford the shield and, when harms remain, least able to prove them. The challenge for the law shifts from detecting and remedying algorithmic discrimination to governing agent-mediated anonymity as civil rights infrastructure: ensuring access to privacy-preserving agents, regulating abuse without forced identification, and deciding whether retailers may refuse to deal with agents at all.

cs.CY

Effective hyperuniformity in time-integrated stochastic Turing patterns

Demographic noise generates stochastic Turing patterns even when reaction-diffusion systems are deterministically stable. We show analytically and verify numerically in the Levin-Segel model that temporal integration of configurations reveals emergent large-scale organization. The intensive number variance in a window of size $R \gg 1$ approaches a finite reaction-kinetic floor as $1/R$, over a spatial range growing by orders of magnitude near the Turing instability. This yields an effectively hyperuniform, fine-tuning-free regime previously unidentified in non-conserved multispecies stochastic systems.

cond-mat.stat-mech

Fluid Agency in AI Systems: A Case for Functional Equivalence in Copyright, Patent, and Tort

Modern Artificial Intelligence (AI) systems lack human-like consciousness or culpability, yet they exhibit fluid agency: behavior that is (i) stochastic (probabilistic and path-dependent), (ii) dynamic (co-evolving with user interaction), and (iii) adaptive (able to reorient across contexts). Fluid agency generates valuable outputs but collapses attribution, irreducibly entangling human and machine inputs. This fundamental unmappability fractures doctrines that assume traceable provenance -- authorship, inventorship, and liability -- yielding ownership gaps and moral "crumple zones." This Article argues that only functional equivalence stabilizes doctrine. Where provenance is indeterminate, legal frameworks must treat human and AI contributions as equivalent for allocating rights and responsibility -- not as a claim of moral or economic parity but as a pragmatic default. This principle stabilizes doctrine across domains, offering administrable rules: in copyright, vesting ownership in human orchestrators without parsing inseparable contributions; in patent, tying inventor-of-record status to human orchestration and reduction to practice, even when AI supplies the pivotal insight; and in tort, replacing intractable causation inquiries with enterprise-level and sector-specific strict or no-fault schemes. The contribution is both descriptive and normative: fluid agency explains why origin-based tests fail, while functional equivalence supplies an outcome-focused framework to allocate rights and responsibility when attribution collapses.

cs.CY

Beyond Pairwise Comparisons: A Distributional Test of Distinctiveness for Machine-Generated Works in Intellectual Property Law

Key doctrines, including novelty (patent), originality (copyright), and distinctiveness (trademark), turn on a shared empirical question: whether a body of work is meaningfully distinct from a relevant reference class. Yet analyses typically operationalize this set-level inquiry using item-level evidence: pairwise comparisons among exemplars. That unit-of-analysis mismatch may be manageable for finite corpora of human-created works, where it can be bridged by ad hoc aggregations. But it becomes acute for machine-generated works, where the object of evaluation is not a fixed set of works but a generative process with an effectively unbounded output space. We propose a distributional alternative: a two-sample test based on maximum mean discrepancy computed on semantic embeddings to determine if two creative processes-whether human or machine-produce statistically distinguishable output distributions. The test requires no task-specific training-obviating the need for discovery of proprietary training data to characterize the generative process-and is sample-efficient, often detecting differences with as few as 5-10 images and 7-20 texts. We validate the framework across three domains: handwritten digits (controlled images), patent abstracts (text), and AI-generated art (real-world images). We reveal a perceptual paradox: even when human evaluators distinguish AI outputs from human-created art with only about 58% accuracy, our method detects distributional distinctiveness. Our results present evidence contrary to the view that generative models act as mere regurgitators of training data. Rather than producing outputs statistically indistinguishable from a human baseline-as simple regurgitation would predict-they produce outputs that are semantically human-like yet stochastically distinct, suggesting their dominant function is as a semantic interpolator within a learned latent space.

cs.CY

Generic power laws in higher-dimensional lattice models with multidirectional hopping

We show that, on a $d-$dimensional hypercubic lattice with $d>1$, conserved-mass transport processes, with {\it multidirectional} hopping that respect all symmetries of the lattice, exhibit power-law correlations for generic parameter values $-$ even {\it far} from phase transition point, if any. The key idea for generating the algebraic decay is the notion of {\it multidirectional} hopping, which means that several chunks of masses, or several particles, can hop out simultaneously from a lattice site in multiple directions, consequently breaking detailed balance. Notably, the systems we consider are described by a continuous-time Markov process, are diffusive, {\it lattice-rotation symmetric}, spatially homogeneous and thus have {\it no} net mass current. Using hydrodynamic and exact microscopic theory, we show that, for spatial dimensions $d > 1$, the steady-state static density-density and ``activity''-density correlation functions in the thermodynamic limit typically decay as $\sim 1/r^{(d+2)}$ at large distance $r=|{\bf r}|$; the strength of the power law is exactly calculated for several models and expressed in terms of the density-dependent bulk-diffusion coefficient and Onsager matrix (or, mobility tensor). In particular, our theory explains why center-of-mass-conserving dynamics, used to model novel disordered {\it hyperuniform} state of matter, result in generic long-ranged correlations. However, in a restricted parameter regime, the correlations can also be short ranged and are characterized through the Onsager matrix.

cond-mat.stat-mech

Efficient Label Refinement for Face Parsing Under Extreme Poses Using 3D Gaussian Splatting

Accurate face parsing under extreme viewing angles remains a significant challenge due to limited labeled data in such poses. Manual annotation is costly and often impractical at scale. We propose a novel label refinement pipeline that leverages 3D Gaussian Splatting (3DGS) to generate accurate segmentation masks from noisy multiview predictions. By jointly fitting two 3DGS models, one to RGB images and one to their initial segmentation maps, our method enforces multiview consistency through shared geometry, enabling the synthesis of pose-diverse training data with only minimal post-processing. Fine-tuning a face parsing model on this refined dataset significantly improves accuracy on challenging head poses, while maintaining strong performance on standard views. Extensive experiments, including human evaluations, demonstrate that our approach achieves superior results compared to state-of-the-art methods, despite requiring no ground-truth 3D annotations and using only a small set of initial images. Our method offers a scalable and effective solution for improving face parsing robustness in real-world settings.

cs.CV

Tensor Factorized Hamiltonian Downfolding To Optimize The Scaling Complexity Of The Electronic Correlations Problem on Classical and Quantum Computers

Achieving chemical accuracy for strongly correlated molecules is a defining milestone for first-generation, fault-tolerant quantum computers, yet the factorial growth of three, four, and six-index tensor contractions in coupled-cluster CCSD(T), full configuration interaction (FCI), and multireference CI (MRCI) makes current classical and quantum approaches prohibitive. We introduce tensor-factorized Hamiltonian downfolding (TFHD) and its quantum analogue, qubitized downfolding (QD)- a hybrid classical-quantum framework that collapses every high-rank object to rank-2 networks executed in depth-optimal, block-encoded circuits. The complexity of these operations scales exponentially with the system size. We aim to find properties of chemical systems by optimizing this scaling through mathematical transformations on the Hamiltonian and the state space. By defining a bi-partition of the many-body Hilbert space into electronoccupied and electron-unoccupied blocks for a given orbital, we perform a downfolding transformation that decouples the electron-occupied block from its complement. We factorize high-rank electronic integrals and cluster amplitude tensors into low-rank tensor factors of a downfolding transformation, mapping the full many-body Hamiltonian into a smaller dimensional block-Hamiltonians. This reduces the computational complexity of solving the residual equations for Hamiltonian downfolding from O(N7) for CCSD(T) and O(N9) - O(N10) for CI and MRCI to O(N3). This operations can be implemented as a family of tensor networks solely made from two-rank tensors. Additionally, we create block-encoding quantum circuits of the tensor networks, generating circuits of O(N2) depth with O(logN) qubits. We demonstrate super-quadratic speedups of expensive quantum chemistry algorithms on both classical and quantum computers.

quant-ph

Charting the Parrot's Song: A Maximum Mean Discrepancy Approach to Measuring AI Novelty, Originality, and Distinctiveness

Current intellectual property frameworks struggle to evaluate the novelty of AI-generated content, relying on subjective assessments ill-suited for comparing effectively infinite AI outputs against prior art. This paper introduces a robust, quantitative methodology grounded in Maximum Mean Discrepancy (MMD) to measure distributional differences between generative processes. By comparing entire output distributions rather than conducting pairwise similarity checks, our approach directly contrasts creative processes--overcoming the computational challenges inherent in evaluating AI outputs against unbounded prior art corpora. Through experiments combining kernel mean embeddings with domain-specific machine learning representations (LeNet-5 for MNIST digits, CLIP for art), we demonstrate exceptional sensitivity: our method distinguishes MNIST digit classes with 95% confidence using just 5-6 samples and differentiates AI-generated art from human art in the AI-ArtBench dataset (n=400 per category; p<0.0001) using as few as 7-10 samples per distribution despite human evaluators' limited discrimination ability (58% accuracy). These findings challenge the "stochastic parrot" hypothesis by providing empirical evidence that AI systems produce outputs from semantically distinct distributions rather than merely replicating training data. Our approach bridges technical capabilities with legal doctrine, offering a pathway to modernize originality assessments while preserving intellectual property law's core objectives. This research provides courts and policymakers with a computationally efficient, legally relevant tool to quantify AI novelty--a critical advancement as AI blurs traditional authorship and inventorship boundaries.

cs.CY

Stochastic, Dynamic, Fluid Autonomy in Agentic AI: Implications for Authorship, Inventorship, and Liability

Agentic Artificial Intelligence (AI) systems, exemplified by OpenAI's DeepResearch, autonomously pursue goals, adapting strategies through implicit learning. Unlike traditional generative AI, which is reactive to user prompts, agentic AI proactively orchestrates complex workflows. It exhibits stochastic, dynamic, and fluid autonomy: its steps and outputs vary probabilistically (stochastic), it evolves based on prior interactions (dynamic), and it operates with significant independence within human-defined parameters, adapting to context (fluid). While this fosters complex, co-evolutionary human-machine interactions capable of generating uniquely synthesized creative outputs, it also irrevocably blurs boundaries--human and machine contributions become irreducibly entangled in intertwined creative processes. Consequently, agentic AI poses significant challenges to legal frameworks reliant on clear attribution: authorship doctrines struggle to disentangle ownership, intellectual property regimes strain to accommodate recursively blended novelty, and liability models falter as accountability diffuses across shifting loci of control. The central issue is not the legal treatment of human versus machine contributions, but the fundamental unmappability--the practical impossibility in many cases--of accurately attributing specific creative elements to either source. When retroactively parsing contributions becomes infeasible, applying distinct standards based on origin becomes impracticable. Therefore, we argue, legal and policy frameworks may need to treat human and machine contributions as functionally equivalent--not for moral or economic reasons, but as a pragmatic necessity.

cs.CY

Hyperuniformity in mass transport processes with center-of-mass conservation: Some exact results

We characterize steady-state static and dynamic properties in a broad class of mass transport processes on a periodic hypercubic lattice of volume $L^d$, where both mass and {\it center-of-mass} (CoM) remain conserved and detailed balance is violated in the bulk; we specifically consider these models in $d=1$ and $2$ dimensions. Using a microscopic approach, we exactly determine the decay (or, growth) exponents for various dynamic and static correlation functions. We show that, despite constrained dynamics due to the CoM conservation (CoMC), the density relaxation is indeed diffusive. However, fluctuation properties are strikingly different from that in the diffusive systems with a single (mass) conservation law. In the thermodynamic limit, the steady-state variance $\langle {\cal Q}^2(T) \rangle_c$ of time-integrated bond current ${\cal Q}(T)$ across a bond in time interval $T$ exhibits the following long-time behavior: $\langle {\cal Q}^2(T) \rangle_c \simeq A_1 T + A_2 + A_3 T^{-d/2}$. Remarkably, depending on dimensions and microscopic details, the prefactor $A_1$ can vanish (e.g., for $d=1$), causing the variance to eventually {\it saturate}. The exponents governing the small-frequency behavior of the power spectrum $S_J(f) \sim f^{ψ_J}$ for bond current are exactly determined as $ψ_J=3/2$ and $2$ in $d=1$ and $2$ dimensions, respectively, implying a ``dynamic hyperuniformity''. We also compute the static structure factor $S(q)$, which, in the small-$q$ limit, varies as the square of wave number $q$, i.e., $S(q) \sim q^2$. Indeed, both dynamic and static fluctuations are anomalously suppressed, resulting in an extreme form of (``class I'') hyperuniformity in the systems.

cond-mat.stat-mech

Agentic AI: Autonomy, Accountability, and the Algorithmic Society

Agentic Artificial Intelligence (AI) can autonomously pursue long-term goals, make decisions, and execute complex, multi-turn workflows. Unlike traditional generative AI, which responds reactively to prompts, agentic AI proactively orchestrates processes, such as autonomously managing complex tasks or making real-time decisions. This transition from advisory roles to proactive execution challenges established legal, economic, and creative frameworks. In this paper, we explore challenges in three interrelated domains: creativity and intellectual property, legal and ethical considerations, and competitive effects. Central to our analysis is the tension between novelty and usefulness in AI-generated creative outputs, as well as the intellectual property and authorship challenges arising from AI autonomy. We highlight gaps in responsibility attribution and liability that create a "moral crumple zone"--a condition where accountability is diffused across multiple actors, leaving end-users and developers in precarious legal and ethical positions. We examine the competitive dynamics of two-sided algorithmic markets, where both sellers and buyers deploy AI agents, potentially mitigating or amplifying tacit collusion risks. We explore the potential for emergent self-regulation within networks of agentic AI--the development of an "algorithmic society"--raising critical questions: To what extent would these norms align with societal values? What unintended consequences might arise? How can transparency and accountability be ensured? Addressing these challenges will necessitate interdisciplinary collaboration to redefine legal accountability, align AI-driven choices with stakeholder values, and maintain ethical safeguards. We advocate for frameworks that balance autonomy with accountability, ensuring all parties can harness agentic AI's potential while preserving trust, fairness, & societal welfare.

cs.CY

SemUV: Deep Learning based semantic manipulation over UV texture map of virtual human heads

Designing and manipulating virtual human heads is essential across various applications, including AR, VR, gaming, human-computer interaction and VFX. Traditional graphic-based approaches require manual effort and resources to achieve accurate representation of human heads. While modern deep learning techniques can generate and edit highly photorealistic images of faces, their focus remains predominantly on 2D facial images. This limitation makes them less suitable for 3D applications. Recognizing the vital role of editing within the UV texture space as a key component in the 3D graphics pipeline, our work focuses on this aspect to benefit graphic designers by providing enhanced control and precision in appearance manipulation. Research on existing methods within the UV texture space is limited, complex, and poses challenges. In this paper, we introduce SemUV: a simple and effective approach using the FFHQ-UV dataset for semantic manipulation directly within the UV texture space. We train a StyleGAN model on the publicly available FFHQ-UV dataset, and subsequently train a boundary for interpolation and semantic feature manipulation. Through experiments comparing our method with 2D manipulation technique, we demonstrate its superior ability to preserve identity while effectively modifying semantic features such as age, gender, and facial hair. Our approach is simple, agnostic to other 3D components such as structure, lighting, and rendering, and also enables seamless integration into standard 3D graphics pipelines without demanding extensive domain expertise, time, or resources.

cs.CV

OMuSense-23: A Multimodal Dataset for Contactless Breathing Pattern Recognition and Biometric Analysis

In the domain of non-contact biometrics and human activity recognition, the lack of a versatile, multimodal dataset poses a significant bottleneck. To address this, we introduce the Oulu Multi Sensing (OMuSense-23) dataset that includes biosignals obtained from a mmWave radar, and an RGB-D camera. The dataset features data from 50 individuals in three distinct poses -- standing, sitting, and lying down -- each featuring four specific breathing pattern activities: regular breathing, reading, guided breathing, and apnea, encompassing both typical situations (e.g., sitting with normal breathing) and critical conditions (e.g., lying down without breathing). In our work, we present a detailed overview of the OMuSense-23 dataset, detailing the data acquisition protocol, describing the process for each participant. In addition, we provide, a baseline evaluation of several data analysis tasks related to biometrics, breathing pattern recognition and pose identification. Our results achieve a pose identification accuracy of 87\% and breathing pattern activity recognition of 83\% using features extracted from biosignals. The OMuSense-23 dataset is publicly available as resource for other researchers and practitioners in the field.

cs.CV

CAVIAR: Categorical-Variable Embeddings for Accurate and Robust Inference

Social science research often hinges on the relationship between categorical variables and outcomes. We introduce CAVIAR, a novel method for embedding categorical variables that assume values in a high-dimensional ambient space but are sampled from an underlying manifold. Our theoretical and numerical analyses outline challenges posed by such categorical variables in causal inference. Specifically, dynamically varying and sparse levels can lead to violations of the Donsker conditions and a failure of the estimation functionals to converge to a tight Gaussian process. Traditional approaches, including the exclusion of rare categorical levels and principled variable selection models like LASSO, fall short. CAVIAR embeds the data into a lower-dimensional global coordinate system. The mapping can be derived from both structured and unstructured data, and ensures stable and robust estimates through dimensionality reduction. In a dataset of direct-to-consumer apparel sales, we illustrate how high-dimensional categorical variables, such as zip codes, can be succinctly represented, facilitating inference and analysis.

econ.EM

AI Knowledge and Reasoning: Emulating Expert Creativity in Scientific Research

We investigate whether modern AI can emulate expert creativity in complex scientific endeavors. We introduce novel methodology that utilizes original research articles published after the AI's training cutoff, ensuring no prior exposure, mitigating concerns of rote memorization and prior training. The AI are tasked with redacting findings, predicting outcomes from redacted research, and assessing prediction accuracy against reported results. Analysis on 589 published studies in four leading psychology journals over a 28-month period, showcase the AI's proficiency in understanding specialized research, deductive reasoning, and evaluating evidentiary alignment--cognitive hallmarks of human subject matter expertise and creativity. These findings suggest the potential of general-purpose AI to transform academia, with roles requiring knowledge-based creativity become increasingly susceptible to technological substitution.

cs.AI