SearcharxivSearch

arXiv subjects

Victor Chen

Publications and source records attributed to Victor Chen.

17 recordsLinked to original sources

A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction

This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that prioritizes privacy over distributional fidelity. Our findings confirm that RDS-based PPML is technically feasible for educational settings and offers a practical alternative to federated learning for small-scale inter-institutional collaborations. The code is available at https://github.com/jtfields/NAIRR240195-Privacy-Preserving-Machine-Learning.

cs.CR

Memento: Personalized RAG-Style Long-Retention Data Scaling for META Ads Recommendation

Modeling of long history data suffers from long-context window attention dilution, system efficiency and catastrophic forgetting problems, where naive linear scaling approach like LastN would fail. We introduce Memento, a personalized retrieval-augmented framework that treats historical user engagements as a document corpus and ad requests as queries, retrieving relevant interactions via Maximal Marginal Relevance (MMR) to balance similarity with diversity. We identify two complementary applications: Representation Memento, which retrieves historical embeddings for feature augmentation, and Data Memento, which retrieves past training examples for multipass training. Through infrastructure co-design -- temporal chunking, INT8 quantization, and asynchronous serving -- Memento achieves 5-10$\times$ resource efficiency over linear scaling. Memento processes daily requests with sub-10ms latency, yielding 0.25-0.3% Normalized Entropy gain on both click-through and conversion prediction. In production, Memento delivers a 1% CTR lift on Facebook Feed and Reels and a 1.2% CVR lift, scaling personalization to 365+ days of history.

cs.IR

&inator: Correct, Precise C-to-Rust Interface Translation

Automatically translating system software from C to Rust is an appealing but challenging problem, as it requires whole-program reasoning to satisfy Rust's ownership and borrowing discipline. A key enabling step in whole-program translation is interface translation, which produces Rust declarations for the C program's top-level declarations (i.e., structs and function signatures), enabling modular and incremental code translation. This paper introduces correct, precise C-to-Rust interface translation, called &inator. &inator employs a novel constraint-based formulation of semantic equivalence and type correctness including borrow-checking rules to produce a Rust interface that is correct (i.e., the interface admits a semantics-preserving implementation in safe Rust) and precise (i.e., it uses the simplest, least costly types). Our results show &inator produces correct, precise Rust interfaces for real C programs, but support for certain C features and scaling to large programs are challenges left for future work. This work advances the state of the art by being the first correct, precise approach to C-to-Rust interface translation.

cs.PL

Efficient Auxiliary-Field Quantum Monte Carlo using Isometric Tensor Hypercontraction

Auxiliary Field Quantum Monte Carlo (AFQMC) has emerged as a powerful framework for treating strongly correlated electronic systems, offering a favorable balance between computational cost and accuracy. In this paper, we present a novel AFQMC method that uses the isometric tensor hypercontraction (ITHC) technique to diagonalize the two-body Coulomb interaction of molecular electronic Hamiltonians by introducing additional fictitious fermionic modes. Our method shows reduced theoretical complexity and better practical performance for both propagation and local energy evaluation compared to the standard AFQMC method. We demonstrate the efficacy of this approach by computing the ground-state energies of a linear $\ce{H10}$-chain and the benzene molecule. Our results show that the extended-basis AFQMC recovers many-body correlations with a precision comparable to that of high-level wavefunction methods such as Coupled Clusters (CC) or Density Matrix Renormalization Group (DMRG), while offering significantly improved scaling.

physics.chem-ph

Computation of the multiplicities of zigzags

In this note, we explore various cohomological invariants on double complexes with the aim of finding their decomposition into irreducible parts, which are of square and zigzag shape. By studying the growth rate of the number of invariants given by the multiplicities of zigzags in the double complex of an n-dimensional complex manifold, we show that the De Rham, Dolbeault, Bott-Chern, Aeppli, and Varouchas cohomologies do not suffice to distinguish non-isomorphic double complexes. We also describe the zigzags counted by the Bigolin cohomology, and show how their dimensions are related to the multiplicities of odd zigzags. A special class of complex manifolds is given by the nilmanifolds. For a nilmanifold, the double complex of left-invariant forms is quasi-isomorphic to the double complex of differential forms. In dimension 6, we compute the double complex of forms of two nilmanifolds having the same Betti, Hodge and Bott-Chern numbers, but whose double complexes are non-isomorphic. We also compute the double complexes of a subclass of almost abelian nilmanifolds, which exist in any dimension.

math.DG

Efficient enumeration of quadratic lattices

We present an algorithm to enumerate isometry classes of integral quadratic lattices of a given rank and determinant, and analyze its running time by giving bounds on the number of genus symbols for a fixed rank and determinant. We build on previous work of Kirschmer, Brandhorst, Hanke, and Dubey and Holenstein. We analyze the running times of their respective algorithms and compare the practical performance of their implementations with our own. Our implementations are publicly available.

math.NT

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.

cs.LG

Global sections of the positively twisted Green-Griffiths bundles

With various jet orders $k$ and weights $n$, let $E_{k,n}^{\rm GG}$ be the Green-Griffiths bundles over the projective space $\mathbb{P}^N (\mathbb{C})$. Denote by $\mathcal{O} (d)$ the tautological line bundle over $\mathbb{P}^N (\mathbb{C})$. Although only negative twists are of interest for applications to complex hyperbolicity (above general type projective submanifolds $Y \subset \mathbb{P}^N (\mathbb{C})$), it is known that the positive twists $E_{k,n}^{\rm GG} \otimes \mathcal{O} (d)$ enjoy nontrivial global sections. In this article, we establish that for every $d \geqslant 1$ and for every jet order $k \geqslant d-1$: \[ \dim\, H^0 \bigg( \mathbb{P}^N,\,\, \bigoplus_{n=1}^{\infty} E_{k, n}^{\text{GG}} \otimes \mathcal{O}(d) \bigg) = (N+1)^d. \] This theorem is actually a corollary of a recent work of Etesse, devoted to a proof, from the point of view of differentially homogeneous polynomials, of the so-called Schmidt-Kolchin-Reinhart conjecture, by means of (advanced) Representation Theory. As Etesse discovered a (simple) tight link with the Green-Griffiths bundles, both statements are in fact equivalent. Our objective is to set up an alternative proof of the above precise dimension estimate, from the Green-Griffiths point of view (only). More precisely, we find an explicit description of all concerned global sections. Our arguments are elementary, and use only determinants, linear algebra, monomial orderings. One old hope is to discover some explicit formulas for global sections of negatively twisted Green-Griffiths bundles over projective general type submanifolds $Y \subset \mathbb{P}^N (\mathbb{C})$, a problem still open.

math.AG

A New Era: Intelligent Tutoring Systems Will Transform Online Learning for Millions

Despite artificial intelligence (AI) having transformed major aspects of our society, less than a fraction of its potential has been explored, let alone deployed, for education. AI-powered learning can provide millions of learners with a highly personalized, active and practical learning experience, which is key to successful learning. This is especially relevant in the context of online learning platforms. In this paper, we present the results of a comparative head-to-head study on learning outcomes for two popular online learning platforms (n=199 participants): A MOOC platform following a traditional model delivering content using lecture videos and multiple-choice quizzes, and the Korbit learning platform providing a highly personalized, active and practical learning experience. We observe a huge and statistically significant increase in the learning outcomes, with students on the Korbit platform providing full feedback resulting in higher course completion rates and achieving learning gains 2 to 2.5 times higher than both students on the MOOC platform and students in a control group who don't receive personalized feedback on the Korbit platform. The results demonstrate the tremendous impact that can be achieved with a personalized, active learning AI-powered system. Making this technology and learning experience available to millions of learners around the world will represent a significant leap forward towards the democratization of education.

cs.CY

Computer-assisted construct classification of organizational performance concerning different stakeholder groups

The number of research articles in business and management has dramatically increased along with terminology, constructs, and measures. Proper classification of organizational performance constructs from research articles plays an important role in categorizing the literature and understanding to whom its research implications may be relevant. In this work, we classify constructs (i.e., concepts and terminology used to capture different aspects of organizational performance) in research articles into a three-level categorization: (a) performance and non-performance categories (Level 0); (b) for performance constructs, stakeholder group-level of performance concerning investors, customers, employees, and the society (community and natural environment) (Level 1); and (c) for each stakeholder group-level, subcategories of different ways of measurement (Level 2). We observed that increasing contextual information with features extracted from surrounding sentences and external references improves classification of disaggregate-level labels, given limited training data. Our research has implications for computer-assisted construct identification and classification - an essential step for research synthesis.

cs.CL

Understanding the Design Space of Mouth Microgestures

As wearable devices move toward the face (i.e. smart earbuds, glasses), there is an increasing need to facilitate intuitive interactions with these devices. Current sensing techniques can already detect many mouth-based gestures; however, users' preferences of these gestures are not fully understood. In this paper, we investigate the design space and usability of mouth-based microgestures. We first conducted brainstorming sessions (N=16) and compiled an extensive set of 86 user-defined gestures. Then, with an online survey (N=50), we assessed the physical and mental demand of our gesture set and identified a subset of 14 gestures that can be performed easily and naturally. Finally, we conducted a remote Wizard-of-Oz usability study (N=11) mapping gestures to various daily smartphone operations under a sitting and walking context. From these studies, we develop a taxonomy for mouth gestures, finalize a practical gesture set for common applications, and provide design guidelines for future mouth-based gesture interactions.

cs.HC

My Team Will Go On: Differentiating High and Low Viability Teams through Team Interaction

Understanding team viability -- a team's capacity for sustained and future success -- is essential for building effective teams. In this study, we aggregate features drawn from the organizational behavior literature to train a viability classification model over a dataset of 669 10-minute text conversations of online teams. We train classifiers to identify teams at the top decile (most viable teams), 50th percentile (above a median split), and bottom decile (least viable teams), then characterize the attributes of teams at each of these viability levels. We find that a lasso regression model achieves an accuracy of .74--.92 AUC ROC under different thresholds of classifying viability scores. From these models, we identify the use of exclusive language such as `but' and `except', and the use of second person pronouns, as the most predictive features for detecting the most viable teams, suggesting that active engagement with others' ideas is a crucial signal of a viable team. Only a small fraction of the 10-minute discussion, as little as 70 seconds, is required for predicting the viability of team interaction. This work suggests opportunities for teams to assess, track, and visualize their own viability in real time as they collaborate.

cs.CY

Dimension-Robust MCMC in Bayesian Inverse Problems

The methodology developed in this article is motivated by a wide range of prediction and uncertainty quantification problems that arise in Statistics, Machine Learning and Applied Mathematics, such as non-parametric regression, multi-class classification and inversion of partial differential equations. One popular formulation of such problems is as Bayesian inverse problems, where a prior distribution is used to regularize inference on a high-dimensional latent state, typically a function or a field. It is common that such priors are non-Gaussian, for example piecewise-constant or heavy-tailed, and/or hierarchical, in the sense of involving a further set of low-dimensional parameters, which, for example, control the scale or smoothness of the latent state. In this formulation prediction and uncertainty quantification relies on efficient exploration of the posterior distribution of latent states and parameters. This article introduces a framework for efficient MCMC sampling in Bayesian inverse problems that capitalizes upon two fundamental ideas in MCMC, non-centred parameterisations of hierarchical models and dimension-robust samplers for latent Gaussian processes. Using a range of diverse applications we showcase that the proposed framework is dimension-robust, that is, the efficiency of the MCMC sampling does not deteriorate as the dimension of the latent state gets higher. We showcase the full potential of the machinery we develop in the article in semi-supervised multi-class classification, where our sampling algorithm is used within an active learning framework to guide the selection of input data to manually label in order to achieve high predictive accuracy with a minimal number of labelled data.

stat.ME

Property Testing via Set-Theoretic Operations

Given two testable properties $\mathcal{P}_{1}$ and $\mathcal{P}_{2}$, under what conditions are the union, intersection or set-difference of these two properties also testable? We initiate a systematic study of these basic set-theoretic operations in the context of property testing. As an application, we give a conceptually different proof that linearity is testable, albeit with much worse query complexity. Furthermore, for the problem of testing disjunction of linear functions, which was previously known to be one-sided testable with a super-polynomial query complexity, we give an improved analysis and show it has query complexity $O(1/\eps^2)$, where $\eps$ is the distance parameter.

cs.DS

Efficient and Error-Correcting Data Structures for Membership and Polynomial Evaluation

We construct efficient data structures that are resilient against a constant fraction of adversarial noise. Our model requires that the decoder answers most queries correctly with high probability and for the remaining queries, the decoder with high probability either answers correctly or declares "don't know." Furthermore, if there is no noise on the data structure, it answers all queries correctly with high probability. Our model is the common generalization of a model proposed recently by de Wolf and the notion of "relaxed locally decodable codes" developed in the PCP literature. We measure the efficiency of a data structure in terms of its length, measured by the number of bits in its representation, and query-answering time, measured by the number of bit-probes to the (possibly corrupted) representation. In this work, we study two data structure problems: membership and polynomial evaluation. We show that these two problems have constructions that are simultaneously efficient and error-correcting.

cs.DS

Testing Linear-Invariant Non-Linear Properties

We consider the task of testing properties of Boolean functions that are invariant under linear transformations of the Boolean cube. Previous work in property testing, including the linearity test and the test for Reed-Muller codes, has mostly focused on such tasks for linear properties. The one exception is a test due to Green for "triangle freeness": a function $f:\cube^{n}\to\cube$ satisfies this property if $f(x),f(y),f(x+y)$ do not all equal 1, for any pair $x,y\in\cube^{n}$. Here we extend this test to a more systematic study of testing for linear-invariant non-linear properties. We consider properties that are described by a single forbidden pattern (and its linear transformations), i.e., a property is given by $k$ points $v_{1},...,v_{k}\in\cube^{k}$ and $f:\cube^{n}\to\cube$ satisfies the property that if for all linear maps $L:\cube^{k}\to\cube^{n}$ it is the case that $f(L(v_{1})),...,f(L(v_{k}))$ do not all equal 1. We show that this property is testable if the underlying matroid specified by $v_{1},...,v_{k}$ is a graphic matroid. This extends Green's result to an infinite class of new properties. Our techniques extend those of Green and in particular we establish a link between the notion of "1-complexity linear systems" of Green and Tao, and graphic matroids, to derive the results.

math.CO

A Hypergraph Dictatorship Test with Perfect Completeness

A hypergraph dictatorship test is first introduced by Samorodnitsky and Trevisan and serves as a key component in their unique games based $\PCP$ construction. Such a test has oracle access to a collection of functions and determines whether all the functions are the same dictatorship, or all their low degree influences are $o(1).$ Their test makes $q\geq3$ queries and has amortized query complexity $1+O(\frac{\log q}{q})$ but has an inherent loss of perfect completeness. In this paper we give an adaptive hypergraph dictatorship test that achieves both perfect completeness and amortized query complexity $1+O(\frac{\log q}{q})$.

math.CO