SearcharxivSearch

arXiv subjects

Sameer Jain

Publications and source records attributed to Sameer Jain.

7 recordsLinked to original sources

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention. However, directly optimizing retention is difficult because return signals are sparse, delayed, and only partially attributable to earlier recommendations. Prior work has addressed this challenge with sequential modeling and reinforcement learning, but these approaches typically require task specific reward engineering, substantial computational overhead, and surface specific implementations that are difficult to generalize. In this paper, we present a unified, model-agnostic downstream reward framework for optimizing long-term user value in large-scale recommendation systems. First, we formulate the downstream reward learning problem and develop an offline screening framework to identify session level behaviors that are both observable early and predictive of future retention. We then propose several model-agnostic downstream rewards signals derived from observed user action patterns across multiple sources. We further discuss the engineering effort to productionize the proposed rewards derivations and challenges we faced when adding them to our ranking models. Online A/B experiments demonstrate consistent improvements in engagement and retention-related metrics, and the framework has been deployed across multiple Pinterest surfaces, including Homefeed, Related Pins, Search, and Notifications.

cs.LG

Study of the subleading twist GTMD $E_{21}$ for proton in light-front quark-diquark model

In this work, we study the generalized transverse momentum dependent distribution (GTMD) $E_{21}$ for proton using light-front quark-diquark model. We construct the expression of $E_{21}$ GTMD using the overlap equation in light-front wave functions obtained from the GTMD correlator with Dirac matrix structure $\Gamma=1$, in both situations of scalar and vector diquark. The $3$-dimensional plots of GTMD $E_{21}$ have been analyzed with respect to its variables by taking two variables at a time while holding others constant.

hep-ph

Unraveling sub-leading twist GTMDs of proton using LFQDM

This study investigates the sub-leading twist generalized transverse momentum dependent distributions (GTMDs) of a proton within the light-front quark-diquark model (LFQDM) framework. We solve the parametrization equations for the Dirac matrix structure to yield the explicit GTMD expressions for scalar as well as vector diquark configurations for active $u$ and $d$ quarks. This analysis addresses the multi-dimensional nature of GTMDs by exploring their dependencies on one or two variables while keeping others fixed. Additionally, we extract transverse momentum dependent form factors (TMFFs) from GTMDs by integrating over the longitudinal momentum fraction $x$. The study uses $3$-dimensional plots to illustrate the variation of TMFFs with the quark's transverse momentum $p_{\perp}$ and the transverse momentum transfer to the proton $\Delta_{\perp}$.

hep-ph

Deciphering Twist-3 Chiral-Even GPDs in the Light-Front Quark-Diquark Model

We investigate quantum chromodynamics (QCD) in this study by computing chiral-even generalized parton distributions (GPDs) at twist-$3$ using the light-front quark-diquark model (LFQDM), particularly when the longitudinal momentum transfer is zero. We provide a detailed analysis of the twist-$3$ chiral-even GPD's dependence on the longitudinal momentum fraction ($x$) and the momentum transfer ($t$) by illustrating their behavior through extensive two-dimensional ($2$-D) and three-dimensional ($3$-D) visualizations. Our investigation also reveals the intricate relationships between these GPDs and other distribution functions (DFs) such as generalized transverse-momentum dependent distributions (GTMDs), transverse momentum-dependent parton distributions (TMDs), and parton distribution functions (PDFs). Our study also includes the connected form factors (FFs) which are crucial in understanding the internal structure of hadrons. Additionally, we provide impact parameter GPD plots to offer insights into the spatial distribution of partons.

hep-ph

Where It Really Matters: Few-Shot Environmental Conservation Media Monitoring for Low-Resource Languages

Environmental conservation organizations routinely monitor news content on conservation in protected areas to maintain situational awareness of developments that can have an environmental impact. Existing automated media monitoring systems require large amounts of data labeled by domain experts, which is only feasible at scale for high-resource languages like English. However, such tools are most needed in the global south where news of interest is mainly in local low-resource languages, and far fewer experts are available to annotate datasets sustainably. In this paper, we propose NewsSerow, a method to automatically recognize environmental conservation content in low-resource languages. NewsSerow is a pipeline of summarization, in-context few-shot classification, and self-reflection using large language models (LLMs). Using at most 10 demonstration example news articles in Nepali, NewsSerow significantly outperforms other few-shot methods and achieves comparable performance with models fully fine-tuned using thousands of examples. The World Wide Fund for Nature (WWF) has deployed NewsSerow for media monitoring in Nepal, significantly reducing their operational burden, and ensuring that AI tools for conservation actually reach the communities that need them the most. NewsSerow has also been deployed for countries with other languages like Colombia.

cs.CL

Multi-Dimensional Evaluation of Text Summarization with In-Context Learning

Evaluation of natural language generation (NLG) is complex and multi-dimensional. Generated text can be evaluated for fluency, coherence, factuality, or any other dimensions of interest. Most frameworks that perform such multi-dimensional evaluation require training on large manually or synthetically generated datasets. In this paper, we study the efficacy of large language models as multi-dimensional evaluators using in-context learning, obviating the need for large training datasets. Our experiments show that in-context learning-based evaluators are competitive with learned evaluation frameworks for the task of text summarization, establishing state-of-the-art on dimensions such as relevance and factual consistency. We then analyze the effects of factors such as the selection and number of in-context examples on performance. Finally, we study the efficacy of in-context learning based evaluators in evaluating zero-shot summaries written by large language models such as GPT-3.

cs.CL

Four Years of FAccT: A Reflexive, Mixed-Methods Analysis of Research Contributions, Shortcomings, and Future Prospects

Fairness, Accountability, and Transparency (FAccT) for socio-technical systems has been a thriving area of research in recent years. An ACM conference bearing the same name has been the central venue for scholars in this area to come together, provide peer feedback to one another, and publish their work. This reflexive study aims to shed light on FAccT's activities to date and identify major gaps and opportunities for translating contributions into broader positive impact. To this end, we utilize a mixed-methods research design. On the qualitative front, we develop a protocol for reviewing and coding prior FAccT papers, tracing their distribution of topics, methods, datasets, and disciplinary roots. We also design and administer a questionnaire to reflect the voices of FAccT community members and affiliates on a wide range of topics. On the quantitative front, we use the full text and citation network associated with prior FAccT publications to provide further evidence about topics and values represented in FAccT. We organize the findings from our analysis into four main dimensions: the themes present in FAccT scholarship, the values that underpin the work, the impact of the contributions both within academic circles and beyond, and the practices and informal norms of the community that has formed around FAccT. Finally, our work identifies several suggestions on directions for change, as voiced by community members.

cs.CY