SearcharxivSearch

arXiv subjects

Soumyadip Banerjee

Publications and source records attributed to Soumyadip Banerjee.

7 recordsLinked to original sources

CASTLE: Contrastive and Seed-Guided Training for Cold-Start Natural Language Search

Deploying natural language search systems presents a critical cold-start challenge: no real user queries to learn linguistic patterns, and no relevance labels to train ranking models. We present CASTLE (Contrastive And Seed-guided Training for natural Language sEarch), an LLM-based framework for generating synthetic queries and relevance labels from structured catalog data, powering Airbnb's natural language search across its full lifecycle. CASTLE makes three contributions. First, we generate realistic queries by combining structure-guided prompting with seed queries from user research, using template, few-shot, and attribute-grounded prompt variants together with explicit variety mechanisms to prevent query collapse. Second, we produce relevance labels by construction via contrastive listing pairs derived from booking sessions, achieving near-zero false positives without LLM judgment. Third, CASTLE's structured input design is flexible: incorporating richer signals such as guest reviews and photo captions alongside listing attributes enables generation of niche, long-tail queries that reflect subjective user preferences (e.g., "cozy cabin with fireplace") beyond what catalog attributes alone can express. Compared against InPars-style, Promptagator, and contrastive-only baselines, CASTLE achieves KL 1.01 vs. real users -- a 9.2x improvement over the best baseline (9.33) -- and the lowest attribute-type KL divergence (0.08), outperforming even survey seed queries (0.09). A human evaluation on 200 sampled triplets confirms label quality: annotators agree with CASTLE labels at 91-93%. We deploy production pipelines generating synthetic examples daily for embedding-based retrieval and ranking evaluation. Synthetic data remains valuable beyond cold-start: it targets tail queries underrepresented in organic traffic and extends naturally to multi-turn conversational search.

cs.IR

High Precision Audience Expansion via Extreme Classification in a Two-Sided Marketplace

Airbnb search must balance a worldwide, highly varied supply of homes with guests whose location, amenity, style, and price expectations differ widely. Meeting those expectations hinges on an efficient retrieval stage that surfaces only the listings a guest might realistically book, before resource intensive ranking models are applied to determine the best results. Unlike many recommendation engines, our system faces a distinctive challenge, location retrieval, that sits upstream of ranking and determines which geographic areas are queried in order to filter inventory to a candidate set. The preexisting approach employs a deep bayesian bandit based system to predict a rectangular retrieval bounds area that can be used for filtering. The purpose of this paper is to demonstrate the methodology, challenges, and impact of rearchitecting search to retrieve from the subset of most bookable high precision rectangular map cells defined by dividing the world into 25M uniform cells.

cs.IR

Applying Embedding-Based Retrieval to Airbnb Search

The goal of Airbnb search is to match guests with the ideal accommodation that fits their travel needs. This is a challenging problem, as popular search locations can have around a hundred thousand available homes, and guests themselves have a wide variety of preferences. Furthermore, the launch of new product features, such as \textit{flexible date search,} significantly increased the number of eligible homes per search query. As such, there is a need for a sophisticated retrieval system which can provide high-quality candidates with low latency in a way that integrates with the overall ranking stack. This paper details our journey to build an efficient and high-quality retrieval system for Airbnb search. We describe the key unique challenges we encountered when implementing an Embedding-Based Retrieval (EBR) system for a two sided marketplace like Airbnb -- such as the dynamic nature of the inventory, a lengthy user funnel with multiple stages, and a variety of product surfaces. We cover unique insights when modeling the retrieval problem, how to build robust evaluation systems, and design choices for online serving. The EBR system was launched to production and powers several use-cases such as regular search, flexible date and promotional emails for marketing campaigns. The system demonstrated statistically-significant improvements in key metrics, such as booking conversion, via A/B testing.

cs.IR

Learning to Rank for Maps at Airbnb

As a two-sided marketplace, Airbnb brings together hosts who own listings for rent with prospective guests from around the globe. Results from a guest's search for listings are displayed primarily through two interfaces: (1) as a list of rectangular cards that contain on them the listing image, price, rating, and other details, referred to as list-results (2) as oval pins on a map showing the listing price, called map-results. Both these interfaces, since their inception, have used the same ranking algorithm that orders listings by their booking probabilities and selects the top listings for display. But some of the basic assumptions underlying ranking, built for a world where search results are presented as lists, simply break down for maps. This paper describes how we rebuilt ranking for maps by revising the mathematical foundations of how users interact with search results. Our iterative and experiment-driven approach led us through a path full of twists and turns, ending in a unified theory for the two interfaces. Our journey shows how assumptions taken for granted when designing machine learning algorithms may not apply equally across all user interfaces, and how they can be adapted. The net impact was one of the largest improvements in user experience for Airbnb which we discuss as a series of experimental validations.

cs.IR

ALMA detection of water vapour in the low mass protostar IRAS 16293$-$2422

The low mass protostar IRAS 16293$-$2422 is a well-known young stellar system that is observed in the L1689N molecular cloud in the constellation of Ophiuchus. In the interstellar medium and solar system bodies, water is a necessary species for the formation of life. We present the spectroscopic detection of the rotational emission line of water (H$_{2}$O) vapour from the low mass protostar IRAS 16293$-$2422 using the Atacama Large Millimeter/submillimeter Array (ALMA) band 5 observation. The emission line of H$_{2}$O is detected at frequency $ν$ = 183.310 GHz with transition J=3$_{1,3}$$-$2$_{2,2}$. The statistical column density of the emission line of water vapour is $N$(H$_{2}$O) = 4.2$\times$10$^{16}$ cm$^{-2}$ with excitation temperature ($T_{ex}$) = 124$\pm$10 K. The fractional abundance of H$_{2}$O with respect to H$_{2}$ is 1.44$\times$10$^{-7}$ where $N$(H$_{2}$) = 2.9$\times$10$^{23}$ cm$^{-2}$.

astro-ph.GA

On the length scale dependence of DNA conformational change under local perturbation

Conformational change of a DNA molecule is frequently observed in multiple biological processes and has been modelled using a chain of strongly coupled oscillators with a nonlinear bistable potential. While the mechanism and properties of conformational change in the model have been investigated and several reduced order models developed, the conformational dynamics as a function of the length of the oscillator chain is relatively less clear. To address this, we used a modified Lindstedt-Poincare method and numerical computations. We calculate a perturbation expansion of the frequency of the model's nonzero modes, finding that approximating these modes with their unperturbed dynamics, as in a previous reduced order model, may not hold when the length of the DNA model increases. We investigate the conformational change to local perturbation in models of varying lengths, finding that for chosen input and parameters, there are two regions of DNA length in the model, first where the minimum energy required to undergo the conformational change increases with DNA length; and second, where it is almost independent of the length of the DNA model. We analyze the conformational change in these models by adding randomness to the local perturbation, finding that the tendency of the system to remain in a stable conformation against random perturbation decreases with an increase in the DNA length. These results should help to understand the role of the length of a DNA molecule in influencing its conformational dynamics.

physics.bio-ph

Robustness of a Biomolecular Oscillator to Pulse Perturbations

Biomolecular oscillators can function robustly in the presence of environmental perturbations, which can either be static or dynamic. While the effect of different circuit parameters and mechanisms on the robustness to steady perturbations has been investigated, the scenario for dynamic perturbations is relatively unclear. To address this we use a benchmark three protein oscillator design - the repressilator - and investigate its robustness to pulse perturbations, computationally as well as using analytical tools of Floquet theory. We find that the metric provided by direct computations of the time it takes for the oscillator to settle after a pulse perturbation is applied, correlates well with the metric provided by Floquet theory. We investigate the parametric dependence of the Floquet metric, finding that the parameters that increase the effective delay enhance robustness to pulse perturbation. We find that the structural changes such as increasing the number of proteins in a ring oscillator as well as adding positive feedback, both of which increase effective delay, facilitates such robustness. These results highlight such design principles, especially the role of delay, for designing an oscillator that is robust to pulse perturbation.

q-bio.MN