Searcharxiv⌕ Search

arXiv subjects

Vladimir Vlasov

Publications and source records attributed to Vladimir Vlasov.

16 recordsLinked to original sources

Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation

Modern sequential recommender systems commonly use transformer-based models for next-item prediction. While these models demonstrate a strong balance between efficiency and quality, integrating interleaving features - such as the query context (e.g., browse category) under which next-item interactions occur - poses challenges. Effectively capturing query context is crucial for refining ranking relevance and enhancing user engagement, as it provides valuable signals about user intent within a session. Unlike item features, historical query context is typically not aligned with item sequences and may be unavailable at inference due to privacy constraints or feature store limitations - making its integration into transformers both challenging and error-prone. This paper analyzes different strategies for incorporating query context into transformers trained with a causal language modeling procedure as a case study. We propose a new method that effectively fuses the item sequence with query context within the attention mechanism. Through extensive offline and online experiments on a large-scale online platform and open datasets, we present evidence that our proposed method is an effective approach for integrating query context to improve model ranking quality in terms of relevance and diversity.

cs.IR↗

Reducing Popularity Influence by Addressing Position Bias

Position bias poses a persistent challenge in recommender systems, with much of the existing research focusing on refining ranking relevance and driving user engagement. However, in practical applications, the mitigation of position bias does not always result in detectable short-term improvements in ranking relevance. This paper provides an alternative, practically useful view of what position bias reduction methods can achieve. It demonstrates that position debiasing can spread visibility and interactions more evenly across the assortment, effectively reducing a skew in the popularity of items induced by the position bias through a feedback loop. We offer an explanation of how position bias affects item popularity. This includes an illustrative model of the item popularity histogram and the effect of the position bias on its skewness. Through offline and online experiments on our large-scale e-commerce platform, we show that position debiasing can significantly improve assortment utilization, without any degradation in user engagement or financial metrics. This makes the ranking fairer and helps attract more partners or content providers, benefiting the customers and the business in the long term.

cs.IR↗

Efficient slot labelling

Slot labelling is an essential component of any dialogue system, aiming to find important arguments in every user turn. Common approaches involve large pre-trained language models (PLMs) like BERT or RoBERTa, but they face challenges such as high computational requirements and dependence on pre-training data. In this work, we propose a lightweight method which performs on par or better than the state-of-the-art PLM-based methods, while having almost 10x less trainable parameters. This makes it especially applicable for real-life industry scenarios.

cs.CL↗

UNICON: A unified framework for behavior-based consumer segmentation in e-commerce

Data-driven personalization is a key practice in fashion e-commerce, improving the way businesses serve their consumers needs with more relevant content. While hyper-personalization offers highly targeted experiences to each consumer, it requires a significant amount of private data to create an individualized journey. To alleviate this, group-based personalization provides a moderate level of personalization built on broader common preferences of a consumer segment, while still being able to personalize the results. We introduce UNICON, a unified deep learning consumer segmentation framework that leverages rich consumer behavior data to learn long-term latent representations and utilizes them to extract two pivotal types of segmentation catering various personalization use-cases: lookalike, expanding a predefined target seed segment with consumers of similar behavior, and data-driven, revealing non-obvious consumer segments with similar affinities. We demonstrate through extensive experimentation our framework effectiveness in fashion to identify lookalike Designer audience and data-driven style segments. Furthermore, we present experiments that showcase how segment information can be incorporated in a hybrid recommender system combining hyper and group-based personalization to exploit the advantages of both alternatives and provide improvements on consumer experience.

cs.IR↗

DIET: Lightweight Language Understanding for Dialogue Systems

Large-scale pre-trained language models have shown impressive results on language understanding benchmarks like GLUE and SuperGLUE, improving considerably over other pre-training methods like distributed representations (GloVe) and purely supervised approaches. We introduce the Dual Intent and Entity Transformer (DIET) architecture, and study the effectiveness of different pre-trained representations on intent and entity prediction, two common dialogue language understanding tasks. DIET advances the state of the art on a complex multi-domain NLU dataset and achieves similarly high performance on other simpler datasets. Surprisingly, we show that there is no clear benefit to using large pre-trained models for this task, and in fact DIET improves upon the current state of the art even in a purely supervised setup without any pre-trained embeddings. Our best performing model outperforms fine-tuning BERT and is about six times faster to train.

cs.CL↗

Dialogue Transformers

We introduce a dialogue policy based on a transformer architecture, where the self-attention mechanism operates over the sequence of dialogue turns. Recent work has used hierarchical recurrent neural networks to encode multiple utterances in a dialogue context, but we argue that a pure self-attention mechanism is more suitable. By default, an RNN assumes that every item in a sequence is relevant for producing an encoding of the full sequence, but a single conversation can consist of multiple overlapping discourse segments as speakers interleave multiple topics. A transformer picks which turns to include in its encoding of the current dialogue state, and is naturally suited to selectively ignoring or attending to dialogue history. We compare the performance of the Transformer Embedding Dialogue (TED) policy to an LSTM and to the REDP, which was specifically designed to overcome this limitation of RNNs.

cs.CL↗

Where is the context? -- A critique of recent dialogue datasets

Recent dialogue datasets like MultiWOZ 2.1 and Taskmaster-1 constitute some of the most challenging tasks for present-day dialogue models and, therefore, are widely used for system evaluation. We identify several issues with the above-mentioned datasets, such as history independence, strong knowledge base dependence, and ambiguous system responses. Finally, we outline key desiderata for future datasets that we believe would be more suitable for the construction of conversational artificial intelligence.

cs.CL↗

Few-Shot Generalization Across Dialogue Tasks

Machine-learning based dialogue managers are able to learn complex behaviors in order to complete a task, but it is not straightforward to extend their capabilities to new domains. We investigate different policies' ability to handle uncooperative user behavior, and how well expertise in completing one task (such as restaurant reservations) can be reapplied when learning a new one (e.g. booking a hotel). We introduce the Recurrent Embedding Dialogue Policy (REDP), which embeds system actions and dialogue states in the same vector space. REDP contains a memory component and attention mechanism based on a modified Neural Turing Machine, and significantly outperforms a baseline LSTM classifier on this task. We also show that both our architecture and baseline solve the bAbI dialogue task, achieving 100% test accuracy.

cs.CL↗

Thermodynamics of network model fitting with spectral entropies

An information theoretic approach inspired by quantum statistical mechanics was recently proposed as a means to optimize network models and to assess their likelihood against synthetic and real-world networks. Importantly, this method does not rely on specific topological features or network descriptors, but leverages entropy-based measures of network distance. Entertaining the analogy with thermodynamics, we provide a physical interpretation of model hyperparameters and propose analytical procedures for their estimate. These results enable the practical application of this novel and powerful framework to network model inference. We demonstrate this method in synthetic networks endowed with a modular structure, and in real-world brain connectivity networks.

cond-mat.stat-mech↗

Hub induced remote synchronization and desynchronization in complex networks of Kuramoto oscillators

The concept of "remote synchronization" (RS) was introduced in [Phys. Rev. E 85, 026208 (2012)], where synchronization in a star network of Stuart-Landau oscillators was investigated. In the RS regime therein described, the central hub served as a transmitter of information between peripheral nodes, while maintaining independent dynamics that were asynchronous with the rest of the network. One of the key conclusions of that paper was that RS cannot occur in pure phase-oscillator networks. Here, we show that the RS regime can exist in networks of Kuramoto oscillators, and that hub nodes can actively drive remote synchronization even in the presence of a repulsive mean field. We apply this model to study the synchronization dynamics in complex networks endowed with hub-nodes, an ubiquitous feature of many natural networks. We show that a change in the natural frequency of a hub can alone reshape synchronization patterns, and switch from direct to remote synchronization, or to hub-driven desynchronization. We discuss the potential role of this phenomenon in real-world networks, including the Karate-club and brain connectivity networks.

nlin.AO↗

Dynamics of weakly inhomogeneous oscillator populations: Perturbation theory on top of Watanabe-Strogatz integrability

As has been shown by Watanabe and Strogatz (WS) [Phys. Rev. Lett., 70, 2391 (1993)], a population of identical phase oscillators, sine-coupled to a common field, is a partially integrable system for any size: its dynamics reduces to equations for several collective variables. Here we develop a perturbation approach for weakly nonidentical ensembles. We calculate corrections to the WS dynamics for two types of perturbations: due to a distribution of natural frequencies and of forcing terms, and due to small white noise. We demonstrate, that in both cases the complex mean field for which the dynamical equations are written, is close up to the leading order in the perturbation to the Kuramoto order parameter. This supports validity of the dynamical reduction suggested by Ott and Antonsen [Chaos, 18, 037113 (2008)] for weakly inhomogeneous populations.

nlin.CD↗

Star-type oscillatory networks with generic Kuramoto-type coupling: a model for "Japanese drums synchrony"

We analyze star-type networks of phase oscillators by virtue of two methods. For identical oscillators we adopt the Watanabe-Strogatz approach, that gives full an- alytical description of states, rotating with constant frequency. For nonidentical oscillators, such states can be obtained by virtue of the self-consistent approach in a parametric form. In this case stability analysis cannot be performed, however with the help of direct numerical simulations we show which solutions are stable and which not. We consider this system as a model for a drum orchestra, where we assume that the drummers follow the signal of the leader without listening to each other and the coupling parameters are determined by a geometrical organization of the orchestra.

nlin.AO↗

Explosive Synchronization is Discontinuous

Spontaneous explosive is an abrupt transition to collective behavior taking place in heterogeneous networks when the frequencies of the nodes are positively correlated to the node degree. This explosive transition was conjectured to be discontinuous. Indeed, numerical investigations reveal a hysteresis behavior associated with the transition. Here, we analyze explosive synchronization in star graphs. We show that in the thermodynamic limit the transition to (and out) collective behavior is indeed discontinuous. The discontinuous nature of the transition is related to the nonlinear behavior of the order parameter, which in the thermodynamic limit exhibits multiple fixed points. Moreover, we unravel the hysteresis behavior in terms of the graph parameters. Our numerical results show that finite size graphs are well described by our predictions.

nlin.AO↗

Synchronization transitions in ensembles of noisy oscillators with bi-harmonic coupling

We describe synchronization transitions in an ensemble of globally coupled phase oscillators with a bi-harmonic coupling function, and two sources of disorder - diversity of intrinsic oscillatory frequencies and external independent noise. Based on the self-consistent formulation, we derive analytic solutions for different synchronous states. We report on various non-trivial transitions from incoherence to synchrony where possible scenarios include: simple supercritical transition (similar to classical Kuramoto model), subcritical transition with large area of bistability of incoherent and synchronous solutions, and also appearance of symmetric two-cluster solution which can coexist with regular synchronous state. Remarkably, we show that the interplay between relatively small white noise and finite-size fluctuations can lead to metastable asynchronous solution.

nlin.AO↗

Synchronization of oscillators in a Kuramoto-type model with generic coupling

We study synchronization properties of coupled oscillators on networks that allow description in terms of global mean field coupling. These models generalize the standard Kuramoto-Sakaguchi model, allowing for different contributions of oscillators to the mean field and to different forces from the mean field on oscillators. We present the explicit solutions of self-consistency equations for the amplitude and frequency of the mean field in a parametric form, valid for noise-free and noise-driven oscillators. As an example we consider spatially spreaded oscillators, for which the coupling properties are determined by finite velocity of signal propagation.

nlin.CD↗

Synchronization of a Josephson junction array in terms of global variables

We consider an array of Josephson junctions with a common LCR-load. Application of the Watanabe-Strogatz approach [Physica D, v. 74, p. 197 (1994)] allows us to formulate the dynamics of the array via the global variables only. For identical junctions this is a finite set of equations, analysis of which reveals the regions of bistability of the synchronous and asynchronous states. For disordered arrays with distributed parameters of the junctions, the problem is formulated as an integro-differential equation for the global variables, here stability of the asynchronous states and the properties of the transition synchrony-asynchrony are established numerically.

nlin.CD↗