Searcharxiv⌕ Search

arXiv subjects

Daniel Li

Publications and source records attributed to Daniel Li.

At least 37 records · Page 2Linked to original sources

Reward Shaping for User Satisfaction in a REINFORCE Recommender

How might we design Reinforcement Learning (RL)-based recommenders that encourage aligning user trajectories with the underlying user satisfaction? Three research questions are key: (1) measuring user satisfaction, (2) combatting sparsity of satisfaction signals, and (3) adapting the training of the recommender agent to maximize satisfaction. For measurement, it has been found that surveys explicitly asking users to rate their experience with consumed items can provide valuable orthogonal information to the engagement/interaction data, acting as a proxy to the underlying user satisfaction. For sparsity, i.e, only being able to observe how satisfied users are with a tiny fraction of user-item interactions, imputation models can be useful in predicting satisfaction level for all items users have consumed. For learning satisfying recommender policies, we postulate that reward shaping in RL recommender agents is powerful for driving satisfying user experiences. Putting everything together, we propose to jointly learn a policy network and a satisfaction imputation network: The role of the imputation network is to learn which actions are satisfying to the user; while the policy network, built on top of REINFORCE, decides which items to recommend, with the reward utilizing the imputed satisfaction. We use both offline analysis and live experiments in an industrial large-scale recommendation platform to demonstrate the promise of our approach for satisfying user experiences.

cs.IR↗

Boundedness of composition operators on general weighted Hardy spaces of analytic functions

We characterize the (essentially) decreasing sequences of positive numbers $β$ = ($β$ n) for which all composition operators on H 2 ($β$) are bounded, where H 2 ($β$) is the space of analytic functions f in the unit disk such that $\infty$ n=0 |c n | 2 $β$ n < $\infty$ if f (z) = $\infty$ n=0 c n z n. We also give conditions for the boundedness when $β$ is not assumed essentially decreasing.

math.FA↗

BOP2-DC: Bayesian optimal phase II designs with dual-criterion decision making

The conventional phase II trial design paradigm is to make the go/no-go decision based on the hypothesis testing framework. Statistical significance itself alone, however, may not be sufficient to establish that the drug is clinically effective enough to warrant confirmatory phase III trials. We propose the Bayesian optimal phase II trial design with dual-criterion decision making (BOP2-DC), which incorporates both statistical significance and clinical relevance into decision making. Based on the posterior probability that the treatment effect reaches the lower reference value (statistical significance) and the clinically meaningful value (clinical significance), BOP2-DC allows for go/consider/no-go decisions, rather than a binary go/no-go decision, and it is optimized to maximize the probability of a go decision when the treatment is effective or minimize the sample size when the treatment is futile. BOP2-DC is highly flexible and accommodates various types of endpoints, including binary, continuous, time-to-event, multiple, and co-primary endpoints, in single-arm and randomized trials. Simulation studies show that the BOP2-DC design yields desirable operating characteristics. The software to implement BOP2-DC is freely available at \url{www.trialdesign.org}.

stat.ME↗

Hierarchical Summarization for Longform Spoken Dialog

Every day we are surrounded by spoken dialog. This medium delivers rich diverse streams of information auditorily; however, systematically understanding dialog can often be non-trivial. Despite the pervasiveness of spoken dialog, automated speech understanding and quality information extraction remains markedly poor, especially when compared to written prose. Furthermore, compared to understanding text, auditory communication poses many additional challenges such as speaker disfluencies, informal prose styles, and lack of structure. These concerns all demonstrate the need for a distinctly speech tailored interactive system to help users understand and navigate the spoken language domain. While individual automatic speech recognition (ASR) and text summarization methods already exist, they are imperfect technologies; neither consider user purpose and intent nor address spoken language induced complications. Consequently, we design a two stage ASR and text summarization pipeline and propose a set of semantic segmentation and merging algorithms to resolve these speech modeling challenges. Our system enables users to easily browse and navigate content as well as recover from errors in these underlying technologies. Finally, we present an evaluation of the system which highlights user preference for hierarchical summarization as a tool to quickly skim audio and identify content of interest to the user.

cs.CL↗

Compactification and decompactification by weights on Bergman spaces

We characterize the symbols $Φ$ for which there exists a weight w such that the weighted composition operator M w C $Φ$ is compact on the weighted Bergman space B 2 $α$. We also characterize the symbols for which there exists a weight w such that M w C $Φ$ is bounded but not compact. We also investigate when there exists w such that M w C $Φ$ is Hilbert-Schmidt on B 2 $α$.

math.FA↗

Sentence Boundary Augmentation For Neural Machine Translation Robustness

Neural Machine Translation (NMT) models have demonstrated strong state of the art performance on translation tasks where well-formed training and evaluation data are provided, but they remain sensitive to inputs that include errors of various types. Specifically, in the context of long-form speech translation systems, where the input transcripts come from Automatic Speech Recognition (ASR), the NMT models have to handle errors including phoneme substitutions, grammatical structure, and sentence boundaries, all of which pose challenges to NMT robustness. Through in-depth error analysis, we show that sentence boundary segmentation has the largest impact on quality, and we develop a simple data augmentation strategy to improve segmentation robustness.

cs.CL↗

Central Limit Theorems for Compound Paths on the 2-Dimensional Lattice

Zeckendorf proved that every integer can be written uniquely as a sum of non-consecutive Fibonacci numbers $\{F_n\}$, and later researchers showed that the distribution of the number of summands needed for such decompositions of integers in $[F_n, F_{n+1})$ converges to a Gaussian as $n\to\infty$. Decomposition problems have been studied extensively for a variety of different sequences and notions of a legal decompositions; for the Fibonacci numbers, a legal decomposition is one for which each summand is used at most once and no two consecutive summands may be chosen. Recently, Chen et al. [CCGJMSY] generalized earlier work to $d$-dimensional lattices of positive integers; there, a legal decomposition is a path such that every point chosen had each component strictly less than the component of the previous chosen point in the path. They were able to prove Gaussianity results despite the lack of uniqueness of the decompositions; however, their results should hold in the more general case where some components are identical. The strictly decreasing assumption was needed in that work to obtain simple, closed form combinatorial expressions, which could then be well approximated and led to the limiting behavior. In this work we remove that assumption through inclusion-exclusion arguments. These lead to more involved combinatorial sums; using generating functions and recurrence relations we obtain tractable forms in $2$ dimensions and prove Gaussianity again; a more involved analysis should work in higher dimensions.

math.NT↗

Statistical Issues and Recommendations for Clinical Trials Conducted During the COVID-19 Pandemic

The COVID-19 pandemic has had and continues to have major impacts on planned and ongoing clinical trials. Its effects on trial data create multiple potential statistical issues. The scale of impact is unprecedented, but when viewed individually, many of the issues are well defined and feasible to address. A number of strategies and recommendations are put forward to assess and address issues related to estimands, missing data, validity and modifications of statistical analysis methods, need for additional analyses, ability to meet objectives and overall trial interpretability.

q-bio.OT↗

Comparison of singular numbers of composition operators on different Hilbert spaces of analytic functions

We compare the rate of decay of singular numbers of a given composition operator acting on various Hilbert spaces of analytic functions on the unit disk $\D$. We show that for the Hardy and Bergman spaces, our results are sharp. We also give lower and upper estimates of the singular numbers of the composition operator with symbol the ``cusp map'' and the lens maps, acting on weighted Dirichlet spaces.

math.FA↗

Compactification, and beyond, of composition operators on Hardy spaces by weights

We study when multiplication by a weight can turn a non-compact composition operator on H 2 into a compact operator, and when it can be in Schatten classes. The q-summing case in H p is considered. We also study when this multiplication can turn a compact composition operator into a non-compact one. MSC 2010 primary: 47B33 ; secondary: 46B28

math.FA↗

Composition operators with surjective symbol and small approximation numbers

We give a new proof of the existence of a surjective symbol whose associated composition operator on H 2 (D) is in all Schatten classes, with the improvement that its approximation numbers can be, in some sense, arbitrarily small. We show, as an application, that, contrary to the 1-dimensional case, for N $\ge$ 2, the behavior of the approximation numbers a n = a n (C $Φ$), or rather of $β$ -- N = lim inf n$\rightarrow$$\infty$ [a n ] 1/n 1/N or $β$ + N = lim sup n$\rightarrow$$\infty$ [a n ] 1/n 1/N , of composition operators on H 2 (D N) cannot be determined by the image of the symbol. MSC 2010 Primary: 47B33 Secondary: 32A35 ; 46B28

math.FA↗

Pluricapacity and approximation numbers of composition operators

For suitable bounded hyperconvex sets $Ω$ in $\mathbb{C}^N$, in particular the ball or the polydisk, we give estimates for the approximation numbers of composition operators $C_ϕ\colon H^2 (Ω) \to H^2 (Ω)$ when $ϕ(Ω)$ is relatively compact in $Ω$, involving the Monge-Ampère capacity of $ϕ(Ω)$.

math.FA↗

Some examples of composition operators and their approximation numbers on the Hardy space of the bi-disk

We give examples of composition operators $C\_Φ$ on $H^2 (\D^2)$ showing that the condition $\|Φ\|\_\infty = 1$ is not sufficient for their approximation numbers $a\_n (C\_Φ)$ to satisfy $\lim\_{n \to \infty} [a\_n (C\_Φ) ]^{1/\sqrt{n}} = 1$, contrary to the $1$-dimensional case. We also give a situation where this implication holds. We make a link with the Monge-Ampère capacity of the image of $Φ$.

math.FA↗

Adaptive Memory Networks

We present Adaptive Memory Networks (AMN) that processes input-question pairs to dynamically construct a network architecture optimized for lower inference times for Question Answering (QA) tasks. AMN processes the input story to extract entities and stores them in memory banks. Starting from a single bank, as the number of input entities increases, AMN learns to create new banks as the entropy in a single bank becomes too high. Hence, after processing an input-question(s) pair, the resulting network represents a hierarchical structure where entities are stored in different banks, distanced by question relevance. At inference, one or few banks are used, creating a tradeoff between accuracy and performance. AMN is enabled by dynamic networks that allow input dependent network creation and efficiency in dynamic mini-batching as well as our novel bank controller that allows learning discrete decision making with high accuracy. In our results, we demonstrate that AMN learns to create variable depth networks depending on task complexity and reduces inference times for QA tasks.

cs.AI↗

Approximation numbers of weighted composition operators

We study the approximation numbers of weighted composition operators $f\mapsto w\cdot(f\circφ)$ on the Hardy space $H^2$ on the unit disc. For general classes of such operators, upper and lower bounds on their approximation numbers are derived. For the special class of weighted lens map composition operators with specific weights, we show how much the weight $w$ can improve the decay rate of the approximation numbers, and give sharp upper and lower bounds. These examples are motivated from applications to the analysis of relative commutants of special inclusions of von Neumann algebras appearing in quantum field theory (Borchers triples).

math.FA↗