Searcharxiv⌕ Search

arXiv subjects

Bin Fu

Publications and source records attributed to Bin Fu.

107 records · Page 6Linked to original sources

A Model for Donation Verification

In this paper, we introduce a model for donation verification. A randomized algorithm is developed to check if the money claimed being received by the collector is $(1-ε)$-approximation to the total amount money contributed by the donors. We also derive some negative results that show it is impossible to verify the donations under some circumstances.

cs.DS↗

Deep & Cross Network for Ad Click Predictions

Feature engineering has been the key to the success of many prediction models. However, the process is non-trivial and often requires manual feature engineering or exhaustive searching. DNNs are able to automatically learn feature interactions; however, they generate all the interactions implicitly, and are not necessarily efficient in learning all types of cross features. In this paper, we propose the Deep & Cross Network (DCN) which keeps the benefits of a DNN model, and beyond that, it introduces a novel cross network that is more efficient in learning certain bounded-degree feature interactions. In particular, DCN explicitly applies feature crossing at each layer, requires no manual feature engineering, and adds negligible extra complexity to the DNN model. Our experimental results have demonstrated its superiority over the state-of-art algorithms on the CTR prediction dataset and dense classification dataset, in terms of both model accuracy and memory usage.

cs.LG↗

Concentration Independent Random Number Generation in Tile Self-Assembly

In this paper we introduce the \emph{robust random number generation} problem where the goal is to design an abstract tile assembly system (aTAM system) whose terminal assemblies can be split into $n$ partitions such that a resulting assembly of the system lies within each partition with probability 1/$n$, regardless of the relative concentration assignment of the tile types in the system. First, we show this is possible for $n=2$ (a \emph{robust fair coin flip}) within the aTAM, and that such systems guarantee a worst case $\mathcal{O}(1)$ space usage. We accompany our primary construction with variants that show trade-offs in space complexity, initial seed size, temperature, tile complexity, bias, and extensibility, and also prove some negative results. As an application, we combine our coin-flip system with a result of Chandran, Gopalkrishnan, and Reif to show that for any positive integer $n$, there exists a $\mathcal{O}(\log n)$ tile system that assembles a constant-width linear assembly of expected length $n$ for any concentration assignment. We then extend our robust fair coin flip result to solve the problem of robust random number generation in the aTAM for all $n$. Two variants of robust random bit generation solutions are presented: an unbounded space solution and a bounded space solution which incurs a small bias. Further, we consider the harder scenario where tile concentrations change arbitrarily at each assembly step and show that while this is not possible in the aTAM, the problem can be solved by exotic tile assembly models from the literature.

cs.FL↗

Partial Sublinear Time Approximation and Inapproximation for Maximum Coverage

We develop a randomized approximation algorithm for the classical maximum coverage problem, which given a list of sets $A_1,A_2,\cdots, A_m$ and integer parameter $k$, select $k$ sets $A_{i_1}, A_{i_2},\cdots, A_{i_k}$ for maximum union $A_{i_1}\cup A_{i_2}\cup\cdots\cup A_{i_k}$. In our algorithm, each input set $A_i$ is a black box that can provide its size $|A_i|$, generate a random element of $A_i$, and answer the membership query $(x\in A_i?)$ in $O(1)$ time. Our algorithm gives $(1-{1\over e})$-approximation for maximum coverage problem in $O(p(m))$ time, which is independent of the sizes of the input sets. No existing $O(p(m)n^{1-ε})$ time $(1-{1\over e})$-approximation algorithm for the maximum coverage has been found for any function $p(m)$ that only depends on the number of sets, where $n=\max(|A_1|,\cdots,| A_m|)$ (the largest size of input sets). The notion of partial sublinear time algorithm is introduced. For a computational problem with input size controlled by two parameters $n$ and $m$, a partial sublinear time algorithm for it runs in a $O(p(m)n^{1-ε})$ time or $O(q(n)m^{1-ε})$ time. The maximum coverage has a partial sublinear time $O(p(m))$ constant factor approximation. On the other hand, we show that the maximum coverage problem has no partial sublinear $O(q(n)m^{1-ε})$ time constant factor approximation algorithm. It separates the partial sublinear time computation from the conventional sublinear time computation by disproving the existence of sublinear time approximation algorithm for the maximum coverage problem.

cs.DS↗

Order O(1) algorithm for first-principles transient current through open quantum systems

In the study the response time of ultrafast transistor and peak transient current to prevent melt down of nano-chips, the first principles transient current calculation plays an essential role in nanoelectronics. The first principles calculation of transient current through nano-devices for a period of time T is known to be extremely time consuming with the best scaling TN^3 where N is the dimension of the device. In this work, we provide an order O(1) algorithm that reduces the computational complexity to T^0 N^3 for large systems. Benchmark calculation has been done on graphene nanoribbons with N = 10^4 confirming the O(1) scaling. This breakthrough allows us to tackle many large scale transient problems including magnetic tunneling junctions and ferroelectric tunneling junctions that cannot be touched before.

cond-mat.mes-hall↗

Derandomizing Polynomial Identity over Finite Fields Implies Super-Polynomial Circuit Lower Bounds for NEXP

We show that derandomizing polynomial identity testing over an arbitrary finite field implies that NEXP does not have polynomial size boolean circuits. In other words, for any finite field F(q) of size q, $PIT_q\in NSUBEXP\Rightarrow NEXP\not\subseteq P/poly$, where $PIT_q$ is the polynomial identity testing problem over F(q), and NSUBEXP is the nondeterministic subexpoential time class of languages. Our result is in contract to Kabanets and Impagliazzo's existing theorem that derandomizing the polynomial identity testing in the integer ring Z implies that NEXP does have polynomial size boolean circuits or permanent over Z does not have polynomial size arithmetic circuits.

cs.CC↗

Sublinear Time Motif Discovery from Multiple Sequences

A natural probabilistic model for motif discovery has been used to experimentally test the quality of motif discovery programs. In this model, there are $k$ background sequences, and each character in a background sequence is a random character from an alphabet $Σ$. A motif $G=g_1g_2...g_m$ is a string of $m$ characters. Each background sequence is implanted a probabilistically generated approximate copy of $G$. For a probabilistically generated approximate copy $b_1b_2...b_m$ of $G$, every character $b_i$ is probabilistically generated such that the probability for $b_i\neq g_i$ is at most $α$. We develop three algorithms that under the probabilistic model can find the implanted motif with high probability via a tradeoff between computational time and the probability of mutation. The methods developed in this paper have been used in the software implementation. We observed some encouraging results that show improved performance for motif detection compared with other softwares.

cs.DS↗

Sublinear Time Approximate Sum via Uniform Random Sampling

We investigate the approximation for computing the sum $a_1+...+a_n$ with an input of a list of nonnegative elements $a_1,..., a_n$. If all elements are in the range $[0,1]$, there is a randomized algorithm that can compute an $(1+ε)$-approximation for the sum problem in time ${O({n(\log\log n)\over\sum_{i=1}^n a_i})}$, where $ε$ is a constant in $(0,1)$. Our randomized algorithm is based on the uniform random sampling, which selects one element with equal probability from the input list each time. We also prove a lower bound $Ω({n\over \sum_{i=1}^n a_i})$, which almost matches the upper bound, for this problem.

cs.DS↗

On the Complexity of Approximate Sum of Sorted List

We consider the complexity for computing the approximate sum $a_1+a_2+...+a_n$ of a sorted list of numbers $a_1\le a_2\le ...\le a_n$. We show an algorithm that computes an $(1+ε)$-approximation for the sum of a sorted list of nonnegative numbers in an $O({1\over ε}\min(\log n, {\log ({x_{max}\over x_{min}})})\cdot (\log {1\over ε}+\log\log n))$ time, where $x_{max}$ and $x_{min}$ are the largest and the least positive elements of the input list, respectively. We prove a lower bound $Ω(\min(\log n,\log ({x_{max}\over x_{min}}))$ time for every O(1)-approximation algorithm for the sum of a sorted list of nonnegative elements. We also show that there is no sublinear time approximation algorithm for the sum of a sorted list that contains at least one negative number.

cs.DS↗

Self-Assembly with Geometric Tiles

In this work we propose a generalization of Winfree's abstract Tile Assembly Model (aTAM) in which tile types are assigned rigid shapes, or geometries, along each tile face. We examine the number of distinct tile types needed to assemble shapes within this model, the temperature required for efficient assembly, and the problem of designing compact geometric faces to meet given compatibility specifications. Our results show a dramatic decrease in the number of tile types needed to assemble $n \times n$ squares to $Θ(\sqrt{\log n})$ at temperature 1 for the most simple model which meets a lower bound from Kolmogorov complexity, and $O(\log\log n)$ in a model in which tile aggregates must move together through obstacle free paths within the plane. This stands in contrast to the $Θ(\log n / \log\log n)$ tile types at temperature 2 needed in the basic aTAM. We also provide a general method for simulating a large and computationally universal class of temperature 2 aTAM systems with geometric tiles at temperature 1. Finally, we consider the problem of computing a set of compact geometric faces for a tile system to implement a given set of compatibility specifications. We show a number of bounds on the complexity of geometry size needed for various classes of compatibility specifications, many of which we directly apply to our tile assembly results to achieve non-trivial reductions in geometry size.

cs.CG↗

A Dense Hierarchy of Sublinear Time Approximation Schemes for Bin Packing

The bin packing problem is to find the minimum number of bins of size one to pack a list of items with sizes $a_1,..., a_n$ in $(0,1]$. Using uniform sampling, which selects a random element from the input list each time, we develop a randomized $O({n(\log n)(\log\log n)\over \sum_{i=1}^n a_i}+({1\over ε})^{O({1\overε})})$ time $(1+ε)$-approximation scheme for the bin packing problem. We show that every randomized algorithm with uniform random sampling needs $Ω({n\over \sum_{i=1}^n a_i})$ time to give an $(1+ε)$-approximation. For each function $s(n): N\rightarrow N$, define $\sum(s(n))$ to be the set of all bin packing problems with the sum of item sizes equal to $s(n)$. For a constant $b\in (0,1)$, every problem in $\sum(n^{b})$ has an $O(n^{1-b}(\log n)(\log\log n)+({1\over ε})^{O({1\overε})})$ time $(1+ε)$-approximation for an arbitrary constant $ε$. On the other hand, there is no $o(n^{1-b})$ time $(1+ε)$-approximation scheme for the bin packing problems in $\sum(n^{b})$ for some constant $ε>0$.

cs.CC↗

Multivariate Polynomial Integration and Derivative Are Polynomial Time Inapproximable unless P=NP

We investigate the complexity of integration and derivative for multivariate polynomials in the standard computation model. The integration is in the unit cube $[0,1]^d$ for a multivariate polynomial, which has format $f(x_1,\cdots, x_d)=p_1(x_1,\cdots, x_d)p_2(x_1,\cdots, x_d)\cdots p_k(x_1,\cdots, x_d)$, where each $p_i(x_1,\cdots, x_d)=\sum_{j=1}^d q_j(x_j)$ with all single variable polynomials $q_j(x_j)$ of degree at most two and constant coefficients. We show that there is no any factor polynomial time approximation for the integration $\int_{[0,1]^d}f(x_1,\cdots,x_d)d_{x_1}\cdots d_{x_d}$ unless $P=NP$. For the complexity of multivariate derivative, we consider the functions with the format $f(x_1,\cdots, x_d)=p_1(x_1,\cdots, x_d)p_2(x_1,\cdots, x_d)\cdots p_k(x_1,\cdots, x_d),$ where each $p_i(x_1,\cdots, x_d)$ is of degree at most $2$ and $0,1$ coefficients. We also show that unless $P=NP$, there is no any factor polynomial time approximation to its derivative ${\partial f^{(d)}(x_1,\cdots, x_d)\over \partial x_1\cdots \partial x_d}$ at the origin point $(x_1,\cdots, x_d)=(0,\cdots,0)$. Our results show that the derivative may not be easier than the integration in high dimension. We also give some tractable cases of high dimension integration and derivative.

cs.CC↗

NE is not NP Turing Reducible to Nonexpoentially Dense NP Sets

A long standing open problem in the computational complexity theory is to separate NE from BPP, which is a subclass of $NP_T(NP\cap P/poly)$. In this paper, we show that $NE\not\subseteq NP_(NP \cap$ Nonexponentially-Dense-Class), where Nonexponentially-Dense-Class is the class of languages A without exponential density (for each constant c>0,$|A^{\le n}|\le 2^{n^c}$ for infinitely many integers n). Our result implies $NE\not\subseteq NP_T({pad(NP, g(n))})$ for every time constructible super-polynomial function g(n) such as $g(n)=n^{\ceiling{\log\ceiling{\log n}}}$, where Pad(NP, g(n)) is class of all languages $L_B=\{s10^{g(|s|)-|s|-1}:s\in B\}$ for $B\in NP$. We also show $NE\not\subseteq NP_T(P_{tt}(NP)\cap Tally)$.

cs.CC↗

The Complexity of Testing Monomials in Multivariate Polynomials

The work in this paper is to initiate a theory of testing monomials in multivariate polynomials. The central question is to ask whether a polynomial represented by certain economically compact structure has a multilinear monomial in its sum-product expansion. The complexity aspects of this problem and its variants are investigated with two folds of objectives. One is to understand how this problem relates to critical problems in complexity, and if so to what extent. The other is to exploit possibilities of applying algebraic properties of polynomials to the study of those problems. A series of results about $ΠΣΠ$ and $ΠΣ$ polynomials are obtained in this paper, laying a basis for further study along this line.

cs.CC↗

Algorithms for Testing Monomials in Multivariate Polynomials

This paper is our second step towards developing a theory of testing monomials in multivariate polynomials. The central question is to ask whether a polynomial represented by an arithmetic circuit has some types of monomials in its sum-product expansion. The complexity aspects of this problem and its variants have been investigated in our first paper by Chen and Fu (2010), laying a foundation for further study. In this paper, we present two pairs of algorithms. First, we prove that there is a randomized $O^*(p^k)$ time algorithm for testing $p$-monomials in an $n$-variate polynomial of degree $k$ represented by an arithmetic circuit, while a deterministic $O^*(6.4^k + p^k)$ time algorithm is devised when the circuit is a formula, here $p$ is a given prime number. Second, we present a deterministic $O^*(2^k)$ time algorithm for testing multilinear monomials in $Π_mΣ_2Π_t\times Π_kΠ_3$ polynomials, while a randomized $O^*(1.5^k)$ algorithm is given for these polynomials. The first algorithm extends the recent work by Koutis (2008) and Williams (2009) on testing multilinear monomials. Group algebra is exploited in the algorithm designs, in corporation with the randomized polynomial identity testing over a finite field by Agrawal and Biswas (2003), the deterministic noncommunicative polynomial identity testing by Raz and Shpilka (2005) and the perfect hashing functions by Chen {\em at el.} (2007). Finally, we prove that testing some special types of multilinear monomial is W[1]-hard, giving evidence that testing for specific monomials is not fixed-parameter tractable.

cs.CC↗

Approximating Multilinear Monomial Coefficients and Maximum Multilinear Monomials in Multivariate Polynomials

This paper is our third step towards developing a theory of testing monomials in multivariate polynomials and concentrates on two problems: (1) How to compute the coefficients of multilinear monomials; and (2) how to find a maximum multilinear monomial when the input is a $ΠΣΠ$ polynomial. We first prove that the first problem is \#P-hard and then devise a $O^*(3^ns(n))$ upper bound for this problem for any polynomial represented by an arithmetic circuit of size $s(n)$. Later, this upper bound is improved to $O^*(2^n)$ for $ΠΣΠ$ polynomials. We then design fully polynomial-time randomized approximation schemes for this problem for $ΠΣ$ polynomials. On the negative side, we prove that, even for $ΠΣΠ$ polynomials with terms of degree $\le 2$, the first problem cannot be approximated at all for any approximation factor $\ge 1$, nor {\em "weakly approximated"} in a much relaxed setting, unless P=NP. For the second problem, we first give a polynomial time $λ$-approximation algorithm for $ΠΣΠ$ polynomials with terms of degrees no more a constant $λ\ge 2$. On the inapproximability side, we give a $n^{(1-ε)/2}$ lower bound, for any $ε>0,$ on the approximation factor for $ΠΣΠ$ polynomials. When terms in these polynomials are constrained to degrees $\le 2$, we prove a $1.0476$ lower bound, assuming $P\not=NP$; and a higher $1.0604$ lower bound, assuming the Unique Games Conjecture.

cs.CC↗

XML Reconstruction View Selection in XML Databases: Complexity Analysis and Approximation Scheme

Query evaluation in an XML database requires reconstructing XML subtrees rooted at nodes found by an XML query. Since XML subtree reconstruction can be expensive, one approach to improve query response time is to use reconstruction views - materialized XML subtrees of an XML document, whose nodes are frequently accessed by XML queries. For this approach to be efficient, the principal requirement is a framework for view selection. In this work, we are the first to formalize and study the problem of XML reconstruction view selection. The input is a tree $T$, in which every node $i$ has a size $c_i$ and profit $p_i$, and the size limitation $C$. The target is to find a subset of subtrees rooted at nodes $i_1,\cdots, i_k$ respectively such that $c_{i_1}+\cdots +c_{i_k}\le C$, and $p_{i_1}+\cdots +p_{i_k}$ is maximal. Furthermore, there is no overlap between any two subtrees selected in the solution. We prove that this problem is NP-hard and present a fully polynomial-time approximation scheme (FPTAS) as a solution.

cs.DS↗