{"title": "Near Optimal Sketching of Low-Rank Tensor Regression", "book": "Advances in Neural Information Processing Systems", "page_first": 3466, "page_last": 3476, "abstract": "We study the least squares regression problem $\\min_{\\Theta \\in \\RR^{p_1 \\times \\cdots \\times p_D}} \\| \\cA(\\Theta) -  b \\|_2^2$, where $\\Theta$ is a low-rank tensor, defined as $\\Theta = \\sum_{r=1}^{R} \\theta_1^{(r)} \\circ \\cdots \\circ \\theta_D^{(r)}$, for vectors $\\theta_d^{(r)} \\in \\mathbb{R}^{p_d}$ for all $r \\in [R]$ and $d \\in [D]$.    %$R$ is small compared with $p_1,\\ldots,p_D$,   Here, $\\circ$ denotes the outer product of vectors, and $\\cA(\\Theta)$ is a linear function on $\\Theta$. This problem is motivated by the fact that the number of parameters in $\\Theta$ is only $R \\cdot \\sum_{d=1}^D p_D$, which is significantly smaller than the $\\prod_{d=1}^{D} p_d$ number of parameters in ordinary least squares regression. We consider the above CP decomposition model of tensors $\\Theta$, as well as the Tucker decomposition. For both models we show how to apply data dimensionality reduction techniques based on {\\it sparse} random projections $\\Phi \\in \\RR^{m \\times n}$, with $m \\ll n$, to reduce the problem to a much smaller problem $\\min_{\\Theta} \\|\\Phi \\cA(\\Theta) - \\Phi b\\|_2^2$, for which $\\|\\Phi \\cA(\\Theta) - \\Phi b\\|_2^2 = (1 \\pm \\varepsilon) \\| \\cA(\\Theta) -  b \\|_2^2$ holds simultaneously for all $\\Theta$. We obtain a significantly smaller dimension and sparsity in the randomized linear mapping $\\Phi$ than is possible for ordinary least squares regression. Finally, we give a number of numerical simulations supporting our theory.", "full_text": "Near Optimal Sketching of Low-Rank Tensor Regression\n\nJarvis Haupt1\n\njdhaupt@umn.edu\n\nXingguo Li1,2\n\nlixx1661@umn.edu\n\nDavid P. Woodruff 3\n\ndwoodruf@cs.cmu.edu \u21e4\n\n1 University of Minnesota\n\n2Georgia Tech\n\n3Carnegie Mellon University\n\nAbstract\nWe study the least squares regression problem\n\nmin\n\nr=1 \u2713(r)\n\n1 \u00b7\u00b7\u00b7 \u2713(r)\n\n\u21e52Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD kA(\u21e5)  bk2\n2,\nwhere \u21e5 is a low-rank tensor, de\ufb01ned as \u21e5= PR\nD , for vectors\n\u2713(r)\nd 2 Rpd for all r 2 [R] and d 2 [D]. Here,  denotes the outer product of\nvectors, and A(\u21e5) is a linear function on \u21e5. This problem is motivated by the fact\nthat the number of parameters in \u21e5 is only R \u00b7PD\nd=1 pd, which is signi\ufb01cantly\nsmaller than theQD\nd=1 pd number of parameters in ordinary least squares regression.\nWe consider the above CP decomposition model of tensors \u21e5, as well as the\nTucker decomposition. For both models we show how to apply data dimensionality\nreduction techniques based on sparse random projections  2 Rm\u21e5n, with m \u2327 n,\nto reduce the problem to a much smaller problem min\u21e5 kA(\u21e5)bk2\n2, for which\n2 holds simultaneously for all \u21e5. We obtain\nkA(\u21e5) bk2\na signi\ufb01cantly smaller dimension and sparsity in the randomized linear mapping \nthan is possible for ordinary least squares regression. Finally, we give a number of\nnumerical simulations supporting our theory.\n\n2 = (1\u00b1 \")kA(\u21e5) bk2\n\nIntroduction\n\n1\nFor a sequence of D-way design tensors Ai 2 Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD, i 2 [n] , {1, . . . , n}, suppose we observe\nnoisy linear measurements of an unknown D-way tensor \u21e5 2 Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD, given by\n(1)\nwhere A(\u00b7) : Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD ! Rn is a linear function with Ai(\u21e5) = hAi, \u21e5i = vec(Ai)>vec(\u21e5) for\nall i 2 [n], vec(X) is the vectorization of a tensor X, and z = [z1, . . . , zn]> corresponds to the\nobservation noise. Given the design tensors {Ai}n\ni=1 and noisy observations b = [b1, . . . , bn]>, a\nnatural approach for estimating the parameter \u21e5 is to use the Ordinary Least Square (OLS) estimation\nfor the tensor regression problem, i.e., to solve\n\nb = Ai(\u21e5) + z, b, z 2 Rn,\n\nmin\n\n\u21e52Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD kA(\u21e5)  bk2\n2.\n\n(2)\n\nTensor regression has been widely studied in the literature. Applications include computer vision\n[8, 19, 34], data mining [5], multi-model ensembles [32], neuroimaging analysis [15, 36], multitask\nlearning [21, 31], and multivariate spatial-temporal data analysis [1, 11]. In these applications,\nmodeling the unknown parameters as a tensor is what is needed, as it allows for learning data that\nhas multi-directional relations, such as in climate prediction [33], inherent structure learning with\nmulti-dimensional indices [21], and hand movement trajectory decoding [34].\n\n\u21e4The authors are listed in alphabetical order. Correspondence to: Xingguo Li <lixx1661@umn.edu>.\nThe authors acknowledge support from University of Minnesota Startup Funding and Doctoral Dissertation\nFellowship from University of Minnesota.\n\n31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.\n\n\fDue to the high dimensionality of tensor data, structured learning based on low-rank tensor de-\ncompositions, such as CANDECOMP/PARAFAC (CP) decomposition and Tucker decomposition\nmodels [13, 24], have been proposed in order to obtain tractable tensor regression problems. As\ndiscussed more below, requiring the unknown tensor to be low-rank signi\ufb01cantly reduces the number\nof unknown parameters.\nWe consider low-rank tensor regression problems based on the CP decomposition and Tucker\ndecomposition models. For simplicity, we \ufb01rst focus on the CP model, and later extend our analysis\nto the Tucker model. Suppose that \u21e5 admits a rank-R CP decomposition, that is,\n\n\u21e5=\n\n\u2713(r)\n1 \u00b7\u00b7\u00b7 \u2713(r)\nD ,\n\n(3)\n\nRXr=1\n\n(4)\n\n(5)\n\nwhere \u2713(r)\nreparameterize the set of low-rank tensors by its matrix slabs/factors:\n\nd 2 Rpd for all r 2 [R], d 2 [D], and  is the outer product of vectors. For convenience, we\nSD,R ,n[[\u21e51, . . . , \u21e5D]] | \u21e5d = [\u2713(1)\n\n] 2 Rpd\u21e5R, for all d 2 [D]o.\n\nThen we can rewrite model (1) in a compact form\n\nd , . . . ,\u2713 (R)\n\nd\n\nb = A(\u21e5D \u00b7\u00b7\u00b7 \u21e51)1R + z,\n\nwhere A = [vec(A1),\u00b7\u00b7\u00b7 , vec(An)]> 2 Rn\u21e5QD\nd=1 pd is the matricization of all design tensors,\n1R = [1, . . . , 1] 2 RR is a vector of all 1s, \u2326 is the Kronecker product, and  is the Khatri-Rao\nproduct. In addition, the OLS estimation for tensor regression (2) can be rewritten as the following\nnonconvex problem in terms of low-rank tensor parameters [[\u21e51, . . . , \u21e5D]],\n\nmin\n\n2, where\n\n#2SD,R kA#  bk2\nSD,R ,n(\u21e5D \u00b7\u00b7\u00b7 \u21e51)1R 2 RQD\n\nd=1 pd [[\u21e51, . . . , \u21e5D]] 2S D,Ro.\n\nThe number of parameters for a general tensor \u21e5 2 Rp1\u21e5\u00b7\u00b7\u00b7\u21e5pD isQD\nd=1 pd, which may be prohibitive\nd=1. The bene\ufb01t of the low-rank tensor model (3) is that\nfor estimation even for small values of {pd}D\nit dramatically reduces the degrees of freedom of the unknown tensor fromQD\nd=1 pd,\nwhere we are typically interested in the case when R \uf8ff pd for all d 2 [D]. For example, a typical\nMRI image has size 2563 \u21e1 1.7 \u21e5 107, while using the low-rank model with R = 10, we reduce the\nnumber of unknown parameters to 256 \u21e5 3 \u21e5 10 \u21e1 8 \u21e5 103 \u2327 107. This signi\ufb01cantly increases the\napplicability of the tensor regression model in practice.\nNevertheless, solving the tensor regression problem (5) is still expensive in terms of both computation\nand memory requirements, for typical settings, when n  R\u00b7PD\nd=1 pd. In particular, the per iteration\ncomplexity is at least linear in n for popular algorithms such as block alternating minimization and\nblock gradient descent [27, 28]. In addition, in order to store A, it takes n \u00b7QD\nd=1 pd words of\nmemory. Both of these aspects are undesirable when n is large. This motivates us to consider data\ndimensionality reduction techniques, also called sketching, for the tensor regression problem.\nInstead of solving (5), we consider the simple Sketched Ordinary Least Square (SOLS) problem:\n\nd=1 pd to R \u00b7PD\n\nmin\n\n#2SD,R kA#  bk2\n2,\n\n(6)\n\nwhere  2 Rm\u21e5n is a random matrix (speci\ufb01ed in Section 2). Importantly,  will satisfy two\nproperties, namely (1) m \u2327 n so that we signi\ufb01cantly reduce the size of the problem, and (2)  will\nbe very sparse so that v can be computed very quickly for any v 2 Rn.\nNa\u00efvely applying existing analyses of sketching techniques for least squares regression requires\nm =\u2326( QD\nd=1 pd), which is prohibitive (for a survey, see, e.g., [30]). In this paper, our main\ncontribution is to show that it is possible to use a sparse Johnson-Lindenstrauss transformation as\nour sketching matrix for the CP model of low-rank tensor regression, with constant column sparsity\nand dimension m = R \u00b7PD\nd=1 pd, up to poly-logarithmic (polylog) factors. Note that our dimension\nmatches the number of intrinsic parameters in the CP model. Further, we stress that we do not assume\nanything about the tensor, such as orthogonal matrix slabs/factor, or incoherence; our dimensionality\n\n2\n\n\freduction works for arbitrary tensors. We show, with the above sparsity and dimenion, that with\nconstant probability, simultaneously for all # 2S ,D,R,\n\nkA#  bk2\n\n2 = (1 \u00b1 \")kA#  bk2\n2.\n\nreduction for this problem, i.e., dimensionality reduction better thanQD\n\nThis implies that any solution to (6) has the same cost as in (5) up to a (1 + \")-factor. In partic-\nular, by solving (6) we obtain a (1 + \")-approximation to (5). We note that our dimensionality\nreduction technique is not tied to any particular algorithm; that is, if one runs any algorithm or\nheuristic on the reduced (sketched) problem, obtaining an \u21b5-approximate solution #, then # is also a\n(1+\u270f)\u21b5-approximate solution to the original problem. Our result is the \ufb01rst non-trivial dimensionality\nd=1 pd, which is trivial by ig-\nnoring the low-rank structure of the tensor, and which achieves a relative error (1 + \")-approximation.\nWhile it may be possible to apply dimensionality reduction methods directly in alternating minimiza-\ntion methods for solving tensor regression, unlike our method, such methods do not have provable\nguarantees and it is not clear how errors propagate across iterations. However, since we reduce the\noriginal problem to a smaller version of itself with a provable guarantee, one could further apply\ndimensionality reduction techniques as heuristics for alternating minimization on the smaller problem.\nOur proof is based on a careful characterization of Talagrand\u2019s functional for the parameter space\nof low-rank tensors, providing a highly nontrivial analysis for what we consider to be a simple and\npractical algorithm. One of the main dif\ufb01culties is dealing with general, non-orthogonal tensors, for\nwhich we are able to provide a careful re-parameterization in order to bound the so-called Finsler\nmetric; interestingly, for non-orthogonal tensors it is always possible to partially orthogonalize them,\nand this partial orthogonalization turns out to suf\ufb01ce for our analysis. We give precise details below.\nWe also provide numerical evaluations on both synthetic and real data to demonstrate the empirical\nperformance of our algorithm.\nNotation. For scalars x, y 2 R, let x = (1 \u00b1 \")y if x 2 [(1  \")y, (1 + \")y], x . (&)y if\nx \uf8ff ()c1y, poly(x) = xc2 and polylog(x, y) = (log x)c3 \u00b7 (log y)c4 for some universal constants\nc1, c2, c3, c4 > 0. We also use standard asymptotic notations O(\u00b7) and \u2326(\u00b7). Given a matrix\nA 2 Rm\u21e5n, we denote kAk2 as the spectral norm, span(A) \u2713 Rm as the subspace spanned by the\ncolumns of A, max(A) and min(A) as the largest and smallest singular values of A, respectively,\nand \uf8ffA = max(A)/min(A) as the condition number. We use nnz(A) to denote the number\nof nonzero entries of A, and PA as the projection operator onto span(A). Given two matrices\nA = [a1, . . . , an] 2 Rm\u21e5n and B = [b1, . . . , bq] 2 Rp\u21e5q, A\u2326B = [a1\u2326B, . . . , an\u2326B] 2 Rmp\u21e5nq\ndenotes the Kronecker product, and AB = [a1\u2326b1, . . . , an\u2326bn] 2 Rmp\u21e5n denotes the Khatri-Rao\nproduct with n = q. We let Bn \u21e2 Rn be the unit sphere in Rn, i.e., Bn = {x 2 Rn |k xk2 = 1},\nP(\u00b7) be the probability of an event, and E(\u00b7) denotes the expectation of a random variable. Without\nfurther speci\ufb01cation, we denoteQ =QD\nd=1. We further summarize the dimension\nparameters for ease of reference. Given a tensor \u21e5, D is the number of ways, pd is the dimension of\nthe d-th way for d 2 [D]. R is the rank of \u21e5 for all ways under the CP decomposition, and Rd is the\nrank of the d-th way under the Tucker decomposition for d 2 [D]. n is the number of observations\nfor tensor regression. m is the sketching dimension and s is the sparsity of each column in a sparse\nJohnson-Lindenstrauss transformation.\n\nd=1 andP =PD\n\n2 Background\n\nWe start with a few important de\ufb01nitions.\nDe\ufb01nition 1 (Oblivious Subspace Embedding). Suppose \u21e7 is a distribution on m\u21e5 n matrices where\nm is a function of parameters n, d, and \". Further, suppose that with probability at least 1  , for\nany \ufb01xed n \u21e5 d matrix A, a matrix  drawn from \u21e7 has the property that kAxk2\n2 = (1 \u00b1 \")kAxk2\nsimultaneously for all x 2X\u2713 Rd. Then \u21e7 is an (\", ) oblivious subspace embedding (OSE) of X .\nAn OSE  preserves the norm of vectors in a certain set X after linear transformation by A. This is\nwidely studied as a key property for sketching based analyses (see [30] and the references therein).\nWe want to show an analogous property when X is parameterized by low-rank tensors.\nDe\ufb01nition 2 (Leverage Scores). Given A 2 Rn\u21e5d, let Z 2 Rn\u21e5d have orthonormal columns that\nspan the column space of A. Then `2\n\n2 is the i-th leverage score of A.\n\n2\n\ni (A) = ke>i Zk2\n\n3\n\n\fLeverage scores play an important role in randomized matrix algorithms [7, 16, 17]. Calculating\nthe leverage scores na\u00efvely by orthogonalizing A requires O(nd2) time. It is shown in [3] that\nthe leverage scores of A can be approximated individually up to a constant multiplicative factor in\nO(nnz(A) log n + poly(d)) time using sparse subspace embeddings. In our analysis, there will be a\nvery mild dependence on the maximum leverage score of A and the sparsity for the sketching matrix\n. Note that we do not need to calculate the leverage scores.\nDe\ufb01nition 3 (Talagrand\u2019s Functional). Given a (semi-)metric \u21e2 on Rn and a bounded set S\u21e2 Rn,\nTalagrand\u2019s 2-functional is\n\n(7)\n\nsup\nx2S\n\n{Sr}1r=0\n\n2(S, \u21e2) = inf\n\n2r/2 \u00b7 \u21e2(x,Sr),\n\n1Xr=0\nwhere \u21e2(x,Sr) is a distance from x to Sr and the in\ufb01mum is taken over all collections {Sr}1r=0 such\nthat S0 \u21e2S 1 \u21e2 . . . \u21e2S with |S0| = 1 and |Sr|\uf8ff 22r.\nA closely related notion of the 2-functional is the Gausssian mean width: G(S) = Eg supx2Shg, xi,\nwhere g \u21e0N n(0, In). For any bounded S\u21e2 Rn, G(S) and 2(S,k\u00b7k 2) differ multiplicatively by at\nmost a universal constant in Euclidean space [25]. Finding a tight upper bound on the 2-functional\nfor the parameter space of low-rank tensors is key to our analysis.\nDe\ufb01nition 4 (Finsler Metric). Let E, E0 \u21e2 Rn be p-dimensional subspaces. The Finsler metric of\nE and E0 is \u21e2Fin(E, E0) = kPE P E0k2, where PE is the projection onto the subspace E.\nThe Finsler metric is the semi-metric used in the 2-functional in our analysis. Note that \u21e2Fin(E, E0) \uf8ff\n1 always holds for any E and E0 [23].\nDe\ufb01nition 5 (Sparse Johnson-Lindenstrauss Transforms). Let ij be independent Rademacher\nrandom variables, i.e., P(ij = 1) = P(ij = 1) = 1/2, and let ij :\u2326  !{ 0, 1} be random\nvariables, independent of the ij, with the following properties:\n(i) ij are negatively correlated for \ufb01xed j, i.e., for all 1 \uf8ff i1 < . . . < ik \uf8ff m, we have\nE\u21e3Qk\nt=1 it,j\u2318 \uf8ffQk\n(ii) There are s =Pm\n(iii) The vectors (ij)m\nThen  2 Rm\u21e5n is a sparse Johnson-Lindenstrauss transform (SJLT) matrix if ij = 1ps ijij.\nThe SJLT has several bene\ufb01ts [4, 12, 30]. First, the computation of x takes only O(nnz(x)) time\nwhen s is a constant. Second, storing  takes only sn memory instead of mn, which is signi\ufb01cant\nwhen s \u2327 m. This can often further be reduced by drawing the entries of  from a limited\nindependent family of random variables.\nWe will use an SJLT matrix as the sketching matrix  in our analysis and our goal will be to\nshow suf\ufb01cient conditions on the sketching dimension m and per-column sparsity s such that the\nanalogue of the OSE property holds for low-rank tensor regression. Speci\ufb01cally, we provide suf\ufb01cient\nconditions for the SJLT matrix  2 Rm\u21e5n to preserve the cost of all solutions for tensor regression,\ni.e., bounds on m and s for which\n\nt=1 E (it,j) = s\nmk;\ni=1 ij nonzero ij for a \ufb01xed j; and\ni=1 are independent across j 2 [n].\n\nE\n\n\nsup\n\nx2Tkxk2\n\n2  1 <\n\n\"\n10\n\n,\n\n(8)\n\nwhere \" is a given precision and T is a normalized space parameterized as the union of certain\nsubspaces of A, which will be further discussed in the following sections. Note that by linearity, it is\nsuf\ufb01cient to consider x with kxk2 = 1 in the above, which explains the form of (8). Moreover, by\nMarkov\u2019s inequality, (8) implies that simultaneously for all # = vec(\u21e5) 2S D,R, where \u21e5 admits a\nlow-rank tensor decomposition, with probability at least 9/10, we have\n(9)\nwhich allows us to minimize the much smaller sketched problem to obtain parameters # which, when\nplugged into the original objective function, provide a multiplicative (1 + \")-approximation.\n\n2 = (1 \u00b1 \")kA#  bk2\n2,\n\nkA#  bk2\n\n4\n\n\f3 Dimensionality Reduction for CP Decomposition\n\n1\n\n\\1 , . . . , A\u2713(R)\n\nD , where \u2713(r)\n\nd 2 Rpd\n\n1 \u00b7\u00b7\u00b7 \u2713(r)\n\nWe start with the following notation. Given a tensor \u21e5= PR\nfor all d 2 [D] and r 2 [R], we \ufb01x all but \u2713(r)\n\\1o =hA\u2713(1)\nAn\u2713(r)\njD=1 \u00b7\u00b7\u00b7Pp2\nj2=1 A(jD,...,j2)\u2713(i)\n\nr=1 \u2713(r)\nfor r 2 [R], and denote\n\\1 i 2 Rn\u21e5Rp1,\n\nwhere A\u2713(i)\nd , and\nA(jD,...,j2) 2 Rn\u21e5p1 is a column submatrix of A indexed by jD 2 [pD], . . . , j2 2 [p2], i.e.,\nA =\u21e5A(1,...,1), . . . , A(pD,...,p2)\u21e4 2 Rn\u21e5Q pd. The above parameterization allows us to view tensor\nregression as preserving the norms of vectors in an in\ufb01nite union of subspaces, described in more\ndetail in the full version of our paper [10]. Then we rewrite the observation model (4) as\ni> + z.\n\nd,jd is the jd-th entry of \u2713(i)\n\n\\1o \u00b7h\u2713(1)>1\n\n\\1 = PpD\n\n1 + z = An\u2713(r)\n\n\u2713(r)\nD \u2326\u00b7\u00b7\u00b7\u2326 \u2713(r)\n\nD,jD \u00b7\u00b7\u00b7 \u2713(i)\n\n2,j2, \u2713(i)\n\nA\u2713(r)\n\n\\1 \u00b7 \u2713(r)\n\nb = A \u00b7\n\nRXr=1\n\n1 + z =\n\nRXr=1\n\n. . .\u2713 (R)>1\n\nr=1 \u2713(r)\n\nd , (r)\n\nr=1 (r)\n\n[18]. The following theorem provides suf\ufb01cient conditions to guarantee (1 + \")-approximation of the\nobjective for low-rank tensor regression under the CP decomposition model.\n\n3.1 Main Result\nThe parameter space for the tensor regression problem (1) is a subspace of RQ pd, i.e., SD,R \u21e2\nRQ pd. Therefore, a na\u00efve application of sketching requires m &Q pd/\"2 in order for (9) to hold\nTheorem 1. Suppose R \uf8ff maxd pd/2 and maxi2[n] `2\nd 2B pdo\nT = [r2[R],d2[D]n A#A'\nD \u2326\u00b7\u00b7\u00b7\u2326 \u2713(r)\nand let  2 Rm\u21e5n be an SJLT matrix with column sparsity s. Then with probability at least 9/10, (9)\nholds if m and s satisfy, respectively,\n\ni (A) \uf8ff 1/(RPD\n1 ,' =PR\n\nkA#A'k2# =PR\n\nd=2 pd)2. Let\n1 ,\u2713 (r)\n\nD \u2326\u00b7\u00b7\u00b7\u2326 (r)\n\n\u2326(1), up to logarithmic factors, we can guarantee (1 + \")-approximation of the objective. The\nsketching complexity of m is nearly optimal compared with the number of free parameters for the CP\n\nm & RX pd log\u21e3DR\uf8ffAX pd\u2318 polylog(m, n)/\"2 and s & log2\u21e3X pd\u2318 polylog(m, n)/\"2.\nFrom Theorem 1, we have that for an SJLT matrix  2 Rm\u21e5n with m =\u2326( RP pd) and s =\ndecomposition model, i.e., R(P pd  D + 1), up to logarithmic factors. Here wo do not make any\northogonality assumption on the tensor factors \u2713(r)\nd , and show in our analysis that the general tensor\nspace T can be paramterized in terms of an orthogonal one if R \uf8ff maxd pd/2 holds. The condition\nR \uf8ff maxd pd/2 is not restrictive in our setting, as we are interested in low-rank tensors with R \uf8ff pd.\nNote that we achieve a (1 + \")-approximation in objective function value for arbitrary tensors; if one\nwants to achieve closeness of the underlying parameters one needs to impose further assumptions on\nthe model, such as the form of the noise distribution or structural properties of A [20, 36].\nOur maximum leverage score assumption is very mild and much weaker than the standard inco-\nherence assumptions used for example, in matrix completion, which allow for uniform sampling\nbased approaches. For example, our assumption states that the maximum leverage score is at most\n\n1/(RPD\nd=2 pd)2. In the typical overconstrained case, n Q pd, and in order for uniform sampling\nto provide a subspace embedding, one needs the maximum leverage score to be at most RP pd/n\n(see, e.g., Section 2.4 of [30]), which is much less than 1/(RPD\nd=2 pd)2 when n is large, and so\nuniform sampling fails in our setting. Moreover, it is also possible to apply a standard idea to \ufb02atten\nthe leverage scores of a deterministic design A based on the Subsampled Randomized Hadamard\nTransformation (SRHT) using the Walsh-Hadamard matrix [9, 26]. Note that applying the SRHT to\nan n \u21e5 d matrix A only takes O(nd log n) time, which if A is dense, is the same amount of time one\nneeds just to read A (up to a log n factor). Further details are deferred to the full version of our paper\n[10].\n\n5\n\n\f3.2 Proof Sketch of Our Analysis for a Basic Case\n\nWe provide a sketch of our analysis for the case when R = 1 and D = 2, i.e., \u21e5 is rank 1 matrix. The\nanalysis for more general cases is more involved, but with similar intuition. Details of the analyses\nare deferred to the the full version of our paper, where we start with a proof for the most basic cases\nand gradually build up the proof for the most general case.\n\nLet Av = Pp2\ni=1 A(i)vi, where A = [A(1), . . . , A(p2)] 2 Rn\u21e5p2p1 with A(i) 2 Rn\u21e5p1 for all\ni 2 [p2], V =SfW {span[Av1, Av2]}, andfW = {v1, v2 2B p2 with hv1, v2i = 0}. We start with an\nillustration that the set T can be reparameterized to the following set with respect to tensors with\northogonal factors:\n\nT = [E2V\n\n{x 2 E |k xk2 = 1} .\n\n=\n\n=\n\nSuppose hv1, v2i 6= 0. Let v2 = \u21b5v1 + z for some \u21b5,  2 R and a unit vector z 2 Rp2, where\nhv1, zi = 0. Then we have\nAx  Ay\nkAx  Ayk2\n\nAv1(u1  \u21b5u2)  Az(u2)\nkAv1(u1  \u21b5u2)  Az(u2)k2\n\nAv1u1  Av2u2\nkAv1u1  Av2u2k2\n\nwhich is equivalent to hv1, v2i = 0 by reparameterizing z as v2.\nBased on known dimensionality reduction results [2, 6] (see further details in the full version [10]), the\nmain quantities needed for bounding properties of  are the quantities \u21e2V, 2\n2(V,\u21e2 Fin), N (V,\u21e2 Fin,\" 0),\nandR \"0\n0 (log N (V,\u21e2 Fin, t))1/2 dt, where N (V,\u21e2 Fin, t) is the covering number of V under the Finsler\nmetric using balls of radius t and pV = supv1,v22Bp2 ,hv1,v2i=0 dim{span (Av1,v2)}\uf8ff 2p1. Bounding\nthese quantities for the space of low-rank tensors is new and is our main technical contribution. These\nwill be addressed separately as follows.\nPart 1: Bound pV. Let Av1,v2 = [Av1, Av2]. It is straightforward that pV \uf8ff 2p1.\nPart 2: Bound 2\n\n2(V,\u21e2 Fin). By the de\ufb01nition of 2-functional in (7) for the Finsler metric, we have\n\n,\n\n2(V,\u21e2 Fin) = inf\n\n{V k}1k=0\n\nsup\n\nAv1,v22V\n\n2k/2 \u00b7 \u21e2Fin(Av1,v2,V k),\n\n1Xk=0\n\nwhere V k is an \"k-net of V, i,e., for any Av1,v2 2V there exist v1, v2 2B p2 with hv1, v2i = 0,\nkv1  v1k2 \uf8ff \u2318k, and kv2  v2k2 \uf8ff \u2318k, such that Av1,v2 2 V k and \u21e2Fin(Av1,v2, Av1,v2) \uf8ff \"k.\nFrom Lemma 6, we have \u21e2Fin(Av1,v2,V k) \uf8ff 2\uf8ffA\u2318k for kv1  v1k2 \uf8ff \u2318k and kv2  v2k2 \uf8ff\n\u2318k. On the other hand, we have that \u21e2Fin(Av1,v2,V k) \uf8ff 1 always holds. Therefore, we have\n\u21e2Fin(Av1,v2,V k) \uf8ff min{2\uf8ffA\u2318k, 1}. Let k0 be the smallest integer such that 2\uf8ffA\u2318k0 \uf8ff 1. Then\n\n2k/2 +\n\nk0Xk=0\n\n1Xk=0\n\n2(V,\u21e2 Fin) \uf8ff\n\n2k/2\u21e2Fin(Av1,v2,V k).\n\n2k/2\u21e2Fin(Av1,v2,V k) \uf8ff\n\n1Xk=k0+1\nStarting from \u23180 = 1 and |V 0| = 1, for k  1, we have \u2318k < 1 and |V k|\uf8ff (3/\u2318k)p2 [29]. Also from\nthe 2-functional, we require |V k|\uf8ff 22k \uf8ff (3/\u2318k)p2, which implies\n1\n\u2318k0\n\nk0Xk=0\nk such that (3/\u2318k+1)p2 \uf8ff 22k+1. Then we have |V k+1|\uf8ff 22k+1.\n\nFor k > k0, we choose \u2318k+1 = \u23182\nBy choosing k0 to be the smallest integer such that (3/\u2318k0+1)p2 \uf8ff 22k0+1 holds, we have\n\uf8ff 2k0/2 .rp2 log\n\n2k/2 \u00b7 \u21e2Fin(Av1,v2,V k) = 2k0/2 \u00b7\n\n2t/2 \u00b7\u2713 1\n2\u25c62t\n\n.rp2 log\n\n2k0/2\np2  1\n\n2k/2 =\n\n1\n\u2318k0\n\n.\n\n(12)\n\n(10)\n\n(11)\n\n.\n\n1Xk=k0+1\n\n1Xt=1\n\n6\n\n\f.\n\n\uf8ffA\n\"0\n\nthis implies\n\n2(V,\u21e2 Fin) . p2 log\n2\n0 [log N (V,\u21e2 Fin, t)]1/2dt. From our choice from Part 2, \"0 2\n\"0\u23182p2. From direct integration,\n\nCombining (10) \u2013 (12), and choosing a small enough \"0 such that \"0 \uf8ff 2\uf8ffA\u2318k0, we have\nPart 3: Bound N (V,\u21e2 Fin,\" 0) andR \"0\n(0, 1) is a constant. Then it is straightforward that N (V,\u21e2 Fin,\" 0) \uf8ff\u21e3 3\n[log N (V,\u21e2 Fin, t)]1/2dt.\"0rp2 log\n\"0\u2318 \u00b7 polylog(m, n)\n\"0\u2318 \u00b7 polylog(m, n)\n\nCombining the results in Parts 1, 2, and 3, we have that (9) holds if m and s satisfy, respectively\n\nm & \u21e3p2 log \uf8ffA\ns & \u21e3log2 1\n\n0(p1 + p2) log 1\n\n+ p1 + p2 log 1\n\nZ \"0\n\n+ \"2\n\n1\n\"0\n\nand\n\n\"2\n\n\"0\n\n\"0\n\n0\n\n.\n\n.\n\n\"2\n\nWe \ufb01nish the proof by taking \"0 = 1/(p1 + p2).\n\n4 Dimensionality Reduction for Tucker Decomposition\n\nWe start with a formal model description. Suppose \u21e5 admits the following Tucker decomposition:\n\n\u21e5=\n\nR1Xr1=1\n\n\u00b7\u00b7\u00b7\n\nRDXrD=1\n\nG(r1, . . . , rD) \u00b7 \u2713(r1)\n\n1\n\n\u00b7\u00b7\u00b7 \u2713(rD)\nD ,\n\n(13)\n\n\\1\n\n\\1\n\n\\1\n\nrD=1 A\u2713(r1,...,rD )\n\nrD=1 A\u2713(r1,...,rD )\n\nD,jD \u00b7\u00b7\u00b7 \u2713(r2)\n\nj2=1 A(jD,...,j2)\u2713(rD)\n\nrD=1 G(r1, . . . , rD)\u2713(rD)\n\nr2=1\u00b7\u00b7\u00b7PRD\n1 + z = A\u21e2\u2713{rd}\n\n2 Rpd for all rd 2 [Rd] and d 2 [D]. Let\n2,j2 and\n\nG(1, r2,. . ., rD),. . .,PR2\nD \u2326\u00b7\u00b7\u00b7\u2326 \u2713(r1)\n\nwhere G 2 RR1\u21e5\u00b7\u00b7\u00b7\u21e5RD is the core tensor and \u2713(rd)\n=PpD\njD=1\u00b7\u00b7\u00b7Pp2\nA\u2713(r1,...,rD )\nA\u21e2\u2713{rd}\n\\1  =\uf8ffPR2\nr2=1\u00b7\u00b7\u00b7PRD\nThen the observation model (4) can be written as\nr1=1\u00b7\u00b7\u00b7PRD\nb = APR1\nThe following theorem provides suf\ufb01cient conditions to guarantee (1 + \")-approximation of the\nobjective function for low-rank tensor regression under the Tucker decomposition model.\nTheorem 2. Suppose nnz(G) \uf8ff maxd pd/2 and maxi2[n] `2\nLet T = [r2[R],d2[D]n A#  A'\nRDXrD=1\nRDXrD=1\nD \u2326\u00b7\u00b7\u00b7\u2326 (r1)\nG2(r1, . . . , rD) \u00b7 (rD)\n\nD \u2326\u00b7\u00b7\u00b7\u2326 \u2713(r1)\nd 2B pdo\nand  2 Rm\u21e5n be an SJLT matrix with column sparsity s. Then with probability at least 9/10, (9)\nholds if m and s satisfy\n\n\\1 h\u2713(1)>1\ni (A) \uf8ff 1/(PD\nG1(r1, . . . , rD) \u00b7 \u2713(rD)\n\nkA#  A'k2# =\nR1Xr1=1\n\nG(R1, r2,. . ., rD).\ni> + z.\n\nd=2 Rdpd + nnz(G))2.\n\nR1Xr1=1\n\n. . .\u2713 (R1)>1\n\nm & C1 \u00b7 log\u21e3C1D\uf8ffAR1pnnz(G)\u2318 \u00b7 polylog(m, n)/\"2 and s & log2 C1 \u00b7 polylog(m, n)/\"2,\nwhere C1 =P Rdpd + nnz(G).\nFrom Theorem 2, we have that using an SJLT matrix  with m =\u2326( P Rdpd + nnz(G)) and\n\ns = \u2326(1), up to logarithmic factors, we can guarantee (1+\")-approximation of the objective function.\n\n,\u2713 (rd)\n\n, (rd)\n\n\u00b7\u00b7\u00b7\n\n\u00b7\u00b7\u00b7\n\n' =\n\n,\n\n1\n\n1\n\nd\n\nd\n\n7\n\n\fThe sketching complexity of m is near optimal compared with the number of free parameters for\nthe Tucker decomposition model, i.e.,P Rdpd + nnz(G) P R2\nd, up to logarithmic factors. Note\nthat nnz(G) \uf8ff Q Rd, and thus the condition that nnz(G) \uf8ff maxd pd/2 can be more restrictive\nthan R \uf8ff maxd pd/2 in the CP model when nnz(G) > R. This is due to the fact that the Tucker\nmodel is more \u201cexpressive\u201d than the CP model for a tensor of the same dimensions. For example, if\nR1 = \u00b7\u00b7\u00b7 = RD = R, then the CP model (3) can be viewed as special case of the Tucker model (13)\nby setting all off-diagonal entries of the core tensor G to be 0. Moreover, the conditions and results\nin Theorem 2 are essentially of the same order as those in Theorem 1 when nnz(G) = R, which\nindicates the tightness of our analysis.\n\n5 Experiments\nWe study the performance of sketching for tensor regression through numerical experiments over\nboth synthetic and real data sets. For solving the OLS problem for tensor regression (2), we use a\ncyclic block-coordinate minimization algorithm based on a tensor toolbox [35]. Speci\ufb01cally, in a\ncyclic manner for all d 2 [D], we \ufb01x all but one \u21e5d of [[\u21e51, . . . , \u21e5D]] 2S D,R and minimize the\nresulting quadratic loss function (2) with respect to \u21e5i, until the decrease of the objective is smaller\nthan a prede\ufb01ned threshold \u2327. For SOLS, we use the same algorithm after multiplying A and b with\nan SJLT matrix . All results are run on a supercomputer due to the large scale of the data. Note\nthat our result is not tied to any speci\ufb01c algorithm and we can use any algorithm that solves OLS for\nlow-rank tensors for solving SOLS for low-rank tensors.\nFor synthetic data, we generate the low-rank tensor \u21e5 as follows. For each d 2 [D], we generate R\nrandom columns with N (0, 1) entries to form non-orthogonal tensor factors \u21e5d = [\u2713(1)\nd , . . . ,\u2713 (R)\n]\nof [[\u21e51, . . . , \u21e5D]] 2S D,R independently. We also generate R real scalars \u21b51, . . . ,\u21b5 R uniformly\nand independently from [1, 10]. Then \u21e5 is formed by \u21e5= PR\nD . The n tensor\ni=1 are generated independently with i.i.d. N (0, 1) entries for 10% of the entries chosen\ndesigns {Ai}n\nuniformly at random, and the remaining entries are set to zero. We also generate the noise z to have\nz) entries, and the generation of the SJLT matrix  follows De\ufb01nition 5. For both OLS\ni.i.d. N (0, 2\nand SOLS, we use random initializations for \u21e5, i.e., \u21e5d has i.i.d. N (0, 1) entries for all d 2 [D].\nWe compare OLS and SOLS for low-rank tensor regression under both the noiseless and noisy\nscenarios. For the noiseless case, i.e., z = 0, we choose R = 3, p1 = p2 = p3 = 100, m =\n5 \u21e5 R(p1 + p2 + p3) = 4500, and s = 200. Different values of n = 104, 105, and 106 are chosen to\ncompare both statistical and computational performances of OLS and SOLS. For the noisy case, the\nsettings of all parameters are identical to those in the noiseless case, except that z = 1. We provide\na plot of the scaled objective versus the number of iterations for some random trials in Figure 1. The\n2/n for OLS, where\nscaled objective is set to be kA#t\n\n2/n for SOLS and kA#t\n\n1 \u00b7\u00b7\u00b7 \u2713(r)\n\nSOLS  bk2\n\nOLS  bk2\n\nr=1 \u21b5r\u2713(r)\n\nd\n\ne\nv\ni\nt\nc\ne\nj\nb\nO\n\n105\n\n100\n\n10-5\n\n10-10\n\nSOLS n1\nSOLS n2\nSOLS n3\nOLS n1\nOLS n2\nOLS n3\n\n106\n\n104\n\n102\n\n100\n\ne\nv\ni\nt\nc\ne\nj\nb\nO\n\nSOLS n1\nSOLS n2\nSOLS n3\nOLS n1\nOLS n2\nOLS n3\n\n5\n\n10\nIteration\n(a) z = 0\n\n15\n\n20\n\n25\n\n5\n\n10\n\n15\n\n20\n\nIteration\n(b) z = 1\n\nFigure 1: Comparison of SOLS and OLS on synthetic data. The vertical axis corresponds to the scaled\nobjectives kA#t\n2/n for OLS, where #t is the update in the\nt-th iteration. The horizontal axis corresponds to the number of iterations (passes of block-coordinate\nminimization for all blocks). For both the noiseless case z = 0 and noisy case z = 1, we set\nn1 = 104, n2 = 105, and n3 = 106 respectively.\n\n2/n for SOLS and kA#t\n\nSOLS  bk2\n\nOLS  bk2\n\n8\n\n\f2/n and kA#SOLS  bk2\n\nOLS are the updates in the t-th iterations of SOLS and OLS respectively. Note the we are\nSOLS and #t\n#t\n2/n as the objective function for solving the SOLS problem, but looking at\nusing kA#SOLS  bk2\n2/n for the solution of SOLS is ultimately what we are interested\nthe original objective kA#SOLS  bk2\n2/n is very\nin. However, we have that the gap between kA#SOLS  bk2\nsmall in our results (< 1%). The number of iterations is the number of passes of block-coordinate\nminimization for all blocks. We can see that OLS and SOLS require approximately the same number\nof iterations for comparable decrease in objective function value. However, since the SOLS instance\nhas a much smaller size, its per iteration computational cost is much lower than that of OLS.\nWe further provide numerical results on the running time (CPU execution time) and the optimal\nscaled objectives in Table 1. Using the same stopping criterion, we see that SOLS and OLS achieve\ncomparable objectives (within < 5% differences), matching our theory. In terms of the running time,\nSOLS is signi\ufb01cantly faster than OLS, especially when n is large compared to the sketching dimen\nsion m. For example, when n = 106, SOLS is more than 200 times faster than OLS while achieving\na comparable objective function value with OLS. This matches with our theoretical results on the\ncomputational cost of OLS versus SOLS. Note that here we suppose that the rank is known for our\nsimulation, which can be restrictive in practice. We observe that if we choose a moderately larger\nrank than the true rank of the underlying model, then the results are similar to what we discussed\nabove. Smaller values of the rank result in a much deteriorated statistical performance for both OLS\nand SOLS.\nWe also examine sketching of low-rank tensor regression on a real dataset of MRI images [22]. The\ndataset consists of 56 frames of a human brain, each of which is of dimension 128 \u21e5 128 pixels, i.e.,\np1 = p2 = 128 and p3 = 56. The generation of design tensors {Ai}n\ni=1 and linear measurements\nb follows the same settings as for the synthetic data, with z = 0. We choose three values of\nR = 3, 5, 10, and set m = 5 \u21e5 R(p1 + p2 + p3). The sample size is set to n = 104 for all settings of\nR. Analogous to the synthetic data, we provide numerical results for SOLS and OLS on the running\ntime (CPU execution time) and the optimal scaled objectives. The results are provided in Table 2.\nAgain, we have that SOLS is much faster than OLS and they achieve comparable optimal objectives,\nunder all settings of ranks.\n\nTable 1: Comparison of SOLS and OLS on CPU execution time (in seconds) and the optimal scaled\nobjective over different choices of sample sizes and noise levels on synthetic data. The results are\naveraged over 50 random trials, with both the mean values and standard deviations (in parentheses)\nprovided. Note that we terminate the program after the running time exceeds 3 \u21e5 104 seconds.\n\nVariance of Noise\n\nSample Size\n\nTime\n\nObjective\n\nOLS\n\nSOLS\n\nOLS\n\nSOLS\n\nn = 104\n175.37\n(65.784)\n120.34\n(35.711)\n< 1010\n(< 1010)\n< 1010\n(< 1010)\n\nz = 0\nn = 105\n3683.9\n(1496.7)\n128.09\n(37.293)\n< 1010\n(< 1010)\n< 1010\n(< 1010)\n\nn = 106\n> 3 \u21e5 104\n\n(NA)\n132.93\n(38.649)\n< 1010\n(< 1010)\n< 1010\n(< 1010)\n\nn = 104\n168.62\n(24.570)\n121.71\n(34.214)\n0.9153\n(0.0256)\n0.9376\n(0.0261)\n\nz = 1\nn = 105\n2707.3\n(897.14)\n124.84\n(33.774)\n0.9341\n(0.0213)\n0.9817\n(0.0242)\n\nn = 106\n> 3 \u21e5 104\n\n(NA)\n128.65\n(32.863)\n0.9425\n(0.0172)\n0.9901\n(0.0256)\n\nTable 2: Comparison of SOLS and OLS on CPU execution time (in seconds) and the optimal scaled\nobjective over different choices of ranks on the MRI data. The results are averaged over 10 random\ntrials, with both the mean values and standard deviations (in parentheses) provided.\n\nRank\n\nTime\n\nObjective\n\nR = 3\n2824.4\n(768.08)\n16.003\n(0.1378)\n\nOLS\nR = 5\n8137.2\n(1616.3)\n11.164\n(0.1152)\n\nR = 10\n26851\n(8320.1)\n6.8679\n(0.0471)\n\nR = 3\n196.31\n(68.180)\n17.047\n(0.1561)\n\nSOLS\nR = 5\n364.09\n(145.79)\n11.992\n(0.1538)\n\nR = 10\n761.73\n(356.76)\n7.3968\n(0.0975)\n\n9\n\n\fReferences\n[1] Mohammad Taha Bahadori, Qi Rose Yu, and Yan Liu. Fast multivariate spatio-temporal analysis\nvia low-rank tensor learning. In Advances in Neural Information Processing Systems, pages\n3491\u20133499, 2014.\n\n[2] Jean Bourgain, Sjoerd Dirksen, and Jelani Nelson. Toward a uni\ufb01ed theory of sparse dimen-\nsionality reduction in Euclidean space. Geometric and Functional Analysis, 25(4):1009\u20131088,\n2015.\n\n[3] Kenneth L Clarkson and David P Woodruff. Low rank approximation and regression in input\nsparsity time. In Proceedings of the 45th Annual ACM Symposium on Theory of Computing,\npages 81\u201390. ACM, 2013.\n\n[4] Anirban Dasgupta, Ravi Kumar, and Tam\u00e1s Sarl\u00f3s. A sparse Johnson\u2013Lindenstrauss transform.\nIn Proceedings of the 42nd Annual ACM Symposium on Theory of Computing, pages 341\u2013350.\nACM, 2010.\n\n[5] Lieven De Lathauwer, Bart De Moor, and Joos Vandewalle. A multilinear singular value\ndecomposition. SIAM Journal on Matrix Analysis and Applications, 21(4):1253\u20131278, 2000.\n\n[6] Sjoerd Dirksen. Dimensionality reduction with subgaussian matrices: A uni\ufb01ed theory. Foun-\n\ndations of Computational Mathematics, pages 1\u201330, 2015.\n\n[7] Petros Drineas, Malik Magdon-Ismail, Michael W Mahoney, and David P Woodruff. Fast\napproximation of matrix coherence and statistical leverage. Journal of Machine Learning\nResearch, 13(Dec):3475\u20133506, 2012.\n\n[8] Weiwei Guo, Irene Kotsia, and Ioannis Patras. Tensor learning for regression. IEEE Transactions\n\non Image Processing, 21(2):816\u2013827, 2012.\n\n[9] Nathan Halko, Per-Gunnar Martinsson, and Joel A Tropp. Finding structure with randomness:\nProbabilistic algorithms for constructing approximate matrix decompositions. SIAM Review,\n53(2):217\u2013288, 2011.\n\n[10] Jarvis Haupt, Xingguo Li, and David P Woodruff. Near optimal sketching of low-rank tensor\n\nregression. arXiv preprint arXiv:1709.07093, 2017.\n\n[11] Peter D Hoff. Multilinear tensor regression for longitudinal relational data. The Annals of\n\nApplied Statistics, 9(3):1169, 2015.\n\n[12] Daniel M. Kane and Jelani Nelson. Sparser Johnson-Lindenstrauss transforms. Journal of the\n\nACM, 61(1):4:1\u20134:23, 2014.\n\n[13] Tamara G Kolda and Brett W Bader. Tensor decompositions and applications. SIAM Review,\n\n51(3):455\u2013500, 2009.\n\n[14] Bingxiang Li, Wen Li, and Lubin Cui. New bounds for perturbation of the orthogonal projection.\n\nCalcolo, 50(1):69\u201378, 2013.\n\n[15] Xiaoshan Li, Hua Zhou, and Lexin Li. Tucker tensor regression and neuroimaging analysis.\n\narXiv preprint arXiv:1304.5637, 2013.\n\n[16] Michael W Mahoney. Randomized algorithms for matrices and data. Foundations and Trends R\n\nin Machine Learning, 3(2):123\u2013224, 2011.\n\n[17] Michael W Mahoney and Petros Drineas. CUR matrix decompositions for improved data\n\nanalysis. Proceedings of the National Academy of Sciences, 106(3):697\u2013702, 2009.\n\n[18] Jelani Nelson and Huy L Nguyen. Lower bounds for oblivious subspace embeddings. In\nInternational Colloquium on Automata, Languages, and Programming, pages 883\u2013894. Springer,\n2014.\n\n10\n\n\f[19] Sung Won Park and Marios Savvides.\n\nIndividual kernel tensor-subspaces for robust face\nrecognition: A computationally ef\ufb01cient tensor framework without requiring mode factorization.\nIEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 37(5):1156\u20131166,\n2007.\n\n[20] Garvesh Raskutti and Ming Yuan. Convex regularization for high-dimensional tensor regression.\n\narXiv preprint arXiv:1512.01215, 2015.\n\n[21] Bernardino Romera-Paredes, Hane Aung, Nadia Bianchi-Berthouze, and Massimiliano Pontil.\nMultilinear multitask learning. In Proceedings of the 30th International Conference on Machine\nLearning, pages 1444\u20131452, 2013.\n\n[22] Antoine Rosset, Luca Spadola, and Osman Ratib. Osirix: an open-source software for navigating\n\nin multidimensional DICOM images. Journal of Digital Imaging, 17(3):205\u2013216, 2004.\n\n[23] Zhongmin Shen. Lectures on Finsler geometry, volume 2001. World Scienti\ufb01c, 2001.\n[24] Nicholas Sidiropoulos, Lieven De Lathauwer, Xiao Fu, Kejun Huang, Evangelos Papalexakis,\nand Christos Faloutsos. Tensor decomposition for signal processing and machine learning.\nIEEE Transactions on Signal Processing, 2017.\n\n[25] Michel Talagrand. The generic chaining: upper and lower bounds of stochastic processes.\n\nSpringer Science &amp; Business Media, 2006.\n\n[26] Joel A Tropp. Improved analysis of the subsampled randomized hadamard transform. Advances\n\nin Adaptive Data Analysis, 3(01n02):115\u2013126, 2011.\n\n[27] Paul Tseng. Convergence of a block coordinate descent method for nondifferentiable minimiza-\n\ntion. Journal of Optimization Theory and Applications, 109(3):475\u2013494, 2001.\n\n[28] Paul Tseng and Sangwoon Yun. A coordinate gradient descent method for nonsmooth separable\n\nminimization. Mathematical Programming, 117(1-2):387\u2013423, 2009.\n\n[29] Roman Vershynin. Introduction to the non-asymptotic analysis of random matrices. arXiv\n\npreprint arXiv:1011.3027, 2010.\n\n[30] David P Woodruff. Sketching as a tool for numerical linear algebra. Foundations and Trends R\n\nin Theoretical Computer Science, 10(1\u20132):1\u2013157, 2014.\n\n[31] Yongxin Yang and Timothy Hospedales. Deep multi-task representation learning: A tensor\n\nfactorisation approach. arXiv preprint arXiv:1605.06391, 2016.\n\n[32] Rose Yu, Dehua Cheng, and Yan Liu. Accelerated online low-rank tensor learning for multivari-\n\nate spatio-temporal streams. In International Conference on Machine Learning, 2015.\n\n[33] Rose Yu and Yan Liu. Learning from multiway data: Simple and ef\ufb01cient tensor regression. In\n\nInternational Conference on Machine Learning, pages 373\u2013381, 2016.\n\n[34] Qibin Zhao, Cesar F Caiafa, Danilo P Mandic, Zenas C Chao, Yasuo Nagasaka, Naotaka Fujii,\nLiqing Zhang, and Andrzej Cichocki. Higher order partial least squares (hopls): a generalized\nmultilinear regression method. IEEE Transactions on Pattern Analysis and Machine Intelligence,\n35(7):1660\u20131673, 2013.\n\n[35] Hua Zhou. Matlab TensorReg toolbox.\n\ntensorreg/, 2013.\n\nhttp://hua-zhou.github.io/softwares/\n\n[36] Hua Zhou, Lexin Li, and Hongtu Zhu. Tensor regression with applications in neuroimaging\n\ndata analysis. Journal of the American Statistical Association, 108(502):540\u2013552, 2013.\n\n11\n\n\f", "award": [], "sourceid": 1972, "authors": [{"given_name": "Xingguo", "family_name": "Li", "institution": "University of Minnesota"}, {"given_name": "Jarvis", "family_name": "Haupt", "institution": "University of Minnesota"}, {"given_name": "David", "family_name": "Woodruff", "institution": "Carnegie Mellon University"}]}