{"title": "A Linearly Convergent Method for Non-Smooth Non-Convex Optimization on the Grassmannian with Applications to Robust Subspace and Dictionary Learning", "book": "Advances in Neural Information Processing Systems", "page_first": 9442, "page_last": 9452, "abstract": "Minimizing  a non-smooth function over the Grassmannian appears in many applications in machine learning. In this paper we show that if the objective satisfies a certain Riemannian regularity condition with respect to some point in the Grassmannian, then a Riemannian subgradient method with appropriate initialization and geometrically diminishing step size converges at a linear rate to that point. We show that for both the robust subspace learning method Dual Principal Component Pursuit (DPCP) and the Orthogonal Dictionary Learning (ODL) problem, the Riemannian regularity condition is satisfied with respect to appropriate points of interest, namely the subspace orthogonal to the sought subspace for DPCP and the orthonormal dictionary atoms for ODL. Consequently, we obtain in a unified framework significant improvements for the convergence theory of both methods.", "full_text": "A Linearly Convergent Method for Non-Smooth Non-Convex\n\nOptimization on the Grassmannian with Applications to\n\nRobust Subspace and Dictionary Learning\n\nZhihui Zhu\n\nMINDS\n\nTianyu Ding\n\nAMS\n\nJohns Hopkins University\n\nJohns Hopkins University\n\nShanghaiTech University\n\nzzhu29@jhu.edu\n\ntding1@jhu.edu\n\nmtsakiris@shanghaitech.edu.cn\n\nManolis C. Tsakiris\n\nSIST\n\nDaniel P. Robinson\n\nISE\n\nLehigh University\n\ndaniel.p.robinson@lehigh.edu\n\nRen\u00e9 Vidal\n\nMINDS\n\nJohns Hopkins University\n\nrvidal@jhu.edu\n\nAbstract\n\nMinimizing a non-smooth function over the Grassmannian appears in many appli-\ncations in machine learning. In this paper we show that if the objective satis\ufb01es a\ncertain Riemannian regularity condition (RRC) with respect to some point in the\nGrassmannian, then a projected Riemannian subgradient method with appropriate\ninitialization and geometrically diminishing step size converges at a linear rate\nto that point. We show that for both the robust subspace learning method Dual\nPrincipal Component Pursuit (DPCP) and the Orthogonal Dictionary Learning\n(ODL) problem, the RRC is satis\ufb01ed with respect to appropriate points of interest,\nnamely the subspace orthogonal to the sought subspace for DPCP and the orthonor-\nmal dictionary atoms for ODL. Consequently, we obtain in a uni\ufb01ed framework\nsigni\ufb01cant improvements for the convergence theory of both methods.\n\n1\n\nIntroduction\n\nOptimization problems on the Grassmannian G(c, D) (a.k.a. the Grassmann manifold that consists of\nthe set of linear c-dimensional subspaces in RD), such as principal component analysis (PCA), appear\nin a wide variety of applications including subspace tracking [3], system identi\ufb01cation [43], action\nrecognition [36], object categorization [20], dictionary learning [34, 39], robust subspace recovery\n[26, 42], subspace clustering [41], and blind deconvolution [50]. However, a key challenge is that the\nassociated optimization problems are often non-convex since the Grassmannian is a non-convex set.\nOne approach to solving optimization problems on the Grassmanian is to exploit the fact that the\nGrassmannian is a Riemannian manifold and develop generic Riemannian optimization techniques.\nWhen the objective function is twice differentiable, [4] shows that Riemannian gradient descent and\nRiemannian trust-region methods converge to \ufb01rst- and second-order stationary solutions, respectively.\nWhen Riemannian gradient descent is randomly initialized, [23] further shows that it converges to a\nsecond-order stationary solution almost surely, but without any guarantee on the convergence rate.\nNon-smooth trust region algorithms [19], gradient sampling methods [9, 8], and proximal gradient\nmethods [7] have also been proposed for non-smooth manifold optimization when the objective\nfunction is not continuously differentiable. However, the available theoretical results establish\n\n33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada.\n\n\fconvergence to stationary points from an arbitrary initialization with either no rate of convergence\nguarantee, or at best a sublinear rate.1\nOn the other hand, when the constraint set is convex, [11, 10, 27] show that subgradient methods\ncan handle non-smooth and non-convex objective functions as long as the problem satis\ufb01es certain\nregularity conditions called sharpness and weak convexity. In such a case, R-linear convergence1 is\nguaranteed (e.g., see robust phase retrieval [13] and robust low-rank matrix recovery [27]). Analogous\nto other regularity conditions for smooth problems, such as the regularity condition of [6] and the error\nbound condition in [29], sharpness and weak convexity capture regularity properties of non-convex\nand non-smooth optimization problems. However, these two properties have not yet been exploited\nfor solving problems on the Grassmannian, or other non-convex manifolds.\nA related regularity condition, which in this paper is called the Riemannian Regularity Condition\n(RRC), has already been exploited for orthogonal dictionary learning (ODL) [2], which solves an\n(cid:96)1 minimization problem on the sphere, a manifold parameterizing G(1, D). However, under this\nRRC, projected Riemannian subgradient methods have only been proved to converge at a sublinear\nrate. On the other hand, a projected subgradient method has been successfully used and proved to\nconverge at a piecewise linear rate for Dual Principal Component Pursuit (DPCP) [42, 51], a method\nthat \ufb01ts a linear subspace to data corrupted by outliers. However, i) the convergence analysis does not\nreveal where the improvement in the convergence rate comes from and ii) is restricted to optimization\non the sphere (G(1, D)) even for subspaces of codimension higher than one.\nIn this paper we make the following speci\ufb01c contributions:\n\u2022 In Theorem 1 we prove that the projected Riemannian subgradient method for the Grassmannian\n(Algorithm 1), with an appropriate initialization and geometrically diminishing step size, converges\nto a point of interest at an R-linear rate if the problem satis\ufb01es the RRC (De\ufb01nition 1). The\nRRC characterizes a local geometric property of the Grassmannian-constrained non-convex and\nnon-smooth problem relative to a point of interest. Informally, the RRC requires that, in the\nneighborhood of the point of interest, the negative of the Riemannian subgradient should have a\nsuf\ufb01ciently small angle with the direction pointing toward the point of interest (Figure 1).\n\n\u2022 We prove that the optimization problem associated with DPCP satis\ufb01es the RRC, which allows us\nto apply our new result and conclude that the projected Riemannian subgradient method converges\nat an R-linear rate to a basis for the orthogonal complement of the underlying (D \u2212 c)-dimensional\nlinear subspace. This is the \ufb01rst result to extend previous guarantees [42, 51, 12] from codimension\n1 to higher codimensions, enabling us to ef\ufb01ciently \ufb01nd the entire orthogonal basis by solving\nthe learning problem directly on G(c, D), as opposed to the less ef\ufb01cient approach of solving a\nsequence of c problems on G(1, D)[42]. Even for subspaces of codimension 1 (i.e., hyperplanes),\nour result improves upon [51] by allowing for (i) a much simpler step size selection strategy that\nrequires little \ufb01ne-tuning, and (ii) a weaker condition on the required initialization.\n\n\u2022 Together with the already established RRC for ODL in [2], our new result implies that the projected\nRiemannian subgradient method converges at an R-linear rate to atoms of the underlying dictionary,\nthus improving upon [2], which established only a sublinear convergence rate.\n\n2 Background and Notation\n\nIn this paper, we consider minimization problems on the Grassmannian G(c, D). For computations, it\nis desirable to parameterize points on the Grassmannian. An element of G(c, D) can be represented by\nan orthonormal matrix in O(c, D) =: {B \u2208 RD\u00d7c : B\nB = Ic}, which is the well-known Stiefel\nmanifold. When D = c, we denote O(c, c) by O(c), the orthogonal group. This matrix representation\nis not unique since Span(BQ) = Span(B) for any Q \u2208 O(c). Thus, we say A \u2208 G(c, D) is\nequivalent to B if Span(A) = Span(B). With this understanding, we use B to represent the\nequivalence class [B] = {BQ : Q \u2208 O(c)} and consider the parameterized problem [14, 20]\n\n(cid:62)\n\nminimize\nB\u2208O(c,D)\n\nf (B),\n\n(1)\n\n1Suppose the sequence {xk} converges to x(cid:63). We say it converges sublinearly if limk\u2192\u221e (cid:107)xk+1 \u2212\nx(cid:63)(cid:107)/(cid:107)xk \u2212 x(cid:63)(cid:107) = 1, and R-linearly if there exists C > 0, q \u2208 (0, 1) such that (cid:107)xk \u2212 x(cid:63)(cid:107) \u2264 Cqk, \u2200k \u2265 0.\n\n2\n\n\fwhere f : RD\u00d7c \u2192 R is locally Lipschitz, possibly non-convex and non-smooth, and invariant to the\naction of O(c), i.e., f (B) = f (BQ) for any Q \u2208 O(c). Again, the global minimum of (1) is not\nunique as if B(cid:63) is a global minimum, then any point in [B(cid:63)] is also a global minimum.\nFor any A, B \u2208 O(c, D), the principal angles between Span(A) and Span(B) are de\ufb01ned as [38]\nB)) for i = 1, . . . , c, where \u03c3i(\u00b7) denotes the i-th singular value. We\n\u03c6i(A, B) = arccos(\u03c3i(A\nthen de\ufb01ne the distance between A and B as\n\n(cid:62)\n\n(cid:118)(cid:117)(cid:117)(cid:116)2\n\nc(cid:88)\n\ni=1\n\n(cid:0)1 \u2212 cos(\u03c6i(A, B))(cid:1) = min\n\nQ\u2208O(c)(cid:107)B \u2212 AQ(cid:107)F ,\n\ndist(A, B) :=\n\n(2)\n\n(3)\n\nwhere the last term is also known as the orthogonal Procrustes problem. The last equality in (2)\nfollows from the result [21] according to which the optimal rotation matrix Q minimizing (cid:107)B \u2212\n(cid:62)\n(cid:62) is the SVD of A\nB. Thus, dist(A, B) = 0 iff Span(A) =\nAQ(cid:107)F is Q(cid:63) = U V\nSpan(B). We also de\ufb01ne the projection of B onto [A] as\n\n(cid:62), where U \u03a3V\n\nPA(B) = AQ(cid:63), where Q(cid:63) = arg min\n\nQ\u2208O(c) (cid:107)B \u2212 AQ(cid:107)F .\n(cid:62)\n\nHere, AQ(cid:63) is in [A], with Q(cid:63) representing a nonlinear transformation of A\nB, as described above.\nSince f can be non-smooth and non-convex, we utilize the Clarke subdifferential, which generalizes\nthe gradient for smooth functions and the subdifferential in convex analysis. The Clarke subdifferential\nof a locally Lipschitz function f at B is de\ufb01ned as [2]\n\n(cid:110)\ni\u2192\u221e\u2207f (Bi) : Bi \u2192 B, f differentiable at Bi\nlim\n\n(cid:111)\n\n,\n\n\u2202f (B) := conv\n\nwhere conv denotes the convex hull. When f is differentiable at B, its Clarke subdifferential is\nsimply {\u2207f (B)}. When f is not differentiable at B, the Clarke subdifferential is the convex hull of\nthe limit of gradients taken at differentiable points. Note that the Clarke subdifferential \u2202f (B) is a\nnonempty and convex set since a locally Lipschitz function is differentiable almost everywhere.\nSince we consider problems on the Grassmannian, we use tools from Riemannian geometry to\nstate optimality conditions. From [14], the tangent space of the Grassmannian at [B] is de\ufb01ned\nas TB := {W \u2208 RD\u00d7c : W\nB = 0}, and the orthogonal projector onto the tangent space is\nfor any A \u2208 [B]. We generalize the de\ufb01nition of the Clarke subdifferential and denote by(cid:101)\u2202f the\n(cid:62)\nI \u2212 BB\n= BB\n(cid:101)\u2202f (B) := conv\nRiemannian subdifferential of f [2]:\n(cid:62)\nWe say that B is a critical point of (1) if and only if 0 \u2208 (cid:101)\u2202f (B), which is a necessary condition for\n\n(cid:111)\n)\u2207f (Bi) : Bi \u2192 B, f differentiable at Bi \u2208 O(c, D)\n\n(cid:62), which is well-de\ufb01ned and does not depend on the class representative as AA\n\ni\u2192\u221e(I \u2212 BB\nlim\n\n(cid:110)\n\n.\n\n(4)\n\n(cid:62)\n\n(cid:62)\n\nbeing a minimizer to (1).\n\n3 Projected Riemannian Subgradient Method\n\nIn this section, we state our key Riemannian regularitity condition (RRC,\u00a73.1), propose a projected\nRiemannian subgradient method (\u00a73.2) based on RRC, and analyze its convergence properties (\u00a73.3).\n\n(\u03b1, \u0001, B(cid:63))-Riemannian Regularity Condition (RRC)\n\n3.1\nDe\ufb01nition 1. We say that f : RD\u00d7c \u2192 R satis\ufb01es the (\u03b1, \u0001, B(cid:63))-Riemannian regularity condi-\ndist(B, B(cid:63)) \u2264 \u0001, there exists a Riemannian subgradient G(B) \u2208 (cid:101)\u2202f (B) such that\ntion (RRC)2 for parameters {\u03b1, \u0001} > 0 and B(cid:63) \u2208 O(c, D), if for every B \u2208 O(c, D) satisfying\n\n(cid:104)PB(cid:63) (B) \u2212 B,\u2212G(B)(cid:105) \u2265 \u03b1 dist(B, B(cid:63)).\n\n(5)\n\n2Strictly speaking, De\ufb01nition 1 is extrinsic since we view the Grassmannian as embedded in the Euclidean\n\nspace and (5) involves the standard inner product in the Euclidean space.\n\n3\n\n\fRecently, a particular instance of De\ufb01nition 1 was shown\nto hold [2] in the context of ODL (see \u00a74.2). We note\nG(B) \u2208 (cid:101)\u2202f (B), and that the set of allowable Riemannian\nthat \u2212G(B) is not necessarily a descent direction for all\nnorm element from (cid:101)\u2202f (B) even though that one is known\nsubgradients that satisfy (5) need not include the minimum\n\n\u03be := sup{(cid:107)G(B)(cid:107)F : dist(B, B(cid:63)) \u2264 \u0001}\n\nto be a descent direction [18]. In \u00a74, we show that a natu-\nral choice of Riemannian subgradient satis\ufb01es (5) for DPCP\nand ODL, where B(cid:63) is a target solution. As illustrated in\nFigure 1, condition (5) implies that the negative of the cho-\nsen Riemannian subgradient G(B) has small angle to the\nFigure 1: Illustration of De\ufb01nition 1.\ndirection PB(cid:63) (B) \u2212 B. To see this, let\nRed nodes denote [B(cid:63)], with the top\none closest to B. Inequality (5) re-\n(6)\nquires the angle between PB(cid:63) (B) \u2212\ndenote an upper bound on the size of the Riemannian sub-\nB (purple arrow) and \u2212G(B) (blue\ngradients in a neighbrohood of B(cid:63). Assume \u03be < \u221e.\narrow) to be suf\ufb01ciently small.\nFrom (5) we have (cid:104)PB(cid:63) (B) \u2212 B,\u2212G(B)(cid:105)/(cid:107)PB(cid:63) (B) \u2212\nB(cid:107)F(cid:107)G(B)(cid:107)F \u2265 \u03b1/\u03be, which gives a bound on the sum of the cosines of the principal angles\nbetween PB(cid:63) (B) \u2212 B and \u2212G(B) and implies that \u03be \u2265 \u03b1.\nIn \u00a73.3 we prove that if the (\u03b1, \u0001, B(cid:63))-RRC holds, then a projected Riemannian subgradient method\nwill converge to B(cid:63) when an appropriate initialization and step size strategy are used.\nDe\ufb01nition 1 is similar in nature to other regularity conditions that characterize geometric properties of\nthe objective function. Perhaps the most closely related ones for non-smooth functions are sharpness\nand weak convexity. Consider a function h : RD \u2192 R and assume that the set\nof global minima of h is non-empty. Then, h is said to be sharp with parameter \u03bd > 0 (see [5]) if\n\nX := {z \u2208 RD : h(z) \u2264 h(x) for all x \u2208 Rn}\n\nh(x) \u2212 min\nz\u2208RD\n\nh(z) \u2265 \u03bd dist(x,X )\n\n(7)\nholds for all x \u2208 RD. The function h is said to be weakly convex with parameter \u03c4 \u2265 0 if\n2(cid:107)x(cid:107)2 is convex [44]. If h is both sharp and weakly convex, then [10, 27] show that\nx (cid:55)\u2192 h(x) + \u03c4\n(8)\nfor any x \u2208 RD and any d \u2208 \u2202h(x), where PX is the orthogonal projector onto the set X . Note that\n(8) is useful when its right-hand side is nonnegative, i.e., when dist(x,X ) \u2264 (2\u03bd)/\u03c4. Thus, for any\n\u0001 < (2\u03bd)/\u03c4, we have\n(9)\nwhenever x satis\ufb01es dist(x,X ) \u2264 \u0001. Noting the similarity between (9) and (5) (B(cid:63) can be taken as a\nminimizer of h), the RRC (5) can be viewed as a generalization of (9) (the consequence of sharpness\nand weak convexity) to the Riemannian manifold. There are two main differences. First, (5) differs\nfrom (9) in that its left-hand side involves the Riemannian subgradient due to the Grassmannian\nconstraint. Second, (5) is only required to hold for a particular Riemannian subgradient at B, while\n(9) holds for all subgradients, thus imposing a slightly stronger regularity condition on the problem.\n\n(cid:104)PX (x) \u2212 x, d(cid:105) \u2265 \u03bd dist(x,X ) \u2212 \u03c4\n(cid:104)PX (x) \u2212 x, d(cid:105) \u2265(cid:0)\u03bd \u2212 \u03c4\n\n2 \u0001(cid:1) dist(x,X ) for all d \u2208 \u2202h(x)\n\n2 dist2(x,X )\n\n3.2 Projected Riemannian Subgradient Method on the Grassmannian\n\nWe propose to solve (1) using the projected Riemannian subgradient method in Algorithm 1. Given\nthe kth iterate Bk, the next iterate Bk+1 is obtained by \ufb01rst moving in a direction opposite to a\nRiemannian subgradient at Bk that satis\ufb01es the regularity condition in (5), and then performing\northonormalization. In Section 4, we will show that such a projected Riemannian subgradient can\n\nbe easily computed for ODL and DPCP. Note that (cid:98)Bk+1 in (10) always has full column rank since\nmultiple ways to orthonormalize (cid:98)Bk+1, although for our purpose they are all equivalent since they all\nG(Bk) is orthogonal to Bk; see the supplementary material for a formal proof. Also, there are\nfor Span((cid:98)Bk+1). For example, one can compute Bk+1 to be the Gram-Schmidt orthonormalization\ncorrespond to the same subspace. In (10), orth refers to any method that \ufb01nds an orthonormal basis\nof (cid:98)Bk+1, or as the \ufb01rst c left singular vectors of (cid:98)Bk+1. Finally, note that no speci\ufb01c step size rule is\n\nprovided in Algorithm 1, whereas speci\ufb01c choices are made for the convergence analysis in \u00a73.3.\n\n4\n\nB<latexit sha1_base64=\"hyAWEBfAdvzJDhSRMMd5Vr0dovI=\">AAAmDnicrZpLc9y4EYBnN6+18/Imx1xYUVTlddlajTabVKW2tmxLsiRbj9FblulygSSGQ4sgaJJDzZjF/INU5ZT8k9xSueYv5IfkngbYnCHR0CaHzEEi+2uAQKPR3eCMl8ZRXqyv/+uTT7/3/R/88Eef3bv/45/89Gc/f/D5Ly5yOc18fu7LWGZXHst5HCX8vIiKmF+lGWfCi/mld7Op+GXJszySyVkxT/lbwcIkGkc+K0B06orn7x6srK+t649DL4Z4sTLAz+jd57/5txtIfyp4Uvgxy/M3w/W0eFuxrIj8mNf33WnOU+bfsJBX0yzuCd5MokAN9iZ/W01gPFnGx/0WTOSCFZPH8L+YCPUvnwtP/ffy+WNP9LU9U5CyjCm79aUzbam+zCIKM5ZOIn8GUrzMUxhLVa19OY7C/MvaGGocyiyCUd4hjnzziUIZzRCyVC1FX5hPPat8zBJ/PgnMmUQFmH3VFOUyM58Vc3AG0+I8mQpQh1ncdwM+Bk/SpqkClt2k8jbgmRdPeV1loVdX4BmPwT021J/fwUMrt+CzwtrAUR9Qcc4mXGZc5Lp7x4VLUMydyk1klARgEbj0xs5JI1+rXceFdm7Cb4umZaVsXGE39eqbnPvKNm+d1Yd7Y2cup84tg15Q24H5eDDNoNdFzIVgdbWv/3W7WFXDvIX1clCG7dd6zX3wlWpTZjKOWTbvdNDTSjOZ1tUI/so80st3h6K2c6NRbS2vv0P7f9GDx3v68bD9Rd+CsE+nYMNn6l/rWEv8qAoli6FpzBLT8o3RF9Y3jdqa9L6LwryYx7xqFrmvjbLqpGW9Fqlt2K1wOSnwp3HE4wC8KeOgqTYVS4LKPaorV0UOz6uO6qaXBTtZshPF0A9ZrNa+cpbM6cL8LniiWvZ77D5tc/m0zcXTgKYxn1nIZkOcPup2eL3s8Npkh0t2WKst3IUxeLwb83EBV0mog3MXZwpnUTgp4NLCYdO7Hg+jpOIfpjpjmE/nSoXD1Z0KCenjEe0kMXqhKh+6/SQsy9icdvNh2Y9dxQNLLfpRN1k0I900Oqqbpcaqg+uHwYF43wmv3wzfVq6EnMYKmSVMcHD0urF+tTJEO6un9VvuCUtLENpadtv5MnnftBQ3PEsgVX8tpq6EAK+KAZQ+0VLoo3NXdxsoj9nikMczfgBe9CxOJ8zjReNSY5/F1dHBaV35Ip/XFQSXRNlC7UrmRXFUzB0Yi5MXrDCnNbLMqnXUUd3M7mE7uS+MuW3Xd7Tcrk0zBNHYVA5MnZJlpo4S6WACU/Sn8ZQsaNCM/7F+cgYxuF4ZNg1kAvNNitx5eAtpvuCJEyVOJqHZYyca61TEgy8gH5tbCXvi1AfeL+n7GvOw45bcv5VZAEGoN/ZHFQDtk2qhWeZAwcG9jIHgHhDQchJZ6H1kmMEvmkl5EuInVFUyBvfSPTmOdnWf5/0mopjd3cSLpX/jNI809tniUY0fmW2afWXuzfZZvTb3oiThmXK5YOoXxvA0axopf/qm9advDQf4RoXBNsL1Tf+tCoGIwIAFXOeQvfljZ5dnUBdFalnTnE8D+SRKVC3NzS52l6u3a7re2ZKdmSyF/gAHLAxhFjDV8TQJmCoRIb1ACQjFJ1mQw2kcm76sZWbGgylxU7ERGpqnKUtMRS1TI2qlxijASj7pvRHSKFWailpGUlFyY+ppGdnuLCT7XckMvULy1NQ7UzJz9lsjMnkQGVqj0y1TS4kMrTENM2MdZnpaasuaalpm6OVRSNZFy0w9ixrVOrX0dkrVAnV0IuYFWVusxKGDn86OaiqmDsl6BCLEglgDmTkG2AK5Lo3NgYgmZuccoq9KPAUU2rEM54ZFeDf4eE0g6XtlLPNpRjzYh210DyJ8Q81QU/AskmTpQE6qDKn2MZwSDFWPJCbwXMGhX4tH41QhJOhCHyICgz/zPDK2YgCJJ6KmkoLUNWmkD7WmKshNVTiZS6uuAqYyJCE49CzsnfuZNjjYMYQoCafxNk21PalpSTiDiOgjJqlOAXKESo+gNzjtgI5aJXXGrBb39Z0t2Kzfor2/swXLQqGtp/Xh7o/q9jvV2aynzmaLvZEXxMHztWINDrMqtUg402RGVdDsXd2ZZc/mUeIvsLo2400sZSf7ubEWOCtDpy3qGw3D93kUdxup+24bzU3XyWv9N4rbg2NnXwTqbVI7kGbLCWf4DvP3k//bx9hkqji9ag8/ulK9IhtRiY96OuRoCMJnqKH4Mwt/3uHPLXyzw8nZDYRbHU6SBgi3O3zbwl90+AsL3+nwHQvf7XBSoIBwr8P3LPxlh7+08Fcd/srC9zt838IPOvzAwg87/NDCjzrctr6jDiepHYTHHX5s4ScdTg77IDzt8FMLP+twUgSC8LzDzy38osMvLPyywy8t/KrDyR4B4esOf23h1x2u3j/oj3HAqvXxomLk6OUh8QjxkZC4VgZI6EGOIyF1ZjlGMiYkREJqxHKCZEJIhISkxvI9kveE3CAhNWsZIyFVeimQkGxdJkhIViglEpKJyxQJqXLLD0g+EJIhIfVnmSPJCSmQkMqnnCKZElIiIZV/eYvklpAZEvKOppwjIW94yo9IPtLCsmTq7QZyt7kh7sqLhYa+Jh6IWRB12luL2n9TKaG+6HfWkRAnZuolK+o1N8TLmPCChQ7eEbdqV8hN6Bqlk9bvXXVJcL7EOcUBj5e2a26IVZYKNtvOFv3PaPdiMXIxtaztR561G2OdbpmkDRxDHb/67zh0/i9mNPeL50hI1hebSEi+F1tISKYX20hIjhcvkJDsLnaQkLwudpGQjC5eIiG5WrxCQrK02EdC8rM4QEIyszhEQnKyOEJCsrEYISF5WBwjIRlYnCAhuVecIiFZV5whIflWnCMhmVZcICE5VlwiIdlVXCEheVW8RkIyqrhGck3dWOzgPlfc3bHtc7GF+0yrbNn2mdhvQ4HW2beGAnEkeLjQaW7IaumAoBVGNCCIUb7ENCCI0yhczqW5IYuxCIBa6XwR/Uy77EFNoL8lJAWq2J4hov7eBAUNSVAQOihoNqQP9HRJ41lrGk8XDp61cvB0OvOs+czTCc2zZjRPpzSvzWlGiPL0LvSs29DTkcCzhgJP+6fXOqjRa84LFfrgHw19IHyOjAQ/EG4iI+EPhFvISAAE4TYyEgJB+AIZCYIg3EFGwiAId5GR2YNwDxnxGhC+REbcBoSvkJFACcJ9ZCRUgvAAGVklEB4iI+EShEfISMAE4QgZCZkgPEZGgiYIT5CRsAnCU2QkcILwDBkJnSA8R0aCJwgvkJHwCcJLZCSAgvAKGQmhIHyNjARREF4j00cS45V8nqutXeo36I+qJ0Mhaq0F+r6cqld51QGbHeivITZlDNob6+p9m5uykOP30bF6r3b/7p+C6AHFPAlB7EYivI2CYlLrZxBp9dXacOPrpr/uOI8vm9cmyx9RZFydfiAovVFfRTrHl3+onZXhW3P+iSy42bT5mcay7SHo2FtnvIzyO9rTt6U8iArrMLWq1vWkjDlLquJWgjkvml8cNbawo2rM4tz8vrlRqd9swKOicTGBJQUl9ULV1kXtqAHAnw31zbxh2JEcj3vH/nfN6KdppYg5xVGTgyzazZcL0P+7BytD82dR9OJiY2341drG8W9Xnj7Fn0x9NvjV4NeDh4Ph4PeDp4PdwWhwPvAH4eBPg78M/rr659W/rf599R+N6qefYJtfDnqf1X/+B9jw3gM=</latexit>G(B)<latexit sha1_base64=\"+JI18PAoMVk8whbhFR2ZPBDwsRs=\">AAAmFnicrZpLc9y4EYBnN6+18/Imx1xYUVTlddlajTabVKW2tmxLsmRbj9FbluhygSSGQ4sgaJJDzZjF/ItU5ZT8k9xSueaaH5J7GmBzhkRDmxwyB4nsrwECjUZ3gzNeGkd5sb7+r08+/d73f/DDH3127/6Pf/LTn/38wee/OM/lNPP5mS9jmV16LOdxlPCzIipifplmnAkv5hfezabiFyXP8kgmp8U85W8FC5NoHPmsANH1E9dn8c5DVzz/4t2DlfW1df1x6MUQL1YG+Bm9+/w3/3YD6U8FTwo/Znl+PVxPi7cVy4rIj3l9353mPGX+DQt5Nc3inuB6EgVq0Df522oC48oyPu63YCIXrJg8hv/FRKh/+Vx46r+Xzx97oq/tmYKUZUzZry+daYv1ZRZRmLF0EvkzkOJlnsJYqmrty3EU5l/WxlDjUGYRjPIOceSbTxTKaIaQpWpJ+sJ86lnlY5b480lgziQqwOyrpiiXmfmsmINTmBbnyVSAOszivhvwMXiUNk0VsOwmlbcBz7x4yusqC726As94DO6xof78Dh5auQWfFdYGjvqAinM64TLjItfdOy5cgmLuVG4ioyQAi8ClN3aOG/la7ToutHMTfls0LStl4wq7qVevc+4r27x1Vh++HDtzOXVuGfSC2g7Mx4NpBr0uYi4Eq6s9/a/bxaoa5i2sl4MybL/Wa+6Dr1SbMpNxzLJ5p4OeVprJtK5G8FfmkV6+OxS1nRuNamt5/R3a/4sePN7Tj4cwIPoWhH06BRs+U/9ax1riR1UoWQxNY5aYlm+MvrC+adTWpPddFObFPOZVs8h9bZRVxy3rtUhtw26Fy0mBP40jHgfgTRkHTbWpWBJU7mFduSpyeF51WDe9LNjxkh0rhn7IYrX2lbNkThfmd8Fj1bLfY/dpm8unbS6eBjSN+cxCNhvi9FG3w6tlh1cmO1iyg1pt4S6MwePdmI8LuEpCHZy7OFM4i8JJAZcWDpve9XgYJRX/MNWZw3w6Vyocru5USEgfj2gnidELVfnQ7SdhWcbmtJsPy37sKh5YatGPusmiGemm0VHdLDVWHVw/DA7E+455fT18W7kSchorZJYwwcHR68b61coQ7aye1m/5UlhagtDWstvOl8n7pqW44VkCqfprMXUlBHhVFKD0iZZCH527uttAecwWhzye8X3womdxOmEeLxqXGkOVUB3un9SVL/J5XUFwSZQt1K5kXhRHxdyBsTh5wQpzWiPLrFpHHdXN7B62k/vCmNt2fUfL7do0QxCNTeXA1ClZZuookQ4mMEV/Gk/JggbN+B/rJ2cQg+uVYdNAJjDfpMidh7eQ5gueOFHiZBKaPXaisU5FPPgC8rG5lbAnTn3g/ZK+rzEPO27J/VuZBRCEemN/VAHQPqkWmmUOFBzcyxgI7gEBLSeRhd5Hhhn8opmUJyF+QlUlY3Av3ZPjaFf3ed5vIorZ3U28WPo3TvNIY58tHtX4kdmm2Vfm3myf1WtzL0oSnimXC6Z+YQxPs6aR8qdvWn/61nCAb1QYbCNc3/TfqhCICAxYwHUO2Zs/dnZ5BnVRpJY1zfk0kE+iRNXU3Oxid7l6u6brnS7ZqclS6A9wwMIQZgFTHU+TgKkSEdILlIBQfJIFOZjGsenLWmZmPJgSNxUboaF5krLEVNQyNaJWaowCrOST3hshjVKlqahlJBUlN6aelpHtzkKy35XM0CskT029UyUzZ781IpMHkaE1OtkytZTI0BrTMDPWYaanpbasqaZlhl4ehWRdtMzUs6hRrRNLbydULVBHJ2JekLXFShw6+OnsqKZi6pCsRyBCLIg1kJljgC2Q69LYHIhoYnbOIfqqxFNAoR3LcG5YhHeDj9cEkr5XxjKfZsSDfdhG9yDCN9QMNQXPIkmWDuSkypBqH8MpwVD1SGICzxUc+rV4NE4VQoIu9CEiMPgzzyNjKwaQeCJqKilIXZNG+lBrqoLcVIWTubTqKmAqQxKCQ8/C3rmfaYODHUOIknAab9NU25OaloQziIg+YpLqFCCHqPQIeoPTDuioVVJnzGpxX9/Zgs36Ldr7O1uwLBTaelof7v6obr9Tnc166my22Bt5QRw8XyvW4DCrUouEM01mVAXN3tWdWfZsHiX+AqtrM97EUnaynxtrgbMydNqivtEwfJ9HcbeRuu+20dx0nbzWf6O4PTh29kWg3iq1A2m2nHCG7zB/P/m/fYxNporTy/bwoyvVS7IRlfiwp0OOhiB8hhqKP7Pw5x3+3MI3O5yc3UC41eEkaYBwu8O3LfxFh7+w8J0O37Hw3Q4nBQoIX3b4Swt/1eGvLPx1h7+28L0O37Pw/Q7ft/CDDj+w8MMOt63vqMNJagfhUYcfWfhxh5PDPghPOvzEwk87nBSBIDzr8DMLP+/wcwu/6PALC7/scLJHQPimw99Y+FWHq/cP+mMcsGp9vKgYOXp5SDxCfCQkrpUBEnqQ40hInVmOkYwJCZGQGrGcIJkQEiEhqbF8j+Q9ITdISM1axkhIlV4KJCRblwkSkhVKiYRk4jJFQqrc8gOSD4RkSEj9WeZIckIKJKTyKadIpoSUSEjlX94iuSVkhoS8oynnSMgbnvIjko+0sCyZeruB3G1uiLvyYqGhr4kHYhZEnfbWovbfVEqoL/qddSTEiZl6yYp6zQ3xMia8YKGDd8St2hVyE7pG6aT1e1ddEpwvcU5xwOOl7ZobYpWlgs22s0X/M9q9WIxcTC1r+5Fn7cZYp1smaQPHUMev/jsOnf+LGc394jkSkvXFJhKS78UWEpLpxTYSkuPFCyQku4sdJCSvi10kJKOLV0hIrhavkZAsLfaQkPws9pGQzCwOkJCcLA6RkGwsRkhIHhZHSEgGFsdISO4VJ0hI1hWnSEi+FWdISKYV50hIjhUXSEh2FZdISF4Vb5CQjCqukFxRNxY7uM8Vd3ds+1xs4T7TKlu2fSb22lCgdfasoUAcCh4udJobslo6IGiFEQ0IYpQvMQ0I4iQKl3NpbshiLAKgVjpbRD/TLi+hJtDfEpICVWzPEFF/b4KChiQoCB0UNBvSB3q6pPGsNY2nCwfPWjl4Op151nzm6YTmWTOap1Oa1+Y0I0R5ehd61m3o6UjgWUOBp/3Tax3U6DXnhQp98I+GPhA+R0aCHwg3kZHwB8ItZCQAgnAbGQmBIHyBjARBEO4gI2EQhLvIyOxB+BIZ8RoQvkJG3AaEr5GRQAnCPWQkVIJwHxlZJRAeICPhEoSHyEjABOEIGQmZIDxCRoImCI+RkbAJwhNkJHCC8BQZCZ0gPENGgicIz5GR8AnCC2QkgILwEhkJoSB8g4wEURBeIdNHEuOVfJ6rrV3qN+iPqidDIWqtBfq+nKpXedU+m+3rryE2ZQzaG+vqfZubspDj99Gxeq92/+6fgugBxTwJQexGIryNgmJS62cQafXV2nDj66a/7jiPLprXJssfUWRcnX4gKF2rryKdo4s/1M7K8K05/0QW3Gza/Exj2fYAdOytM15G+R3t6dtSHkSFdZhaVet6UsacJVVxK8Gc580vjxpb2FE1ZnFuft/cqNTXG/CoaFxMYElBSb1QtXVRO2oA8GdDfTNvGHYkx+Pesf9dM/ppWiliTnHU5CCLdvPlAvT/7sHK0PxZFL0431gbfrW2cfTbladP8SdTnw1+Nfj14OFgOPj94OlgdzAanA38gRz8afCXwV9X/7z6t9W/r/6jUf30E2zzy0Hvs/rP/wBxSeCk</latexit>PB?(B)<latexit sha1_base64=\"kKZ9oj5kOCvCxqwMVSf/JawnXQs=\">AAAmI3icrZpLb9zIEYBnN6+18/ImQC65EFEEeA1bq9FmEyBYLGxLsiRbj9FblqgITbKHQ4tN0nxpxgzzYwLklPyT3IJccsjPyD3VzeIM2dXazSFzkMj6qpvd1dVV1ZxxkjDI8tXVf3308Xe++73v/+CTBw9/+KMf/+Snjz792VkWF6nLT904jNMLh2U8DCJ+mgd5yC+SlDPhhPzcuV2X/LzkaRbE0Uk+S/i1YH4UjAOX5SC6efQL22Xh6Kayxcs/2FnO0voxXH5282hpdWVVfSx6McSLpQF+Rjef/vo/the7heBR7oYsy66Gq0l+XbE0D9yQ1w/tIuMJc2+Zz6siDXuCq0ngyRncZtfVBAaZpnzcb8FEJlg+eQr/84mQ/7KZcOR/J5s9dURf29EFCUuZNGZfOlXm68sMIj9lySRwpyDFyyyBsVTVyufjwM8+r7Whhn6cBjDKe8SBqz9RSKNpQpbI9ekLs8IxyscscmcTT59JkIPZl3VRFqf6s0IOHqJbnEeFAHWYxUPb42NwL2WaymPpbRLfeTx1woLXVeo7dQWe8RTcY03++S08tLJzPs2NDSz5ARXrZMLjlItMdW/ZcAmKmVXZURxEHlgELp2xddTIV2rbsqGdHfG7vGlZSRtX2E29fJVxV9rm2lp+vDO2ZnFh3THoBbUtmI8D0/R6XYRcCFZXu+pft4tlOcw7WC8LZdh+pdfcBV+p1uM0DkOWzjod9LSSNE7qagR/4yxQy3ePorJzo1FtLK6/Qft/0YPHO+rxEBNE34KwTwuw4Qv5r3WsBX5S+TELoWnIIt3yjdHn1teN2pr0oY3CLJ+FvGoWua+NsuqoZb0WiWnYrXAxKfCnccBDD7wp5aApNxWLvMo+qCG0wW51nOqgbnqZs6MFO5IM/ZCFcu0ra8GsLszug0eyZb/H7tPWF09bnz8NaBLyqYGsN8Tqo26Hl4sOL3W2v2D7tdzCXRiCx9shH+dwFfkqOHdxKnEa+JMcLg0cNr3tcD+IKv6+UGlEfzqXKhyu7lWISB9PaCeR1gtVed/tJ2Jpyma0m/eLfswqDlhq3o+8SYMp6abRkd0sNJYtXD8MDsT7jnh9Nbyu7BhyGsvjNGKCg6PXjfWrpSHaWT6t33JHGFqC0NSy286No3dNS3HL0whS9ZeisGMI8LJCQOkzJYU+Ond1t4H0mA0OeTzle+BFL8JkwhyeNy41hoqhOtg7ritXZLO6guASSVvIXcmcIAzymQVjsaCayPVpjQyzah11VDeze9xO7jNtbpv1PS03a90MXjDWlT1dp4RiR9ORIhVMYIpuERZkQb1m/E/Vk1OIwfXSsGkQRzDfKM+sx3eQ5nMeWUFkpTE0e2oFY5WKuPcZ5GN9K2FPnPrAuwV9V2MetuySu3dx6kEQ6o39SQVA+aRcaJZaUHBwJ2UgeAAEtKwoztU+0szg5s2knBjiJ1RVcQjupXqyLOXqLs/6TUQ+vb+JE8burdU8Uttn80c1fqS3afaVvjfbZ/XaPAiiiKfS5bzCzbXhKdY0kv70VetPX2sO8JUMg22E65v+axkCEYEBc7jOIHvzp9Y2T6EuCuSyJhkvvPhZEMkCm+tdbC9Wb1t3vZMFO9FZAv0B9pjvwyxgquMi8pgsESG9QAkIxSdZkP0iDHVfVjI948GUuK7YCDXN44RFuqKSyRG1Um0UYCWX9N4IaZQqdUUlI6koutX1lIxsd+aT/S5lml4e80TXO5EyffYbIzJ5EGlao+MNXUuKNK0xDTNjFWZ6WnLL6mpKpullgU/WRcl0PYMa1To29HZM1Tx5dCLmBVlbrIS+hZ/Ojmoqpg5JewQixJwYA5k+BtgCmSqN9YGIJmZnHKKvTDw5FNph7M80i/Bu8HGaQNL3yjDOipR4sAvb6AFE+IbqoSbnaRCTpQM5qTJiuY/hlKCpOiQxgecKDv0aPBqnCiFBFfoQERj8mWWBthU9SDwBNVUsSF2TBOpQq6uCXFeFk3ls1JVAV4YkBIeeub0zN1UGBzv6ECXhNN6mqbYnOa0YziAi+IBJqlOAHKDSE+gNTjugI1dJnjGr+X19bws27bdo7+9twVJfKOspfbj7k7z9RnU27amz6XxvZDlx8GwlX4HDrEwtMZxpUq0qaPau6sywZ7MgcudYXuvxJozjTvazQyWwloZWW9Q3Gprv8yDsNpL33TaK666T1epvELYHx86+8OQrpnYgzZYT1vAG8/ez/9tH22SyOL1oDz+qUr0gG1GKD3o65GgIwheoIfkLA3/Z4S8NfL3DydkNhBsdTpIGCDc7fNPAX3X4KwPf6vAtA9/ucFKggHCnw3cM/HWHvzbwNx3+xsB3O3zXwPc6fM/A9zt838APOty0vqMOJ6kdhIcdfmjgRx1ODvsgPO7wYwM/6XBSBILwtMNPDfysw88M/LzDzw38osPJHgHh2w5/a+CXHS7fP6iPdsCq1fGiYuTo5SBxCHGRkLhWekjoQY4jIXVmOUYyJsRHQmrEcoJkQkiAhKTG8h2Sd4TcIiE1axkiIVV6KZCQbF1GSEhWKGMkJBOXCRJS5ZbvkbwnJEVC6s8yQ5IRkiMhlU9ZICkIKZGQyr+8Q3JHyBQJeUdTzpCQNzzlByQfaGFZMvl2A7nd3BB35flcQ10TD8QsiDrtrUHt21RKqC/6nXUkxImZfMmKes0N8TImHG+ug3fErdoVsiO6Rsmk9XtbXhKcLXBGscfDhe2aG2KVhYLJttN5/1PavZiPXBSGtf3A03ZjrNItE7WBY6jiV/8dh8r/+ZTmfvESCcn6Yh0JyfdiAwnJ9GITCcnx4hUSkt3FFhKS18U2EpLRxWskJFeLN0hIlha7SEh+FntISGYW+0hIThYHSEg2FiMkJA+LQyQkA4sjJCT3imMkJOuKEyQk34pTJCTTijMkJMeKcyQku4oLJCSvirdISEYVl0guqRuLLdznkttbpn0uNnCfKZUN0z4Tu20oUDq7xlAgDgT35zrNDVktFRCUwogGBDHKFpgGBHEc+Iu5NDdkMeYBUCmdzqOfbpcdqAnUt4SkQBWbU0TU35ugoCAJCkIFBcWG9IGOKmkcY03jqMLBMVYOjkpnjjGfOSqhOcaM5qiU5rQ5TQtRjtqFjnEbOioSOMZQ4Cj/dFoH1XrNeC5DH/yjoQ+EL5GR4AfCdWQk/IFwAxkJgCDcREZCIAhfISNBEIRbyEgYBOE2MjJ7EO4gI14DwtfIiNuA8A0yEihBuIuMhEoQ7iEjqwTCfWQkXILwABkJmCAcISMhE4SHyEjQBOERMhI2QXiMjAROEJ4gI6EThKfISPAE4RkyEj5BeI6MBFAQXiAjIRSEb5GRIArCS2TqSKK9ks8yubVL9Qb9SfVsKESttEDfjQv5Kq/aY9M99TXEehyC9tqqfN9mJ8zn+H10KN+rPbz/pyBqQCGPfBDbgfDvAi+f1OoZRFp9sTJc+7LprzvOw/PmtcniRxQpl6cfCEpX8qtI6/D897W1NLzW5x/FOdebNj/TWLTdBx1z65SXQXZPe/q2lHtBbhymUlW6ThyHnEVVfheDOc+anyE1tjCjaszCTP++uVGpr9bgUcE4n8CSgpJ8oWrqorbkAODPmvxmXjPsKB6Pe8f+m2b0RVJJok9x1OQgg3bz5QL0f/Noaaj/LIpenK2tDL9YWTv8zdLz5/iTqU8Gvxz8avB4MBz8bvB8sD0YDU4H7uCPgz8P/jr42/Jflv++/I/lfzaqH3+EbX4+6H2W//1fjxbl9Q==</latexit>B?<latexit sha1_base64=\"H2K9ECUuwiUKsam/5pmddRQWDzU=\">AAAmFHicrZpLc9y4EYBnN6+18/Imx1xYUVTlddlajTabVKW2ttaWZEm2HqO3LNNxgSSGQ4sgaZBDzZjF/IlU5ZT8k9xSueaeH5J7GmBzhkRDmz1kDhLZXwMEGo3uBme8LI7yYn393x99/L3v/+CHP/rk3v0f/+SnP/v5g09/cZGnU+nzcz+NU3nlsZzHUcLPi6iI+VUmORNezC+9m03FL0su8yhNzop5xt8IFibROPJZAaJXrnj2RzcvmHz7YGV9bV1/HHoxxIuVAX5Gbz/9zX/cIPWngieFH7M8fz1cz4o3FZNF5Me8vu9Oc54x/4aFvJrKuCd4PYkCNeSb/E01gVFJycf9FkzkghWTx/C/mAj1L58LT/338vljT/S1PVOQMcmU9frSmbZXX2YRhZJlk8ifgRQv8wzGUlVrn4+jMP+8NoYah6mMYJR3iCPffKJQRjOELFML0hfmU88qH7PEn08CcyZRAWZfNUV5Ks1nxRxcwrQ4T6YC1GEW992Aj8GftGmqgMmbLL0NuPTiKa8rGXp1BZ7xGNxjQ/35HTy0cgs+K6wNHPUBFedswlPJRa67d1y4BMXcqdwkjZIALAKX3tg5aeRrteu40M5N+G3RtKyUjSvspl59nXNf2eaNs/pwb+zM06lzy6AX1HZgPh5MM+h1EXMhWF3t63/dLlbVMG9hvRyUYfu1XnMffKXaTGUax0zOOx30tDKZZnU1gr9pHunlu0NR27nRqLaW19+i/V304PGefjwEAdG3IOzTKdjwqfrXOtYSP6rClMXQNGaJafnG6Avrm0ZtTXrfRWFezGNeNYvc10ZZddKyXovMNuxWuJwU+NM44nEA3iQ5aKpNxZKgco/qylWRw/Oqo7rpZcFOluxEMfRDFqu1r5wlc7owvwueqJb9HrtP21w+bXPxNKBZzGcWstkQp4+6HV4vO7w22eGSHdZqC3dhDB7vxnxcwFUS6uDcxVJhGYWTAi4tHDa96/EwSir+fqrzhvl0rlQ4XN2pkJA+HtFOEqMXqvK+20/CpGRz2s37ZT92FQ8stehH3choRrppdFQ3S41VB9cPgwPxvhNevx6+qdwUchorUpkwwcHR68b61coQ7aye1m+5JywtQWhr2W3np8m7pqW44TKBVP2lmLopBHhVEqD0iZZCH527uttAecwWhzwu+QF40dM4mzCPF41LjX0WV0cHp3Xli3xeVxBcEmULtSuZF8VRMXdgLA7UEYU5rZFlVq2jjupmdg/byX1mzG27vqPldm2aIYjGpnJg6pRMmjpKpIMJTNGfxlOyoEEz/sf6yRJicL0ybBqkCcw3KXLn4S2k+YInTpQ4MoVmj51orFMRDz6DfGxuJeyJUx94t6TvaszDjlty/zaVAQSh3tgfVQC0T6qFZtKBgoN7koHgHhDQcpK00PvIMINfNJPyUoifUFWlMbiX7slxtKv7PO83EcXs7iZenPo3TvNIY58tHtX4kdmm2Vfm3myf1WtzL0oSLpXLBVO/MIanWdNI+dNXrT99bTjAVyoMthGub/qvVQhEBAYs4DqH7M0fO7tcQl0UqWXNcj4N0idRoipqbnaxu1y9XdP1zpbszGQZ9Ac4YGEIs4CpjqdJwFSJCOkFSkAoPsmCHE7j2PRlLTMzHkyJm4qN0NA8zVhiKmqZGlErNUYBVvJJ742QRqnSVNQykoqSG1NPy8h2ZyHZ70pm6BUpz0y9MyUzZ781IpMHkaE1Ot0ytZTI0BrTMDPWYaanpbasqaZlhl4ehWRdtMzUs6hRrVNLb6dULVBHJ2JekLXFShw6+OnsqKZi6hDZIxAhFsQayMwxwBbIdWlsDkQ0MTvnEH1V4img0I7TcG5YhHeDj9cEkr5Xxmk+lcSDfdhG9yDCN9QMNQWXUUqWDuSkykjVPoZTgqHqkcQEnis49GvxaJwqhARd6ENEYPBnnkfGVgwg8UTUVKkgdU0W6UOtqQpyUxVO5qlVVwFTGZIQHHoW9s59qQ0OdgwhSsJpvE1TbU9qWimcQUT0AZNUpwA5QqVH0BucdkBHrZI6Y1aL+/rOFmzWb9He39mCyVBo62l9uPuTuv1WdTbrqbPZYm/kBXHwfK1Yg8OsSi0pnGmkURU0e1d3ZtmzeZT4C6yuzXgTp2kn+7mxFjgrQ6ct6hsNw/d5FHcbqftuG81N18lr/TeK24NjZ18E6p1SO5Bmywln+Bbz95P/28fYZKo4vWoPP7pSvSIbUYmPejrkaAjCp6ih+FMLf9bhzyx8s8PJ2Q2EWx1OkgYItzt828Kfd/hzC9/p8B0L3+1wUqCAcK/D9yz8RYe/sPCXHf7Swvc7fN/CDzr8wMIPO/zQwo863La+ow4nqR2Exx1+bOEnHU4O+yA87fBTCz/rcFIEgvC8w88t/KLDLyz8ssMvLfyqw8keAeGrDn9l4dcdrt4/6I9xwKr18aJi5OjlIfEI8ZGQuFYGSOhBjiMhdWY5RjImJERCasRygmRCSISEpMbyHZJ3hNwgITVrGSMhVXopkJBsXSZISFYoUyQkE5cZElLllu+RvCdEIiH1Z5kjyQkpkJDKp5wimRJSIiGVf3mL5JaQGRLyjqacIyFveMoPSD7QwrJk6u0Gcre5Ie7Ki4WGviYeiFkQddpbi9r/Uimhvuh31pEQJ2bqJSvqNTfEy5jwgoUO3hG3alfITegaZZPW7111SXC+xDnFAY+XtmtuiFWWCjbbzhb9z2j3YjFyMbWs7Qcu242xTrdM0gaOoY5f/XccOv8XM5r7xTMkJOuLTSQk34stJCTTi20kJMeL50hIdhc7SEheF7tISEYXL5CQXC1eIiFZWuwjIflZHCAhmVkcIiE5WRwhIdlYjJCQPCyOkZAMLE6QkNwrTpGQrCvOkJB8K86RkEwrLpCQHCsukZDsKq6QkLwqXiEhGVVcI7mmbix2cJ8r7u7Y9rnYwn2mVbZs+0zst6FA6+xbQ4E4Ejxc6DQ3ZLV0QNAKIxoQxChfYhoQxGkULufS3JDFWARArXS+iH6mXfagJtDfEpICVWzPEFF/b4KChiQoCB0UNBvSB3q6pPGsNY2nCwfPWjl4Op151nzm6YTmWTOap1Oa1+Y0I0R5ehd61m3o6UjgWUOBp/3Tax3U6DXnhQp98I+GPhA+Q0aCHwg3kZHwB8ItZCQAgnAbGQmBIHyOjARBEO4gI2EQhLvIyOxBuIeMeA0IXyAjbgPCl8hIoAThPjISKkF4gIysEggPkZFwCcIjZCRggnCEjIRMEB4jI0EThCfISNgE4SkyEjhBeIaMhE4QniMjwROEF8hI+AThJTISQEF4hYyEUBC+QkaCKAivkekjifFKPs/V1i71G/RH1ZOhELXWAn0/napXedUBmx3oryE20xi0N9bV+zY3YyHH76Nj9V7t/t0/BdEDinkSgtiNRHgbBcWk1s8g0uqLteHGl01/3XEeXzavTZY/opBcnX4gKL1WX0U6x5d/qJ2V4Rtz/klacLNp8zONZdtD0LG3lryM8jva07elPIgK6zC1qtb10jTmLKmK2xTMedH87qixhR1VYxbn5vfNjUr9egMeFY2LCSwpKKkXqrYuakcNAP5sqG/mDcOO0vG4d+x/24x+mlWKmFMcNTnIot18uQD9v32wMjR/FkUvLjbWhl+sbRz/duWbb/AnU58MfjX49eDhYDj4/eCbwe5gNDgf+AMx+PPgr4O/rf5l9e+r/1j9Z6P68UfY5peD3mf1X/8FATHgsw==</latexit>\fAlgorithm 1 Projected Riemannian Subgradient Method\nInitialization: set B0 and \u00b50;\n1: for k = 0, 1, . . . do\n2:\n3:\n4:\n\nobtain G(Bk) \u2208 (cid:101)\u2202f (Bk) satisfying (5) with B = Bk;\nupdate the iterate: (cid:98)Bk+1 \u2190 Bk \u2212 \u00b5kG(Bk) and Bk+1 \u2190 orth((cid:98)Bk+1);\n\ncompute a step size \u00b5k according to a certain rule;\n\n5: end for\n\n(10)\n\n3.3 Convergence Analysis\n\nOur convergence analysis for Algorithm 1 relies in the RRC of De\ufb01nition 1. When this regularity\ncondition holds, we show that the iterates of Algorithm 1 exhibit the following properties: (i) they\nconverge to a neighborhood of the set B(cid:63) when a constant step size is used, and (ii) they converge at\nan R-linear rate to B(cid:63) when a geometrically diminishing step size is used.\n\n3.3.1 Constant step size\n\nWe \ufb01rst consider the convergence of Algorithm 1 when a constant step size is used.\nProposition 1. Suppose that for some (\u03b1, \u0001, B(cid:63)) the function f satis\ufb01es the (\u03b1, \u0001, B(cid:63))-RRC in\nDe\ufb01nition 1. Let {Bk} be generated by Algorithm 1 with step size \u00b5k \u2261 \u00b5 \u2264 \u03b1\u0001/\u03be2 and initial\niterate B0 satisfying dist(B0, B(cid:63)) \u2264 \u0001, where \u03be is de\ufb01ned in (6). Then, for all k \u2265 0, it holds that\n(11)\n\n(cid:110)\n\n(cid:111)\n\ndist(Bk, B(cid:63)) \u2264 max\n\ndist(B0, B(cid:63)) \u2212 \u00b5\u03b1k/2, \u00b5\u03be2/\u03b1\n\n.\n\nTowards interpreting Proposition 1, \ufb01rst consider the case dist(B0, B(cid:63)) > \u00b5\u03be2/\u03b1, in which\ncase (11) implies that after at most K = 2(dist(B0, B(cid:63)) \u2212 \u00b5\u03be2/\u03b1)/(\u00b5\u03b1) iterates, the inequal-\nity dist(Bk, B(cid:63)) \u2264 \u00b5\u03be2/\u03b1 will hold for all k \u2265 K. In that sense, Proposition 1 essentially says that\nno further decay of dist(Bk, B(cid:63)) can be guaranteed. This agrees with empirical evidence regarding\nAlgorithm 1 with constant step size (see Section 4). Note that (11) also suggests a tradeoff in selecting\nthe step size \u00b5. A larger step size \u00b5 leads to a faster decrease on the bound but a larger universal\nupper bound of \u00b5\u03be2/\u03b1, which may even exceed dist(B0, B(cid:63)) if \u00b5 is too large.\n\n3.3.2 Geometrically diminishing step size\n\nA useful strategy to balance the tradeoff discussed in the previous paragraph is to use a diminishing\nstep size that starts relatively large and decreases to zero as the iterates proceed. As the universal upper\nbound \u00b5\u03be2\n\u03b1 in (11) is proportional to \u00b5, it is expected that the decay rate of the step size will determine\nthe convergence rate of the iterates. In this section, we consider a geometrically diminishing step\nsize scheme, i.e., we decrease the step size by a \ufb01xed fraction between iterations. Our argument is\ninspired by [10, 27]. Convergence with geometrically diminishing step size is guaranteed by the\nfollowing result, which shows that if we choose the decay rate and initial step size properly, then the\nprojected Riemannian subgradient method converges to B(cid:63) at an R-linear rate.\nTheorem 1. Suppose that f satis\ufb01es the (\u03b1, \u0001, B(cid:63))-RRC in De\ufb01nition 1. Let {Bk} be the sequence\ngenerated by Algorithm 1 with step size\n\n\u00b5k = \u00b50\u03b2k\n\n(12)\n\nand initialization B0 satisfying dist(B0, B(cid:63)) \u2264 \u0001. Assume\n\n(cid:115)\n\n\u03b1 dist(B0, B(cid:63))\n\n2\u03be2\n\n\u00b50 \u2264\n\nand\n\n1 \u2212 2\n\n\u03b1\u00b50\n\ndist(B0, B(cid:63))\n\n+\n\n\u00b52\n0\u03be2\n\ndist2(B0, B(cid:63))\n\n=: \u03b2 \u2264 \u03b2 < 1,\n\n(13)\n\nwhere \u03be is de\ufb01ned in (6). Then, the sequence {Bk} satis\ufb01es\n\ndist(Bk, B(cid:63)) \u2264 dist(B0, B(cid:63))\u03b2k for all k \u2265 0.\n\n(14)\n\n5\n\n\fThe rate at which {dist(Bk, B(cid:63))}k\u22650 tends to zero in (14) is determined by \u03b2 which satis\ufb01es\n(13). Note that \u03b2 is well de\ufb01ned and is strictly less than 1 in (13). To see this, on one hand,\n\u00b50 \u2264 \u03b1 dist(B0, B(cid:63))/2\u03be2 and \u03be \u2265 \u03b1 together imply 1 \u2212 2\u03b1\u00b50/dist(B0, B(cid:63)) \u2265 0. On the\n(cid:112)\n0\u03be2/dist2(B0, B(cid:63)) < 0 is a decreasing function of \u00b50 when\nother hand, \u22122\u03b1\u00b50/dist(B0, B(cid:63)) + \u00b52\nIn particular, when \u00b50 = \u03b1 dist(B0, B(cid:63))/2\u03be2, we have \u03b2 =\n\u00b50 \u2208 (0, \u03b1 dist(B0, B(cid:63))/2\u03be2].\n1 \u2212 3\u03b12/4\u03be2, giving the fastest decaying rate by setting \u03b2 = \u03b2. Finally, if dist(B0, B(cid:63)) is not\nknown a priori, then one can replace it by its upper boud \u0001 in (13) and (14) and the results still hold.\n\n4 Applications\n\nIn this section, we show that Algorithm 1 achieves an R-linear convergence rate for Dual Principal\nComponent Pursuit (DPCP) [42, 51] and Orthogonal Dictionary Learning (ODL) [39, 2].\n\n4.1 DPCP for Robust Subspace Learning\n\nWe begin with the problem of learning a subspace from data corrupted by outliers [25]. Important\nmethods include Random Sampling And Consensus (RANSAC) [17], fast median subspace [24],\ngeodesic gradient descent [31], coherence pursuit [35], and many that solve convex formulations\nbased on (cid:96)1 and nuclear norm optimization [46, 49, 26, 37, 48] , but require either the dimension of the\nsubspace or the number of outliers to be suf\ufb01ciently small. On the other hand, DPCP [40, 41, 42, 51]\nsolves a non-convex problem, can provably handle subspaces of high dimension, and can provably\ntolerate as many outliers as the square of the number of inliers. DPCP has been successfully applied\nin three-view geometry problems [42] and road plane detection from 3D point cloud data [51, 12],\nhas been shown to outperform RANSAC, and has been applied in the multiple-hyperplane case [41].\nThe main principle behind DPCP is the computation of a basis for the orthogonal complement\n\nof the subspace to be learned. Speci\ufb01cally, given a dataset (cid:101)X = [X O]\u0393 \u2208 RD\u00d7L, where the\ncolumns of X \u2208 RD\u00d7N are inlier points spanning a d-dimensional subspace S of RD, the columns\nof O \u2208 RD\u00d7M are outlier points, and \u0393 is an unknown permutation, DPCP solves\n\nf (B) := (cid:107)(cid:101)X (cid:62)\n\nB(cid:107)1,2 \u2261\n\nminimize\nB\u2208O(c,D)\n\nL(cid:88)\n\ni=1\n\n(cid:107)(cid:101)x\n\n(cid:62)\ni B(cid:107)2,\n\n(15)\n\nwhere c = D \u2212 d is the codimension of S. An iterative reweighted least squares (IRLS) algorithm\nhas been empirically utilized to solve (15) in [42], but without formal guarantees.\nVeri\ufb01cation of the regularity condition. We will show that the DPCP problem (15) satis\ufb01es the\nRRC, which will then be used to establish convergence rates. Since the objective function f in (15)\n)\u2202f (B). Also note that the (cid:96)2 norm is\n\nis regular, it follows from [47] that (cid:101)\u2202f (B) = (I \u2212 BB\n\nsubdifferentially regular, thus by the chain rule one choice for the Riemannian subgradient is\n\n(cid:62)\n\nG(B) = (I \u2212 BB\n\n(cid:62)\n\n)\n\nL(cid:88)\n\ni=1\n\n(cid:101)xi sign((cid:101)x\n\n(cid:62)\ni B), where sign(a) :=\n\nif a (cid:54)= 0,\nif a = 0.\n\n(16)\n\n(cid:26)a/(cid:107)a(cid:107)2\n)(cid:80)M\n\n0\n\n(cid:13)(cid:13)(I\u2212BB\n\n(cid:13)(cid:13)F\n\n(cid:62)\n\nM maxB\u2208O(c,D)\n\ni=1 oi sign(o(cid:62)\n\nTo analyze (15), we de\ufb01ne two quantities. The \ufb01rst one characterizes the maximum Riemannian\nsubgradient related to outliers: \u03b7O := 1\n, which\nappears in [51] when B is on O(1, D). The second one is related to the inliers and is given\nN minb\u2208S\u2229O(1,D) (cid:107)X (cid:62)b(cid:107)1, which is referred to as the permeance statistic in [26].\nby cX ,min := 1\nThese quantities re\ufb02ect how well distributed the inliers and outliers are, with larger values of\ncX ,min (respectively, smaller values of \u03b7O) corresponding to a more uniform distributions of inliers\n(cid:1), the DPCP problem (15) satis\ufb01es the (\u03b1, \u0001, S\n(respectively, outliers). One of the key insights in this paper is that the DPCP problem (15) satis\ufb01es\nthe RRC of De\ufb01nition 1 as we now state.\nTheorem 2. For any \u0001 <\nRRC with \u03b1 = ((1 \u2212 \u00012/2)N cX ,min \u2212 M \u03b7O)/\u221a2c and any orthonormal basis S\n\u22a5 for S\n(cid:107)G(B)(cid:107)F \u2264 \u221aN (cid:107)X(cid:107)2 + M \u03b7O for all B \u2208 O(c, D), where (cid:107) \u00b7 (cid:107)2 denotes the spectral norm.\n\n2(cid:0)1 \u2212 M \u03b7O/N cX ,min\n\n\u22a5\n)-\n\u22a5. Also,\n\n(cid:113)\n\ni B)\n\nCombining this with Theorem 1 allows us to conclude the linear convergence of Algorithm 1 to S\n\n6\n\n\u22a5.\n\n\f\u22a5\n\n\u22a5\n\n\u22a5\n\n\u22a5\n\nCorollary 1. Suppose that the initialization B0 satis\ufb01es dist2(B0, S\n) < 2 (1 \u2212 M \u03b7O/N cX ,min) ,\nwhere dist(B0, S\n) is de\ufb01ned in (2). Let {Bk} be the sequence generated by Algorithm 1 for solving\nthe DPCP problem (15) with G(Bk) in (16) and step size \u00b5k = \u00b50\u03b2k, where \u00b50 and \u03b2 satisfy (13)\n), \u03b1 = ((1 \u2212 \u00012/2)N cX ,min \u2212 M \u03b7O)/\u221a2c, and \u03be = \u221aN (cid:107)X(cid:107)2 + M \u03b7O.\n\u22a5\nwith \u0001 = dist(B0, S\n\u22a5 at an R-linear rate, i.e., dist(Bk, S\n) for all k \u2265 0.\nThen, Bk converges to S\nCorollary 1 implies that the Riemannian subgradient method with a good initialization converges to an\n\u22a5 at an R-linear rate. When c = 1, a projected subgradient method was proved\northonormal basis of S\nto have a piecewise linear convergence rate in [51]. In this case, Corollary 1 improves upon [51]\nin three ways: (i) it allows for a simpler strategy for selecting the step size than does the piecewise\ngeometrically diminishing step size, which has two more parameters controlling when and how\noften to decay the step size; (ii) it provides a more transparent convergence analysis since its proof\nfollows directly from the RRC and Theorem 1; and (iii) it places a slightly weaker requirement on the\n\ninitialization, which in practice we compute as the bottom eigenvectors of (cid:101)X (cid:101)X (cid:62) as in [42, 51]. In the\n\n) \u2264 \u03b2k dist(B0, S\n\nsupplementary material, we show this spectral initialization satis\ufb01es the requirement in Corollary 1.\n\n(a) Performance on problem (15)\nwith different step size choices \u00b5k.\n\n(b) Performance on problem (15) for\nthe step size \u00b5k = 0.01\u03b2k.\n\n(c) Performance on problem (17)\nwith different step sizes choices \u00b5k.\n\nFigure 2: Performance of Algorithm 1 on the DPCP problem (15) and the ODL problem (17).\n\nExperiments. Synthetic data for the DPCP problem is generated as follows: randomly sample\na subspace S of co-dimension c = 10 in ambient dimension D = 100, and uniformly at random\nratio is M/(M + N ) = 0.7. As an initialization B0, we use the bottom c eigenvectors of (cid:101)X (cid:101)X (cid:62)\nsample N = 1500 inliers from S \u2229 O(1, D) and M = 3500 outliers from O(1, D) so that the outlier\n[42, 51] as we described before.\nFigure 2a displays the convergence of the projected Riemannian subgradient method with different\nchoices of step size. We observe linear convergence for the geometrically diminishing step size,\nwhich converges much faster than when a constant step size or classical diminishing step size (O(1/k)\nand O(1/\u221ak)) is used. In Figure 2b, we illustrate the effect of the decay factor \u03b2 for Algorithm 1\nwith geometrically diminishing step size \u00b5k = 0.01\u03b2k. First observe that, as expected, \u03b2 controls the\nconvergence speed. When \u03b2 is too small (e.g., \u03b2 \u2208 {0.5, 0.6}) convergence may not occur, which\nagrees with (13) and (14). However, when \u03b2 \u2265 0.7 the algorithm converges at an R-linear rate, with\nlarger values of \u03b2 resulting in slower convergence speeds.\n\n4.2 Orthogonal Dictionary Learning\n\nGiven a dataset (cid:101)X \u2208 RD\u00d7N , DL [32] aims to learn a sparse representation \u0398 \u2208 RM\u00d7N for (cid:101)X by\n\ufb01nding a dictionary A \u2208 RD\u00d7M such that (cid:101)X \u2248 A\u0398 with \u0398 sparse. Several DL methods have been\n\nproposed in the literature, including the method of optimal directions (MOD) [16], K-SVD [15],\nand alternating minimization [1], as well as the Riemannian trust region method [39] and projected\nRiemannian subgradient method [2] for ODL. Here, we consider the ODL problem [39, 2] in which\nthe dictionary is square and orthogonal and the data is generated by the following random model.3\nmatrix. The data is generated as (cid:101)X = A\u0398, where each column of \u0398 \u2208 RD\u00d7D is an i.i.d. Bernoulli-\nDe\ufb01nition 2 (Random model for ODL [2]). Assume A \u2208 RD\u00d7D is a \ufb01xed but unknown orthonormal\nGaussian random vector with parameter \u03c1 \u2208 (0, 1) that controls the sparsity.\n\n3Extensions to other models including deterministic models are the subject of future work.\n\n7\n\n0100200300iteration10-1010-505010015020010-1010-5100 = 0.5 = 0.6 = 0.7 = 0.8 = 0.902004006008001000iteration10-5100\fIf b is a column of A, then A\n\nb is a standard basis vector and (cid:101)X (cid:62)b = \u0398(cid:62)\n\nminimizing (cid:107)(cid:101)X (cid:62)b(cid:107)0 over the sphere is expected to yield the column of A that is least used in the\n\nrepresentation \u0398. For computational reasons, the (cid:96)0 semi-norm is replaced by the (cid:96)1 norm [2] 4 thus\nleading to the problem5\n\nb is sparse. Thus,\n\nA\n\n(cid:62)\n\n(cid:62)\n\nminimize\nb\u2208O(1,D)\n\nf (b) = 1\n\nb(cid:107)1.\n\n(17)\n\nN (cid:107)(cid:101)X (cid:62)\n\nVeri\ufb01cation of the regularity condition. We show that the regularity condition in (5) is satis\ufb01ed\nfor problem (17). The primary difference with our previous analysis is that the data is now random.\nThus, (5) will only be proved to hold with high probability. Towards that end, similar to (16), a\nRiemannian subgradient for (17) is\n\n(cid:0)I \u2212 bb\n\n(cid:62)(cid:1)(cid:16) N(cid:88)\n\ni=1\n\nG(b) = 1\n\nN\n\n(cid:17)\n\nsign((cid:101)x\n\ni b)(cid:101)xi\n\n(cid:62)\n\n.\n\n(18)\n\nb2\ni\n\nmaxj(cid:54)=i b2\n\nj \u2265 1 + \u03b6\n\n(cid:110)\nb \u2208 O(1, D) :\n\nThe projected Riemannian subgradient method has been utilized in [2] for solving (17), but only\nwith a sublinear rate of convergence guarantee, even though the function has been proved to satisfy\n(5) with high probability. Based on this condition, we will show that Algorithm 1 can solve (17)\nmore ef\ufb01ciently, indeed with a linear convergence rate. To describe the RRC for (17), suppose\nwithout loss of generality that the orthonormal dictionary A is the identity matrix. Then, the\n(cid:111)\ngoal is to \ufb01nd the standard basis vectors {\u00b1e1, . . . ,\u00b1eD}, where the sign is irrelevant because\nf (b) = f (\u2212b). We now de\ufb01ne a region of interest that is near each basis vector ei and \u2212ei as\n, where \u03b6 > 0 and bj is the j-th entry of b. Each region I i\nI i\n\u03b6 =\ncontains all unit vectors whose i-th entry is at least \u221a1 + \u03b6 larger (in absolute value) than the other\nentries. The RRC for (17) is then captured by the following result.\nTheorem 3. [2, Theorem 3.6] Assume \u03c1 \u2208 [1/D, 1/2] in the random model of De\ufb01nition 2. There\nexist universal constants C, c > 0 such that if N \u2265 CD4\u03b6\u22122\u03c1\u22122 log(D/\u03b6) for all \u03b6 \u2208 (0, 1), then\nwith probability at least 1 \u2212 exp(\u2212cN \u03c13\u03b6 2D\u22123/ log N ) the ODL problem (17) satis\ufb01es (5) for any\nb \u2208 I i\n\u03b6, but not all b that is \u0001-close to ei.\nNote that Theorem 3 ensures that (5) holds only for all b \u2208 I i\nFortunately, [2, Proposition D.2] ensures the iterates generated by Algorithm 1 do stay within I i\n\u03b6,\nwhich together with Theorem 1 guarantees the convergence of Algorithm 1.\nCorollary 2. Let {bk} be the sequence generated by Algorithm 1 for the ODL problem (17) with\n64 ) and step size \u00b5k = \u00b50\u03b2k, where \u00b50 and \u03b2 satisfy the conditions in Theorem 1\nb0 \u2208 I i\nwith \u03be = 2 and \u0001 = \u221a2, and \u03b1 = 1\n2 . Under the same setup as in Theorem 3, with\nprobability at least 1 \u2212 exp(\u2212cN \u03c13\u03b6 2D\u22123/ log N ), {bk} converges to ei at an R-linear rate, i.e.,\n(19)\n\n\u03b6 with G(b) in (18) and B(cid:63) = ei for any i, and \u03b1 = 1\n\n16 \u03c1(1 \u2212 \u03c1)\u03b6D\u2212 3\n\n16 \u03c1(1 \u2212 \u03c1)\u03b6D\u2212 3\n\n\u03b6 (\u03b6 \u2264 55\n\n2 .\n\n\u03b6\n\ndist(bk, ei) \u2264 \u03b2k dist(b0, ei).\n\nCorollary 2 improves upon [2, Theorem 3.8], according to which under the setup in Corollary 2\nand with step size \u00b5k = O(1/k3/8), it follows that mink(cid:48)\u2264k dist(bk(cid:48), ei) = O(1/k3/8). Indeed, our\nresult (19) gives a direct bound on the k-th iteration and not on the best iteration obtained so far.\n\nExperiments We use the same setup in [2] by \ufb01rst generating a random orthogonal dictionary\nA \u2208 RD\u00d7D with D = 70, sparsity level \u03c1 = 0.3, and number of data points N = 5857 \u2248 10D1.5.\nAs in [2], the initialization b0 is randomly generated from the unit sphere O(1, D) and belongs to one\nof the D sets {I i\n1/5 log D : i = 1, . . . , D} with probability at least one-half [2, Lemma 3.9]. Figure 2c\n4[39] considered a smoothed version of (17), allowing one to use gradient-based algorithms. However, the\n\nobtained solution is perturbed from the targeted one and thus a rounding step is needed.\n\n5All the columns of A can be obtained by repeating this process with the removal of the contribution from\nthe previously learned columns. Alternatively, as will been seen in Corollary 2, the column that Algorithm 1\nconverges to depends on the initialization. Thus, one may simply repeat Corollary 2 with different initializations\n(e.g., random initializations) each time [2]. It is of interest to extend (17) in order to estimate the whole dictionary,\n\ne.g., minimizing (cid:107)(cid:101)X (cid:62)B(cid:107)1, s. t. B \u2208 O(D), where (cid:96)1 counts the sum of the elements of the matrix. Note that\n\nthis is not an optimization on the Grassmannian since the objective is not rotation invariant.\n\n8\n\n\fdisplays the convergence of the Riemannian subgradient method for different choices of the step\nsize for solving (17). We observe linear convergence for geometrically diminishing step size, which\nconverges much faster than the others, in particular when \u00b5k = O( 1\n\nk3/8 ) as is used in [2].\n\n5 Conclusion and Discussion\n\nWe proved that a projected Riemannian subgradient method with geometrically diminishing step\nsizes converges linearly for non-convex and non-smooth problems on the Grassmannian that satisfy\na certain regularity condition on the Riemannian subgradient. We also showed that our regularity\ncondition is satis\ufb01ed by (cid:96)1 co-sparse formulations for orthogonal dictionary learning and robust\nsubspace learning, which led to improved convergence rates when compared to existing results.\nWe conclude this paper by pointing out several interesting directions for future work.\nExtension to intrinsic methods. In this paper we take an extrinsic approach because extrinsic\nmethods are typically easier to implement, e.g. when the projection map is easier to compute than the\ngeodesic distance. Extending the current analysis to an intrinsic optimization method\u2014where the\niterates are taken along a geodesic direction\u2014is worth exploring. For example, the geodesic gradient\ndescent (GGD) [31] has been proved to converge at a piecewise linear rate for the robust subspace\nlearning problem. And extension of the current analysis may allow the GGD to use a simpler step\nsize selection strategy (i.e., a geometrically diminishing step size) to obtain a linear convergence rate.\nExtension to other submanifolds of Euclidean space. Although we focus on optimization problems\nover the Grassmannian, the Riemannian regularity condition can be extended to other submanifolds\nof Euclidean space with an appropriate de\ufb01nition of the Riemanniann metric and distance. Using\nthis condition to analyze the convergence of the projected Riemannian subgradient method for other\nmanifolds (such as the Stiefel manifold) is the subject of ongoing work.\nApplication to other problems. Aside from the robust subspace and dictionary learning problems\nconsidered here, other problems in machine learning and signal processing can be formulated as\nminimizing a non-smooth function over the sphere or Stiefel manifold and thus (potentially) can\nbe ef\ufb01ciently solved by the Riemannian subgradient method. These problems include the (cid:96)1-norm\n(kernel) PCA [22, 30, 45], multi-channel sparse blind deconvolution [28, 33], etc.\n\nAcknowledgment\n\nThis research is supported in part by NSF grant 1704458, ShanghaiTech grant 2017F0203-000-16,\nand the Northrop Grumman Mission Systems Research in Applications for Learning Machines\n(REALM) initiative. Zhihui Zhu would like to thank Xiao Li and Dr. Anthony Man-Cho So (CUHK)\nfor fruitful discussions about regularity conditions for non-smooth optimization problems.\n\nReferences\n[1] S. Arora, R. Ge, T. Ma, and A. Moitra. Simple, ef\ufb01cient, and neural algorithms for sparse coding. Journal\n\nof Machine Learning Research, 2015.\n\n[2] Y. Bai, Q. Jiang, and J. Sun. Subgradient Descent Learns Orthogonal Dictionaries. In International\n\nConference on Learning Representations, 2019.\n\n[3] L. Balzano, R. Nowak, and B. Recht. Online identi\ufb01cation and tracking of subspaces from highly\nIn Communication, Control, and Computing (Allerton), 2010 48th Annual\n\nincomplete information.\nAllerton Conference on, pages 704\u2013711. IEEE, 2010.\n\n[4] N. Boumal, P.-A. Absil, and C. Cartis. Global rates of convergence for nonconvex optimization on\n\nmanifolds. IMA Journal of Numerical Analysis, 2016.\n\n[5] J. V. Burke and M. C. Ferris. Weak sharp minima in mathematical programming. SIAM Journal on Control\n\nand Optimization, 31(5):1340\u20131359, 1993.\n\n[6] E. J. Cand\u00e8s, X. Li, and M. Soltanolkotabi. Phase retrieval via wirtinger \ufb02ow: Theory and algorithms.\n\nIEEE Transactions on Information Theory, 61(4):1985\u20132007, 2015.\n\n9\n\n\f[7] S. Chen, S. Ma, A. M.-C. So, and T. Zhang. Proximal gradient method for manifold optimization. arXiv\n\npreprint arXiv:1811.00980, 2018.\n\n[8] F. E. Curtis, T. Mitchell, and M. L. Overton. A bfgs-sqp method for nonsmooth, nonconvex, constrained\noptimization and its evaluation using relative minimization pro\ufb01les. Optimization Methods and Software,\n32(1):148\u2013181, 2017.\n\n[9] F. E. Curtis and X. Que. A quasi-newton algorithm for nonconvex, nonsmooth optimization with global\n\nconvergence guarantees. Mathematical Programming Computation, 7(4):399\u2013428, 2015.\n\n[10] D. Davis, D. Drusvyatskiy, K. J. MacPhee, and C. Paquette. Subgradient methods for sharp weakly convex\n\nfunctions. arXiv preprint arXiv:1803.02461, 2018.\n\n[11] D. Davis, D. Drusvyatskiy, and C. Paquette. The nonsmooth landscape of phase retrieval. arXiv preprint\n\narXiv:1711.03247, 2017.\n\n[12] T. Ding, Z. Zhu, T. Ding, Y. Yang, D. Robinson, R. Vidal, and M. Tsakiris. Noisy dual principal component\n\npursuit. In Proceedings of the International Conference on Machine learning, 2019.\n\n[13] J. C. Duchi and F. Ruan. Solving (most) of a set of quadratic equalities: Composite optimization for robust\n\nphase retrieval. arXiv preprint arXiv:1705.02356, 2017.\n\n[14] A. Edelman, T. Arias, and S. T. Smith. The geometry of algorithms with orthogonality constraints. SIAM\n\nJournal of Matrix Analysis Applications, 20(2):303\u2013353, 1998.\n\n[15] M. Elad and M. Aharon. Image denoising via sparse and redundant representations over learned dictionaries.\n\nIEEE Transactions on Image Processing, 15(12):3736\u20133745, 2006.\n\n[16] K. Engan, S. O. Aase, and J. H. Husoy. Method of optimal directions for frame design. IEEE International\n\nConference on Acoustics, Speech, and Signal Processing, 1999.\n\n[17] M. A. Fischler and R. C. Bolles. RANSAC random sample consensus: A paradigm for model \ufb01tting with\napplications to image analysis and automated cartography. Communications of the ACM, 26:381\u2013395,\n1981.\n\n[18] J.-L. Gof\ufb01n. Subgradient optimization in nonsmooth optimization (including the Soviet revolution). Groupe\n\nd\u2019\u00e9tudes et de recherche en analyse des d\u00e9cisions, 2012.\n\n[19] P. Grohs and S. Hosseini. Nonsmooth trust region algorithms for locally Lipschitz functions on Riemannian\n\nmanifolds. IMA Journal of Numerical Analysis, 36(3):1167\u20131192, 2015.\n\n[20] J. Hamm and D. D. Lee. Grassmann discriminant analysis: a unifying view on subspace-based learning. In\n\nProceedings of the 25th international conference on Machine learning, pages 376\u2013383. ACM, 2008.\n\n[21] N. Higham and P. Papadimitriou. Matrix procrustes problems. Rapport technique, University of Manchester,\n\n1995.\n\n[22] C. Kim and D. Klabjan. A simple and fast algorithm for l1-norm kernel pca. IEEE transactions on pattern\n\nanalysis and machine intelligence, 2019.\n\n[23] J. D. Lee, I. Panageas, G. Piliouras, M. Simchowitz, M. I. Jordan, and B. Recht. First-order methods almost\n\nalways avoid saddle points. arXiv preprint arXiv:1710.07406, 2017.\n\n[24] G. Lerman and T. Maunu. Fast, robust and non-convex subspace recovery. Information and Inference: A\n\nJournal of the IMA, 7(2):277\u2013336, 2017.\n\n[25] G. Lerman and T. Maunu. An overview of robust subspace recovery. Proceedings of the IEEE, 106(8):1380\u2013\n\n1410, 2018.\n\n[26] G. Lerman, M. B. McCoy, J. A. Tropp, and T. Zhang. Robust computation of linear models by convex\n\nrelaxation. Foundations of Computational Mathematics, 15(2):363\u2013410, 2015.\n\n[27] X. Li, Z. Zhu, A. M.-C. So, and R. Vidal. Nonconvex robust low-rank matrix recovery. arXiv preprint\n\narXiv:1809.09237, 2018.\n\n[28] Y. Li and Y. Bresler. Global geometry of multichannel sparse blind deconvolution on the sphere. In\n\nAdvances in Neural Information Processing Systems, pages 1132\u20131143, 2018.\n\n[29] Z.-Q. Luo and P. Tseng. Error bounds and convergence analysis of feasible descent methods: a general\n\napproach. Annals of Operations Research, 46(1):157\u2013178, 1993.\n\n10\n\n\f[30] P. P. Markopoulos, G. N. Karystinos, and D. A. Pados. Optimal algorithms for l_{1}-subspace signal\n\nprocessing. IEEE Transactions on Signal Processing, 62(19):5046\u20135058, 2014.\n\n[31] T. Maunu, T. Zhang, and G. Lerman. A well-tempered landscape for non-convex robust subspace recovery.\n\nJournal of Machine Learning Research, 20(37):1\u201359, 2019.\n\n[32] B. Olshausen and D. Field. Emergence of simple-cell receptive \ufb01eld properties by learning a sparse code\n\nfor natural images. Nature, 381(6583):607\u2013609, 1996.\n\n[33] Q. Qu, X. Li, and Z. Zhu. A nonconvex approach for exact and ef\ufb01cient multichannel sparse blind\n\ndeconvolution. In Advances in Neural Information Processing Systems, 2019.\n\n[34] Q. Qu, J. Sun, and J. Wright. Finding a sparse vector in a subspace: Linear sparsity using alternating\n\ndirections. In Advances in Neural Information Processing Systems, pages 3401\u20133409, 2014.\n\n[35] M. Rahmani and G. Atia. Coherence pursuit: Fast, simple, and robust principal component analysis. arXiv\n\npreprint arXiv:1609.04789, 2016.\n\n[36] R. Slama, H. Wannous, M. Daoudi, and A. Srivastava. Accurate 3d action recognition using learning on\n\nthe grassmann manifold. Pattern Recognition, 48(2):556\u2013567, 2015.\n\n[37] M. Soltanolkotabi, E. Elhamifar, and E. Cand\u00e8s. Robust subspace clustering. http://arxiv.org/abs/1301.2603,\n\n2013.\n\n[38] G. W. Stewart and J. Sun. Matrix perturbation theory. Academic press, 1990.\n\n[39] J. Sun, Q. Qu, and J. Wright. Complete dictionary recovery over the sphere i: Overview and the geometric\n\npicture. IEEE Transactions on Information Theory, 63(2):853\u2013884, 2017.\n\n[40] M. Tsakiris and R. Vidal. Dual principal component pursuit. In ICCV Workshop on Robust Subspace\n\nLearning and Computer Vision, pages 10\u201318, 2015.\n\n[41] M. C. Tsakiris and R. Vidal. Hyperplane clustering via dual principal component pursuit. In International\n\nConference on Machine Learning, 2017.\n\n[42] M. C. Tsakiris and R. Vidal. Dual principal component pursuit. Journal of Machine Learning Research,\n\n18(19):1\u201350, 2018.\n\n[43] K. Usevich and I. Markovsky. Optimization on a Grassmann manifold with application to system identi\ufb01-\n\ncation. Automatica, 50(6):1656\u20131662, 2014.\n\n[44] J.-P. Vial. Strong and weak convexity of sets and functions. Mathematics of Operations Research,\n\n8(2):231\u2013259, 1983.\n\n[45] P. Wang, H. Liu, and A. M.-C. So. Globally convergent accelerated proximal alternating maximization\nmethod for l1-principal component analysis. In IEEE International Conference on Acoustics, Speech and\nSignal Processing (ICASSP), pages 8147\u20138151. IEEE, 2019.\n\n[46] H. Xu, C. Caramanis, and S. Sanghavi. Robust pca via outlier pursuit. IEEE Transactions on Information\n\nTheory, 5(58):3047\u20133064, 2012.\n\n[47] W. H. Yang, L.-H. Zhang, and R. Song. Optimality conditions for the nonlinear programming problems on\n\nriemannian manifolds. Paci\ufb01c Journal of Optimization, 10(2):415\u2013434, 2014.\n\n[48] C. You, D. P. Robinson, and R. Vidal. Provable self-representation based outlier detection in a union of\n\nsubspaces. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4323\u20134332, 2017.\n\n[49] T. Zhang and G. Lerman. A novel m-estimator for robust pca. The Journal of Machine Learning Research,\n\n15(1):749\u2013808, 2014.\n\n[50] Y. Zhang, H.-W. Kuo, and J. Wright. Structured local minima in sparse blind deconvolution. In Advances\n\nin Neural Information Processing Systems, pages 2322\u20132331, 2018.\n\n[51] Z. Zhu, Y. Wang, D. P. Robinson, D. Naiman, R. Vidal, and M. C. Tsakiris. Dual principal component\n\npursuit: Improved analysis and ef\ufb01cient algorithms. In Neural Information Processing Systems, 2018.\n\n11\n\n\f", "award": [], "sourceid": 5035, "authors": [{"given_name": "Zhihui", "family_name": "Zhu", "institution": "Johns Hopkins University"}, {"given_name": "Tianyu", "family_name": "Ding", "institution": "Johns Hopkins University"}, {"given_name": "Daniel", "family_name": "Robinson", "institution": "Johns Hopkins University"}, {"given_name": "Manolis", "family_name": "Tsakiris", "institution": "ShanghaiTech University"}, {"given_name": "Ren\u00e9", "family_name": "Vidal", "institution": "Mathematical Institute for Data Science Johns Hopkins University"}]}