{"title": "Unsupervised Learning of Human Motion Models", "book": "Advances in Neural Information Processing Systems", "page_first": 1287, "page_last": 1294, "abstract": "", "full_text": "Unsupervised Learning of Human Motion\n\nModels\n\nYang Song, Luis Goncalves, and Pietro Perona\n\nCalifornia Institute of Technology, 136-93, Pasadena, CA 9112 5, USA\n\n yangs,luis,perona\n\n@vision.caltech.edu\n\nAbstract\n\nThis paper presents an unsupervised learning algorithm that can derive\nthe probabilistic dependence structure of parts of an object (a moving hu-\nman body in our examples) automatically from unlabeled data. The dis-\ntinguished part of this work is that it is based on unlabeled data, i.e., the\ntraining features include both useful foreground parts and background\nclutter and the correspondence between the parts and detected features\nare unknown. We use decomposable triangulated graphs to depict the\nprobabilistic independence of parts, but the unsupervised technique is\nnot limited to this type of graph. In the new approach, labeling of the\ndata (part assignments) is taken as hidden variables and the EM algo-\nrithm is applied. A greedy algorithm is developed to select parts and to\nsearch for the optimal structure based on the differential entropy of these\nvariables. The success of our algorithm is demonstrated by applying it\nto generate models of human motion automatically from unlabeled real\nimage sequences.\n\n1 Introduction\n\nHuman motion detection and labeling is a very important but dif\ufb01cult problem in computer\nvision. Given a video sequence, we need to assign appropriate labels to the different regions\nof the image (labeling) and decide whether a person is in the image (detection). In [8, 7],\na probabilistic approach was proposed by us to solve this problem. To detect and label a\nmoving human body, a feature detector/tracker (such as corner detector) is \ufb01rst run to obtain\nthe candidate features from a pair of frames. The combination of features is then selected\nbased on maximum likelihood by using the joint probability density function formed by the\nposition and motion of the body. Detection is performed by thresholding the likelihood.\nThe lower part of Figure 1 depicts the procedure.\n\nOne key factor in the method is the probabilistic model of human motion. In order to avoid\nexponential combinatorial search, we use conditional independence property of body parts.\nIn the previous work[8, 7], the independence structures were hand-crafted. In this paper,\nwe focus on the the previously unresolved problem (upper part of Figure 1): how to learn\nthe probabilistic independence structure of human motion automatically from unlabeled\ntraining data, meaning that the correspondence between the candidate features and the parts\nof the object is unknown. For example when we run a feature detector (such as Lucas-\nTomasi-Kanade detector [10]) on real image sequences, the detected features can be from\n\n\u0001\n\f Feature\n detector/\n tracker \n\nUnsupervised\n Learning \n algorithm\n\n Probabilistic\n Model of\n\nHuman Motion\n\n Unlabeled \nTraining Data\n\n Feature\n detector/\n tracker\n\nDetection \n and \nLabeling\n\nTesting: two frames\n\nFigure 1: Diagram of the system.\n\nPresence of Human? \nLocalization of parts?\n\ntarget objects and background clutter with no identity attached to each feature. This case\nis interesting because the candidate features can be acquired automatically. Our algorithm\nleads to systems able to learn models of human motion completely automatically from real\nimage sequences - unlabeled training features with clutter and occlusion.\n\nWe restrict our attention to triangulated models, since they both account for much corre-\nlation between the random variables that represent the position and motion of each body\npart, and they yield ef\ufb01cient algorithms. Our goal is to learn the best triangulated model,\ni.e., the one that reaches maximum likelihood with respect to the training data. Struc-\nture learning has been studied under the graphical model (Bayesian network) framework\n([2, 4, 5, 6]). The distinguished part of this paper is that it is an unsupervised learning\nmethod based on unlabeled data, i.e., the training features include both useful foreground\nparts and background clutter and the correspondence between the parts and detected fea-\ntures are unknown. Although we work on triangulated models here, the unsupervised tech-\nnique is not limited to this type of graph.\n\nThis paper is organized as follows. In section 2 we summarize the main facts about the\ntriangulated probability model. In section 3 we address the learning problem when the\ntraining features are labeled, i.e., the parts of the model and the correspondence between\nthe parts and observed features are known. In section 4 we address the learning problem\nwhen the training features are unlabeled. In section 5 we present some experimental results.\n\n2 Decomposable triangulated graphs\n\nDiscovering the probability structure (conditional independence) among variables is im-\nportant since it makes ef\ufb01cient learning and testing possible, hence some computationally\nintractable problems become tractable. Trees are good examples of modeling conditional\n(in)dependence [2, 6]. The decomposable triangulated graph is another type of graph which\nhas been demonstrated to be useful for biological motion detection and labeling [8, 1].\n\nA decomposable triangulated graph [1] is a collection of cliques of size three, where there\nis an elimination order of vertices such that when a vertex is deleted, it is only contained\nin one triangle and the remaining subgraph is again a collection of triangles until only one\ntriangle left. Decomposable triangulated graphs are more powerful than trees since each\nnode can be thought of as having two parents. Similarly to trees, ef\ufb01cient algorithms allow\nfast calculation of the maximum likelihood interpretation of a given set of data.\n\nConditional independence among random variables (parts) can be described by a decom-\nposable triangulated graph. Let \n\f\u000b\u000e\r\u0010\u000f\u0012\u0011\u0014\u0013\u0015\u000f\u0017\u0016\u0018\u0013\u001a\u0019\u001b\u0019\u001a\u0019\u001c\u0013\u0015\u000f\u0017\u001d\u001f\u001e be the set of \u001d\nparts, and\n, is the measurement for \u000f+' . If the joint probability density function\n \"!$# , \u0011&%(')%*\u001d\n\u0013\u001b\u0019\u001a\u0019\u001b\u0019\u001b\u0013\n,.-/ \"!\u00180\n \"!$3)4 can be decomposed as a decomposable triangulated graph, it can\n\n \"!21\n\n\n\u0001\n\n\u0002\n\u0003\n\u0002\n\u0003\n\u0001\n\u0003\n\u0004\n\u0005\n\n\u0006\n\u0003\n\u0004\n\u0003\n\u0005\n\n\u0007\n\n\u0004\n\n\u0005\n\u0003\n\u0007\n\u0003\n\b\n\n\b\n\t\n\u0004\n\u0013\n\fbe written as,\n\n\u0013\u001b9\n\n\u0013\u001b:\n\nwhere 8\n\u0013\u001b9\n\n\u0013\u001b:\n\n\u0002\u0001\u0004\u0003\u0006\u0005\b\u0007\n\t\f\u000b\u000e\r\u0010\u000f\u0012\u0011\u0014\u0013\u0015\r\u0016\u000f\u0018\u0017\u0019\u0013\u0014\u001a\u001b\u001a\u0014\u001a\u0015\r\u0016\u000f\u0018\u001c\u001e\u001d\n\u00132\r\n \"!$#\n%\u000e&\n\u001d@?\n#<;\n\u0013\u001b:\n\u0013C9\n\n+,(.-/(0\u000b\u000e\r\n\u0002')(\b*\n%>=\n, \u0011&%\n\u0013\u001a\u0019\u001b\u0019\u001b\u0019\u001a\u0013\n\n(01\n\u0013C9\n\n\u0013\u001b:\n\n')5\u0004+)5\u0004-65\n\n')5\n\u000b\u000e\r\n\u001d43\u0014\n\u0016 , \r\n\u0013\u001a\u0019\u001b\u0019\u001b\u0019\u001b\u0013\n4 are the cliques.\n\n+\u00045\n\u0013C9\n\n\u00132\r\n8BA\n\n-65\n\u00137\r\n\u0013\u0014:\n\u0013\u001b\u0019\u001b\u0019\u001a\u0019\u001b\u0013\n\n(1)\n\n\u000b*\n\n, and\n4 gives\n\nthe elimination order for the decomposable graph.\n\n-/ \n\n\u0013\u001b\u0019\u001a\u0019\u001b\u0019\n\n%LK\n\n\u001e are i.i.d samples from a probability density function,\n%JI\n, are labeled data. We want to \ufb01nd the\n, such that ,.-\nis the\n\n\u0013\u001a\u0019\u001b\u0019\u001b\u0019\u001a\u0013\nbeing the \u2019correct\u2019 one given the observed data D\n\n3 Optimization of the decomposable triangulated graph\n FE\nSuppose D\nwhere HG\n4 ,\n!\u00180\n,.-\ndecomposable triangulated graph M\nprobability of graph M\n. Here we use M\nto denote both the decomposable graph and the conditional (in)dependence depicted by the\ngraph. By Bayes\u2019 rule, ,.-\n4 , therefore if we can assume the\nMON\n4 are equal for different decompositions, then our goal is to \ufb01nd the structure\npriors ,.-\nM which can maximize ,.-\n4 . From the previous section, a decomposable triangulated\nDHN\n\u0013C9\nis represented by -\ngraph M\nDHN\ncan be computed as follows,\n\u001dYX\nZ\\[\n\nMON\n4\u001bP\u0010,.-\n\u0013\u001b:\n+,(\n\n\u0013\u001b\u0019\u001a\u0019\u001b\u0019\u001b\u0013\n-\u0012(\n\n4 , then ,.-\n\nis maximized.\n\n\u0013\u001b9\n'\u0004(\u00191\n\n\u0010\u000bVU\n\n\u0013C9\n\u000b\u000e\n\nDHN\n\nZ_[\n\nMON\n\nQSR\fT\n\n,.-\n\n,.-\n\n\u0013\u0014:\n\n\u0013\u0014:\n\n\u00137\n\n\u0013\u0015\n\n\u000b\u000e\n\n(2)\n\nis differential entropy or conditional differential entropy [3] (we consider\ncontinuous random variables here). Equation (2) is an approximation which converges\ndue to the weak Law of Large numbers and de\ufb01nitions and\nproperties of differential entropy [3, 2, 4, 5, 6]. We want to \ufb01nd the decomposition\n\u0013\u001a\u0019\u001b\u0019\u001b\u0019\u001a\u0013\n4 such that the above equation can be maxi-\n\nwhere `\nto equality for K\n\u0013\u001b9\n\n\u0013\u001b:\n\n\u0013\u001b:\n\n\u0013\u001b:\n\n\u0013C9\n\n\u0013C9\n\n\u0011,^\n\n%\u000e&\n\n-7a\n\nmized.\n\n8BA\n\n3.1 Greedy search\n\n ef\n\n ed\n\nThough for tree cases, the optimal structure can be obtained ef\ufb01ciently by the maximum\nspanning tree algorithm [2, 6], for decomposable triangulated graphs, there is no existing\nalgorithm which runs in polynomial time and guarantees to the optimal solution [9]. We\ndevelop a greedy algorithm to grow the graph by the property of decomposable graphs. For\n\neach possible choice of :\n(the last vertex of the last triangle), \ufb01nd the best 9\nA which\ncan maximize ?\n4 , then get the best child of edge -\n4 as 8\n, i.e., the\nvertex (part) that can maximize ?\n4 . The next vertex is added one by\none to the existing graph by choosing the best child of all the edges (legal parents) of the\nexisting graph until all the vertices are added to the graph. For each choice of :\n, one such\ngraph can be grown, so there are \u001d\ncandidate graphs. The \ufb01nal result is the graph with the\n4 among the \u001d\nhighest hji/k\ngraphs.\nThe above algorithm is ef\ufb01cient. The total search cost is \u001dml\n\n,.-\nDFN\n=u?\n4 , which is on the order of \u001drw . The algorithm is a greedy algorithm, with no\n\nguarantee that the global optimal solution could be found. Its effectiveness will be explored\nthrough experiments.\n\n\u0011porqts\n\n\u001dn?\n\n-/ <g\n\n\u0016\u0016l\n\n\u0013\u0014:\n\n3.2 Computation of differential entropy - translation invariance\n\nIn the greedy search algorithm, we need to compute `\n\u0011.%\nbe computed by 0\n\n4 ,\n. If we assume that they are jointly Gaussian, then the differential entropy can\nis the covariance matrix.\n\n4 and `\n\n%r=\n\n-/ \n\n <g\n\n\u0016\u0018z|{\n\nhxiyk\n\n}\u0010N , where I\n\nis the dimension and }\n\n\u001f\n\u0011\n\u0011\n'\n\n+\n(\n-\n(\n\u001d\n#\n#\n\n'\n\u000b\n8\n0\n\u0013\n8\n1\nA\nA\n\u001e\n-\n8\n0\n0\n0\n4\n\u0013\n-\n8\n1\n1\n1\n4\n-\n8\nA\nA\nA\n-\n8\n0\n\u0013\n8\n1\n8\nA\n\u000b\n\n \n0\n\u0013\n \n1\n\u0013\n\u000b\nG\n \nG\n!\n3\n\u0011\nD\n4\nD\n4\nD\n4\n\u000b\nM\n4\nM\nD\nM\nM\n8\n0\n0\n0\n4\n\u0013\n-\n8\n1\n1\n1\n4\n-\n8\nA\nA\nA\nM\n4\n1\nW\n\u001f\n3\n!\n]\n\n\u001d\n3\n^\n+\n5\n-\n5\n\u001d\n4\nb\nc\n-\n8\n0\n0\n0\n4\n\u0013\n-\n8\n1\n1\n1\n4\n-\nA\nA\nA\n`\n-\n5\n\u0013\n5\n9\nA\nA\nA\n`\n5\nN\n \nd\n5\n\u0013\n \nf\n5\nA\nM\n-\n-\n-\n-\nv\n4\no\n\u0011\n4\nl\nv\n4\n-\n(\n\u0013\n \nd\n(\n\u0013\n \nf\n(\nd\n(\n\u0013\n \nf\n(\nv\n1\n-\n4\nG\nN\n\f\u0013\u001b9\n\n\u0013\u0014:\n\nIn our applications, position and velocity are used as measurements for each body part, but\nhumans can be present at different locations of the scene. In order to make the Gaussian\nassumption reasonable, translations need to be removed. Therefore, we use local coordinate\nsystem for each triangle -\n\ns ) as\nthe origin, and use relative positions for other body parts. More formally, let  denote a\nvector of positions\ns , we obtain,\u0004\u0003\n\u000b\t\b\nrelative to 8\nbe written as\n\n4 , i.e., we can take one body part (for example 8\n\u0013\u0002\u0001\n\u0013\u0002\u0001\n, with\f\n\nIn the greedy search algorithm, the differential entropy of all the possible triplets are\nneeded and different triplets are with different origins. To reduce computational cost, notice\nthat\n\n. Then if we describe positions\n. This can\n\n\u0013\u0005\u0001\nZ\u0012\u0011\u0013\u0011\nZ\u0012\u0011\n\n, where [12]\n\n\u000b\r\f\n\n?\u0007\u0001\n\n?\u0006\u0001\n\n\u0013\u0002\u0001\n\n\u0013\u0002\u0001\n\n.\n\nand\n\n[\u0018\u0017\n\n[\u001b\u0017\nFrom the above equations, we can \ufb01rst estimate the mean \u001d\n(including all the body parts and without removing translation), then take the dimensions\ncorresponding to the triangle and use equations (4) and (5) to get the mean and covariance\n eg\nfor -\ns )) to achieve translation invariant.\ntaken as origin for (9\n4 Unsupervised learning of the decomposable graph\n\nand covariance }\n4 . Similar procedure can be applied to pairs (for example, 9\n\u0013\u0014:\n\ns can be\n\nof \n\n(4)\n\n(5)\n\n\f\u0010\u000f\n[\u001b\u0017\n\n\u0014\u0016\u0015\n\n\u000b\u001f\n\nIn this section, we consider the case when only unlabeled data are available. Assume we\nhave a data set of K\n, is\na group of detected features which contains the target object, but \nis unlabeled, which\nmeans the correspondence between the candidate features and the parts of the object is\nunknown. For example when we run a feature detector (such as Lucas-Tomasi-Kanade\ndetector [10]) on real image sequences, the detected features can be from target objects and\nbackground clutter with no identity attached to each feature. We want to select the useful\n\n\u001e . Each sample \n\nsamples D\n\n%tI&%\n\n\u0013\u001b\u0019\u001a\u0019\u001b\u0019\u001b\u0013\n\n HE\n\n, \u0011\n\ncomposite parts of the object and learn the probability structure from D\n\n.\n\n4.1 All foreground parts observed\n\nG denote the labeling for FG\n\nHere we \ufb01rst assume that all the foreground parts are observed for each sample. If the label-\ning for each \nis taken as a hidden variable, then the EM algorithm can be used to learn\nthe probability structure and parameters. Our method was developed from [11], but here\nwe learn the probabilistic independence structure and all the candidate features are with the\n. If FG contains I\u001f\u001e\n\u001e -dimensional vector with each element taken a value from \n! \nsame type. Let `\nis an\nfeatures, then `\nis the back-\n\u001e ,\n\u0013\u001b\u0019\u001a\u0019\u001b\u0019\u001a\u0013\nground clutter label). The observations for the EM algorithm are D\nthe hidden variables are\"\n0 , and the parameters to optimize are the probability\n(in)dependence structure (i.e. the decomposable triangulated graph) and parameters for its\nto represent both the probability struc-\nassociated probability density function. We use M\nG s are independent from each other and `\nture and the parameters. If we assume that \nonly depends on \n\u001d\u0016&\n\u0010\u000bVU\nQSR\u0019T\n\u0010\u000b\n\u0003('*),+.-\n\n, then the likelihood function to maximize is,\n\n\u0010\u000b\n\u00190/\n\n\u0010\u000bVU\n\n\u001d1&\n\nG$#\n\n\u0010\u000b\n\nQSR\fT\n\nQSR\fT\n\nQSR\u0019T\n\nR\u0019T\n\n(6)\n\n(9\n\n8\ns\ns\ns\n\u000b\n-\n\ng\n(\n\u0013\n\nd\n(\n\u0013\n\nf\n(\ng\n(\nd\n(\nf\n(\n4\nA\n\u000b\n-\n\nd\n(\n?\n\ng\n(\n\u0013\n\nf\n(\n?\n\ng\n(\nd\n(\ng\n(\nf\n(\ng\n(\n4\nA\n\u0003\n\n\n\u001f\n\u000e\n\u000e\n\u001f\n\u000b\n\u000e\n\u000e\n\u0011\n\u000f\n\u001f\n\u0011\n]\n\u0019\n&\n\u0011\n\u001a\n\u0015\n\u0019\n\u001f\n\u0011\n]\n\u0019\n&\n\u0011\n\n\u001a\n\u0019\n\u001f\n\n3\n\u0011\n]\n\u0019\n&\n\u0011\n\u001a\n\u0019\n\u001f\n\n\u0014\n\u001c\n\u0015\n\u001f\n\n\u001c\n\n!\nG\n(\n\u0013\n \nd\n(\n\u0013\n \nf\n(\ns\n \n0\n\u0013\n \n1\nG\nK\nG\nG\nG\nI\n\n9\nM\n\u001e\nM\n\u000b\n\n \n0\n\u0013\n \n1\n \nE\n\u000b\n\n`\nG\n\u001e\nE\nG\nG\n%\n\u001f\n\u0013\nW\n\u001d\n\u001f\n1\nW\nW\n\u001d\n\u001f\n]\n\u0017\n\u0019\n&\n\u0011\n]\n'\n\n\u0019\n\u0013\n^\n\u0019\n\u001f\n^\n1\nW\nQ\nW\n\u001d\n\fis the ' th possible labeling for FG\n\nis the set of all such labelings. Op-\ntimization directly over equation (6) is hard, and the EM algorithm solves the optimization\n\nwhere `\nproblem iteratively. In EM, for each iteration v , we will optimize the function,\n\n, and \n\u0011\u0006\u0007\n\n\u0011\u0006\u0007\n\n# , \n\u000f , then,\n\u0010\u000b\n\n\u00190/\n\n\u0013\u0006\u0005\nQSR\u0019T\n+.-\n+.-\n\n\u0002\u0004\u0003\n\nR\u0019T\n\n\u0010\u000bVU\n\u0002\u0004\u0003\n\n\u0010\u000b\n\u0010\u000b\n\u00190/\n\u0019*/\n\n\u00190/\n\u0010\u000b\n\u0003('\n# is the probability of `\n0 . For each iteration v , \n\n\u00117\u001d\n\u0019*/\n\n'\t\b\n\u0010\u000b\nWe will discuss the computation of ,.-\nesis `\nobject, and background (clutter) \nclutter \n\nwhere \n\ns\f\u000b\nprobability structure M\n# can be computed as,\n\u0010\u000b\n\u0019*/\n\n HG\n\nQSR\u0019T\n# given the observation \n\n(7)\n\n\u0010\u000b\n\nQSR\fT\n\n\u0019*/\n\u0010\u000b\n\nG and the decomposable\n# .\n\n\u00190/\n# is a \ufb01xed number for a hypothesis `\n\u0011\b\u001d\u0006\r\n\u0003('*)\ns\f\u000b\n0\u001b4 below. Under the labeling hypoth-\n\u000e\u0010\u000f , which are parts of the\n\u000e\u0012\u000f are independent of\n\n\u0019*/\n\n(8)\n\n\u0011\b\u001d\n\nis divided into the foreground features \n\n\u000f . If the foreground features \n\n\u0013\u0015\u0014\n\n\u0011C\u001d\n\n\u00190/\n\nQSR\u0019T\n\n%\u000e&\n\n\u001d2\u0010\u000b\n\n\u000e\u0012\u000f\n\n HG\n\n\u001d2\u0010\u000b\n\n\u0010\u000b\n\u0010\u000b\n\n\u00190/\n\u00190/\n\n\u001d2\u0010\u000b\n\u001d2\u0010\u000b\n\n\u0010\u000b\nQSR\fT\n\n\u0013\u001a\u0014\n\u0010\u000b\u000e\n\n\u00190/\n\u0019*/\n\u0016\u0017\u0014\n\u0019*/\n\u00190/\n(9)\n4 are the same for different `\n# , and\nFor simplicity, we will assume the priors ,.-\n4 are the same for different graph structures. If we assume uniform background den-\n,.-\nsities [11, 8], then ,.-\nG\u0019\u0018\n, where 8\nis the volume of the space a\n# . Under probability decomposition\nbackground feature lies in, is the same for different `\n, ,.-\n4 can be computed as in equation (1). Therefore the maximization of\nequation (7) is equivalent to maximizing,\n\u0019*/\n\u00190/\n\n\u001d1&\n\u0019*/\n\u00190/\n-65\nFor most problems, the number of possible labelings is very large (on the order of I\nso it is computationally prohibitive to sum over all the possible `\nHowever, if there is one hypothesis labeling `\u001c\u001b\n# \u2019s, then \n\u001d\u001b\ni.e. \n\u001d\u001b\n# \u2019s as \u001f . Hence equation (10) can be approximated as,\nand other \n\n\u001d1&\n\u00190/! \n\u0010\u000b\u000e\r\n\u0010\u000b\u000e\r\n'\u0004(\nQSR\fT\n\n\u0003('*)\n\u0003('*)\nis much larger than other \n\n\u0019*/\" \n\u0013\u0015\r\n+\u0004(\n\n\u0010\u000b\u000e\r\nQSR\u0019T\n),\n# as in equation (10).\n# that is much better than other hypotheses,,\n# can be taken as \u0011\n\n(11)\nwhere \n# .\nare measurements corresponding to the best labeling `#\u001b\nComparing with equation (2) and also by the weak law of large numbers, we know for\niteration v , if the best hypothesis `\nis used as the \u2019true\u2019 labeling, then the decomposable\ns can be obtained through the algorithm described in section\ntriangulated graph structure M\n3. One approximation we make here is that the best hypothesis labeling `\t\u001b\n\nis\nreally dominant among all the possible labelings so that hard assignment for labelings can\nbe used. This is similar to the situation of K-means vs. mixture of Gaussian for clustering\nproblems. We evaluate this approximation in experiments.\n\n# corresponding to `\u001e\u001b\n\n# for each \n\n\u0019*/\" \n-\u0012(\n\n\u00190/! \n+\u00045\n\n\u0019*/\" \n-\u00125\n\n\u00137\n\n\u0019*/\n\n\u0019*/\n+)5\n\n\u0013\u0015\n\nQSR\fT\n\n%\u000e&\n\nand \n\n(10)\n\n\u0013\u0015\n\nG\n#\nG\n\u0001\n\u000b\nW\n%\n1\nW\n%\n#\n\u0011\n\u001d\n\u001f\nQ\n\u0013\nW\n%\n\u001d\n1\nU\n\u0013\nW\n%\n#\n\u001f\n]\n\u0017\n\u0019\n&\n\u0011\n\n\u0019\n\u0013\n^\n\u0019\n\u0013\nW\n%\n\u001d\n1\n\n\u0019\n\u0013\nW\n%\n#\n\u001f\n]\n\u0017\n\u0019\n&\n\u0011\n]\n\u0003\n'\n)\n'\n^\n\u0019\n\u001f\n^\n1\n\n\u0019\n\u0013\nW\n%\n#\n\u0011\n\u001d\n3\n\n\u0019\n\u0013\n^\n\u0019\n\u001f\n^\n\u0013\nW\n%\n\u001d\n\u001f\n]\n\u0017\n\u0019\n&\n\u0011\n]\n)\n\n\u0019\n\u0013\n^\n\u0019\n\u001f\n^\n\u0013\nW\n%\n\u001d\nG\nG\n\u000b\n`\nG\nG\nG\n\nG\n\b\n\u001f\n^\n1\n\n\u0019\n\u0013\nW\n%\n#\n\u001f\n^\n\u0013\n\n\u0019\n\u0013\nW\n%\n#\n]\n^\n\u0013\n\n\u0019\n\u0013\nW\n%\n#\n`\nG\n#\n\u0013\n\u0013\nM\nG\n\u000b\n`\nG\nG\nG\nG\n\u0011\nG\nG\n\u0011\n^\n\u0013\n\n\u0019\n\u0013\nW\n\u001d\n\u001f\n\n\u0019\n1\n^\n\u0013\nW\n^\n\u0013\nW\n\u001d\n\u001f\n\n\u0019\n1\n^\n\u0013\nW\n\n\u0019\n1\n^\n\u0013\nW\n^\n1\nW\nW\n\u001d\n`\nG\n#\nN\nM\nG\nM\n \nG\n\u0011\n\u000f\nN\n`\nG\n#\n\u0013\nM\n4\n\u000b\n-\n0\ng\n4\n\u000b\n3\nG\nM\nN\n`\nG\n#\n\u0013\nM\n\u0001\n\u000b\nW\n%\n1\nW\n%\n#\nX\n]\n\u0017\n\u0019\n&\n\u0011\n]\n\b\n\u0003\n\n\u0019\n1\n^\n\u0013\nW\n%\n\u001d\n\u0007\n\u001f\n]\n\u0017\n\u0019\n&\n\u0011\n]\n\b\n\u0003\n!\n]\n\u0011\n'\n(\n1\n\n+\n(\n-\n(\n\u001d\n\u0007\n3\n\u001e\nG\nG\nG\nG\n#\nG\nG\nG\n\u0001\n\u000b\nW\n%\n1\nW\n%\n#\n\u0011\n\u001d\nX\n]\n\u0017\n\u0019\n&\n\u0011\n\u0003\n!\n]\n\u0011\n1\n\n\u001d\n\u0007\nG\n#\n\u001b\ng\n(\n\u0013\n \nG\n#\n\u001b\nd\n(\nG\n#\n\u001b\nf\n(\nG\n\u001b\nG\n#\nG\nG\n\fThe whole algorithm can be summarized as follows. Given some random initial guess of\nis from \u0011\n\nand its parameters, then for iteration v , (v\n\n# and then compute the differential\nentropies;\nM step: use the differential entropies to run the greedy graph growing algorithm described\n\nto \ufb01nd the best labeling`\u001c\u001b\n\nuntil the algorithm converges),\n\nthe decomposable graph structure M\u0001\nE step: for each FG\nin section 3 and get M\n\n, useM\ns .\n\ns\f\u000b\n\n4.2 Dealing with missing parts (occlusion)\n\nSo far we assume that all the parts are observed. In the case of some parts missing, the\nmeasurements for the missing parts can be taken as additional hidden variables [11], and\nthe EM algorithm can be modi\ufb01ed to handle the missing parts.\n\n, let \n\nbe\n\n\u000e\u0012\u000f\n\n\u0013\u001a\u0014\n\n\u0013\u0015\u0014\n\n\u000b\u000e\n\n\t2\u0001\n\n\u001d\u001b\u000b\u000e\n\n\u000b\u0005\u0004\n\u000b\u000e\n\nvariables, we can get,\n\nand }\nG and \n\ndenote the measurements of the observed parts, \n\nFor each hypothesis `\nthe measurements for the missing parts, and \nA be the measurements of\n\u0003\u0007\u0006\nthe whole object (to reduce clutter in the notation, we assume that the dimensions can be\nsorted in this way). For each EM iteration, we need to compute\u001d\nto obtain\nG\t\b\u000b\n\ns with its parameters. Taking `\nas hidden\nthe differential entropies and then M\n\t2\u0001\n\u0013\u001a\u0014\n\u0013\u001a\u0014\n\u000b\u000e\r\n\u000b\u000e\n\n\t2\u0001\n\u0013\u0015\u0014\n\u000b\u000e\r\nWhere \u0002\n, and \u0002\n-\ba\n4 are conditional expectations with respect to \ns\f\u000b\n0 . Therefore, \ndecomposable graph structure M\n# . Since M\nforeground parts under `\n`\u001e\u001b\n4 can be computed from observed parts \ns\f\u000b\n0 .\nand covariance matrix of M\n5 Experiments\nWe tested the greedy algorithm on labeled motion capture data (Johansson displays) as in\n[8], and the EM-like algorithm on unlabeled detected features from real image sequences.\n\n.\n# and\nare the measurements of the observed\nis Gaussian distributed, conditional expec-\nand the mean\n\nG\t\b\u000b\n\n\t.\u0001\n\u000b\u000e\n\nAll the expectations\n\n\u001f\u000e\n\n\u0013\u0015\u0014\n\n\u000b\u000e\n\n\u000b\u000e\n\ns\f\u000b\n\n(12)\n\n(13)\n\n\u0013\u001a\u0014\n\n\u001d.\n\ntation\n\n4 and\n\n\t2\u0001\n\n\u0013\u0015\u0014\n\n\t2\u0001\n\u000b\u000e\n\n\u001d\u0010\u000f\n\n5.1 Motion capture data\nOur motion capture data consist of the 3-D positions of 14 markers \ufb01xed rigidly on a\nsubject\u2019s body. These positions were tracked with 1mm accuracy as the subject walked\nback and forth, and projected to 2-D.\n\nUnder Gaussian assumption, we \ufb01rst estimated the joint probability density function (mean\nand covariance) of the data. From the estimated mean and covariance, we can compute\ndifferential entropies for all the possible triplets and pairs and further run the greedy search\nalgorithm to \ufb01nd the approximated best triangulated model. Figure 2(a) shows the expected\nlikelihood (differential entropy) of the estimated joint pdf, of the best triangulated model\nfrom the greedy algorithm, of the hand-constructed model from [8], and of randomly gen-\nerated models. The greedy model is clearly superior to the hand-constructed model and the\nrandom models. The gap to the original joint pdf is partly due to the strong conditional in-\ndependence assumptions of the triangulated model, which are an approximation of the true\ndata\u2019s pdf. Figure 2(b) shows the expected likelihood using 50 synthetic datasets. Since\nthese datasets were generated from 50 triangulated models, the greedy algorithm (solid\n\n0\nG\nG\nG\n\u0002\nG\n\u0003\nG\n \nG\nA\n\u0002\n \nG\nA\nG\n\u0003\n\u0014\n\u0019\n\u001f\n\u0011\n[\n]\n\u0019\n\u0002\n\u0019\n\u001d\n\u001c\n\u0019\n\u001f\n\u0011\n[\n]\n\u0019\n\u0002\n\u0019\nZ\n\u0014\n\u0019\n\u0019\nZ\n\u0014\n\u0019\n\u001d\n!\n\u001f\n\u0011\n[\n]\n\u0019\n\u0002\n\u0019\n\n\u0019\n!\n\u001d\nZ\n\u0014\n\u0019\n\u0014\n!\n\u0019\n\u0019\n\u001d\n\u001f\n\u0003\n\n\u0019\n \n!\n\u0005\n\u0002\n\u0019\n!\n\f\n\u001d\n\u0007\n!\n\u0019\n\n\u0019\n!\n\u001d\n\n\u0019\n \n\u0005\n\n\u0019\n \n!\n\u0005\n\n\u0019\n \n\u0005\n\u0002\n\u0019\n!\n\f\n\u001d\n\u0002\n\u0019\n\f\n\u0019\n \n!\n\u0005\n\u0002\n\u0019\n\f\n\n\u0019\n!\n\f\n\u0011\nG\n\u0013\n`\nG\n\u000b\n`\n\u001b\nG\nG\n\u001b\n\u0002\nG\n\u000b\nG\n0\n\u0011\n-\n \nG\n\u0003\n\u0011\n-\n \nG\n\u0003\n \nG\nA\n\u0003\nG\n\u001b\n\u0002\n\fcurve) can match the true model (dashed curve) extremely well. The solid line with error\nbars are the expected likelihoods of random triangulated models.\n\n\u2212110\n\n\u2212120\n\nd\no\no\nh\n\ni\nl\n\n\u2212130\n\ne\nk\n\ni\nl\n \ng\no\nl\n \nd\ne\nt\nc\ne\np\nx\ne\n\n\u2212140\n\n\u2212150\n\nestimated joint pdf\n\n\u2212135\n\nbest trangulated model from greedy search\ntriangulated model used in previous papers\n\nd\no\no\nh\n\ni\nl\n\ne\nk\n\ni\nl\n \ng\no\nl\n \nd\ne\nt\nc\ne\np\nx\ne\n\n\u2212140\n\n\u2212145\n\n\u2212150\n\n\u2212155\n\n\u2212160\n\nrandomly generated triangulated models\n\n\u2212160\n0\nindex of randomly generated triangulated models\n\n3000\n\n2000\n\n2500\n\n1000\n\n1500\n\n500\n\n(a)\n\n\u2212165\n0\nindex of randomly generated triangulated models\n\n20\n\n40\n\n30\n\n60\n\n50\n\n10\n\n(b)\n\nFigure 2: Evaluation of greedy search.\n\n5.2 Real image sequences\n\nThere are three types of sequences used here: (I) a subject walks from left to right (Figure\n3(a,b)); (II) a subject walks from right to left; (III) a subject rides a bike from left to right\n(Figure3(c,d)). Left-to-right walking motion models were learned from type I sequences\nand tested on all types of sequences to see if the learned model can detect left-to-right\nwalking and label the body parts correctly. The candidate features were obtained from a\nLucas-Tomasi-Kanade algorithm [10] on two frames. We used two frames to simulate the\ndif\ufb01cult situation, where due to extreme body motion or to loose and textured clothing and\nocclusion, tracking is extremely unreliable and each feature\u2019s lifetime is short.\nEvaluation of the EM-like algorithm. As described in section 4.1, one approximation\nwe made is taking the best hypothesis labeling instead of summing over all the possible\nhypotheses (equation (11)). This approximation was evaluated by checking how the log-\nlikelihoods evolve with EM iterations and if they converge. Figure 4(a) shows the results of\nlearning a 12-feature model. We used random initializations, and each curve of Figure 4(a)\ncorresponds to one such random initialization. From Figure 4(a) we can see that generally\nthe log-likelihoods grow and converge well with the iterations of EM.\nModels obtained. Figure 4 (b) and (c) show the best model obtained after we ran the\nEM algorithms for 11 times. Figure 4(b) gives the mean positions and mean velocities\n(shown in arrows) of the parts. Figure 4(c) shows the learned decomposable triangulated\nprobabilistic structure. The letter labels show the body parts correspondence.\n\nFigure 3 shows samples of the results. The red dots (with letter labels) are the maximum\nlikelihood con\ufb01guration from the left-to-right walking model. The horizontal bar at the\nbottom left of each frame shows the likelihood of the best con\ufb01guration. The short ver-\ntical bar gives the threshold where ,\nfor all the test data. If\n\n,\b\u0007\n\n\u0006\u0005\n\n\u0002\u0004\u0003\n\n\u000e\u0001\n\n\u000b*\u0011\u0016?\n\n\t\n\nG\nF\nJ\n\nC\n\nB\nD\n\nH\n\nA\n\nL\n\nE\n\nI\n\nG\n\nF\n\nJ\n\nC\n\nB\nD\n\nH\n\nE\n\nI\n\nA\nL\n\nK\n\nH\n\nF G\n\nB\nC D\nA\n\nL\n\nI\n\nE\n\nH\n\nA\nL\n\nG\n\nB\n\nF\nC D E\n\nI\n\n(a)walking detected\n\n(b)detected\n\n(c)not detected\n\n(d)not detected\n\nFigure 3: Sample frames from left-to-right walking (a-b) and biking sequences (c-d). The dots\n(either \ufb01lled or empty) are the features selected by Tomasi-Kanade algorithm [10] on two frames.\nThe \ufb01lled dots (with letter labels) are the maximum likelihood con\ufb01guration from the left-to-right\nwalking model. The horizontal bar at the bottom left of each frame shows the likelihood of the best\ncon\ufb01guration. The short vertical bar gives the threshold for detection.\n\n\b\n\n\u0002\n\u0003\n\b\ns\n\b\ns\n#\n\u0002\nG\n\f\u2212100\n\n\u2212120\n\n\u2212140\n\n\u2212160\n\n\u2212180\n\n\u2212200\n\nd\no\no\nh\n\ni\nl\n\ne\nk\n\ni\nl\n \ng\no\n\nl\n\n\u2212220\n0\n\n5\n\n(a)\n\n40\n\n60\n\n80\n\n100\n\n120\n\n140\n\n160\n\n180\n\n200\n\n100\n\n10\n\niterations of EM\n\n15\n\n20\n\nG\n\nF\nJ\n\nC\n\nB\nD\n\nH\n\nF\n\nJ\n\nG\n\nC\n\nB\nD\n\nI\n\nE\n\nA\n\nK\n\nL\n\nI\n\nE\n\nK\n\nA\n\nH\n\nL\n\n150\n\n200\n\n250\n\n(b)\n\n(c)\n\nFigure 4: (a) Evaluation of the EM-like algorithm. Log-likelihood vs. iterations of EM for different\nrandom initializations. (b) and (c) show the best model obtained after we ran the EM-like algorithms\nfor 11 times.\nthe likelihood is greater than the threshold, a left-to-right walking person is detected.The\ndetection rate is 100% for the left-to-right walking vs. right-to-left walking, and 87% for\nthe left-to-right walking vs. left-to-right biking.\n\n6 Conclusions\nIn this paper we have described a method for learning the structure and parameters of a\ndecomposable triangulated graph in an unsupervised fashion from unlabeled data. We have\napplied this method to learn models of biological motion that can be used to reliably detect\nand label biological motion.\n\nAcknowledgments\nFunded by the NSF Engineering Research Center for Neuromorphic Systems Engineering\n(CNSE) at Caltech (NSF9402726), and by an NSF National Young Investigator Award to\nPP (NSF9457618). We thank Charless Fowlkes for bringing the Chow and Liu\u2019s paper to\nour attention. We thank Xiaolin Feng for providing the real image sequences.\n\nReferences\n[1] Y. Amit and A. Kong, \u201cGraphical templates for model registration\u201d, IEEE Transactions on\n\nPattern Analysis and Machine Intelligence, 18:225\u2013236, 1996.\n\n[2] C.K. Chow and C.N. Liu, \u201cApproximating discrete probability distributions with dependence\n\ntrees\u201d, IEEE Transactions on Information Theory, 14:462\u2013467, 1968.\n\n[3] T.M. Cover and J.A. Thomas, Elements of Information Theory, John Wiley and Sons, 1991.\n[4] N. Friedman and M. Goldszmidt, \u201cLearning bayesian networks from data\u201d, Technical report,\n\nAAAI 1998 Tutorial, http://robotics.stanford.edu/people/nir/tutorial/, 1998.\n\n[5] M.I. Jordan, editor, Learning in Graphical Models, MIT Press, 1999.\n[6] M. Meila and M.I. Jordan, \u201cLearning with mixtures of trees\u201d, Journal of Machine Learning\n\nRearch, 1:1\u201348, 2000.\n\n[7] Y. Song, X. Feng, and P. Perona, \u201cTowards detection of human motion\u201d, In Proc. IEEE CVPR\n\n2000, volume 1, pages 810\u2013817, June 2000.\n\n[8] Y. Song, L. Goncalves, E. Di Bernardo, and P. Perona, \u201cMonocular perception of biological\nmotion in johansson displays\u201d, Computer Vision and Image Understanding, 81:303\u2013327, 2001.\n[9] Nathan Srebro, \u201cMaximum likelihood bounded tree-width markov networks\u201d, In UAI, pages\n\n504\u2013511, San Francisco, CA, 2001.\n\n[10] C. Tomasi and T. Kanade, \u201cDetection and tracking of point features\u201d, Tech. Rep. CMU-CS-91-\n\n132,Carnegie Mellon University, 1991.\n\n[11] M. Weber, M. Welling, and P. Perona, \u201cUnsupervised learning of models for recognition\u201d, In\n\nProc. ECCV, volume 1, pages 18\u201332, June/July 2000.\n\n[12] Markus Weber, Unsupervised Learning of Models for Object Recognition, Ph.d. thesis, Caltech,\n\nMay 2000.\n\n\f", "award": [], "sourceid": 2106, "authors": [{"given_name": "Yang", "family_name": "Song", "institution": null}, {"given_name": "Luis", "family_name": "Goncalves", "institution": null}, {"given_name": "Pietro", "family_name": "Perona", "institution": null}]}