{"title": "Conditional Models on the Ranking Poset", "book": "Advances in Neural Information Processing Systems", "page_first": 431, "page_last": 438, "abstract": null, "full_text": "Conditional Models on the Ranking Poset\n\nGuy Lebanon\n\nSchool of Computer Science\nCarnegie Mellon University\n\nPittsburgh, PA 15213\n\nJohn Lafferty\n\nSchool of Computer Science\nCarnegie Mellon University\n\nPittsburgh, PA 15213\n\nlebanon@cs.cmu.edu\n\nlafferty@cs.cmu.edu\n\nAbstract\n\nA distance-based conditional model on the ranking poset is presented\nfor use in classi\ufb01cation and ranking. The model is an extension of the\nMallows  model, and generalizes the classi\ufb01er combination methods\nused by several ensemble learning algorithms, including error correcting\noutput codes, discrete AdaBoost, logistic regression and cranking. The\nalgebraic structure of the ranking poset leads to a simple Bayesian inter-\npretation of the conditional model and its special cases. In addition to a\nunifying view, the framework suggests a probabilistic interpretation for\nerror correcting output codes and an extension beyond the binary coding\nscheme.\n\n1 Introduction\n\n. A gener-\nClassi\ufb01cation is the task of associating a single label\na full or partial\nalization of this problem is conditional ranking, the task of assigning to\n. This paper studies the algebraic structure of this problem, and\nranking of the items in\nproposes a combinatorial structure called the ranking poset for building probability models\nfor conditional ranking.\n\nwith a covariate\n\n\u0001\u0003\u0002\u0005\u0004\n\nIn ensemble approaches to classi\ufb01cation and ranking, several base models are combined to\nproduce a single ranker or classi\ufb01er. An important distinction between different ensemble\nmethods is whether they use discrete inputs, ranked inputs, or con\ufb01dence-rated predic-\n, and no\ntions. In the case of discrete inputs, the base models provide a single item in\npreference for a second or third choice is given. In the case of ranked input, the base clas-\nsi\ufb01ers output a full or partial ranking over\n. Of course, discrete input is a special case\nof ranked input, where the partial ranking consists of the single topmost item. In the case\nof con\ufb01dence-rated predictions, the base models again output full or partial rankings, but\nin addition provide a con\ufb01dence score, indicating how much one class should be preferred\nto another. While con\ufb01dence-rated predictions are sometimes preferable as input to an en-\nsemble method, such con\ufb01dence scores are often not available (as is typically the case in\nmetasearch), and even when they are available, the scores may not be well calibrated.\n\nThis paper investigates a unifying algebraic framework for ensemble methods for clas-\nsi\ufb01cation and conditional ranking, focusing on the cases of discrete and ranked inputs.\nOur approach is based on the ranking poset on\n, which consists of\nthe collection of all full and partial rankings equipped with the partial order given by re-\n\nitems, denoted\n\n\b\n\t\n\n\u0006\n\u0006\n\u0004\n\u0004\n\u0004\n\u0007\n\fgives rise to\n\ufb01nement of rankings. The structure of the poset of partial ranking over\nnatural invariant distance functions that generalize Kendall\u2019s Tau and the Hamming dis-\n\ntance. Using these distance functions we de\ufb01ne a conditional model \u0002\u0001\u0004\u0003\u0006\u0005\b\u0007\n\t\f\u000b\u000e\r\u0010\u000f\u0011\u000f\u0011\u000f\u0012\r\u0013\t\u0015\u0014\u0017\u0016\nwhere \u0005\u0018\r\u0019\t\f\u000b\u001a\r\u0010\u000f\u0011\u000f\u0011\u000f\u0012\r\u0013\t\u0015\u0014\n. This conditional model generalizes several existing models\nfor classi\ufb01cation and ranking, and includes as a special case the Mallows  model [11]. In\naddition, the model represents algebraically the way in which input classi\ufb01ers are combined\nin certain ensemble methods, including error correcting output codes [4], several versions\nof AdaBoost [7, 1], and cranking [10].\n\nIn Section 2 we review some basic algebraic concepts and in Section 3 we de\ufb01ne the rank-\ning poset. The new model and its Bayesian interpretation are described in Section 4. A\nderivation of some special cases is given in Section 5, and we conclude with a summary in\nSection 6.\n\n2 Permutations and Cosets\n\nWe begin by reviewing some basic concepts from algebra, with some of the notation and\nde\ufb01nitions borrowed from Critchlow [2].\n\nIdentifying the items to be ranked\n\n\u0007 \u001f , then \u0005!\u0003\u0006\"\n\u0016 denotes the rank given to item \" and \u0005$#\n\npermutation of \u001d\u001e\u001c\u0017\r\u0010\u000f\u0011\u000f\u0011\u000f\f\r\nthe item assigned to rank \" . The collection of all permutations of\nThe subgroup of %\n\nwith the numbers \u001c\u0017\r\u0010\u000f\u0011\u000f\u0011\u000f\u001b\r\n, if \u0005 denotes a\n\u0003\u0006\"\n\u0016 denotes\n. The multiplicative notation \u0005'&\u000e\t)(*\u0005\u0002\t\nconsisting of all permutations that \ufb01x the top + positions is denoted\n\nabelian symmetric group of order\nis used to denote function composition.\n\n\u0011\u000f\u0010\u000f\u0011\u000f\u001b\r\n, denoted %\n\n-items forms the non-\n\n\u0014 ; thus,\n\nThe right coset\n\n\u0014,(-\u001d\u0010\u0005\n\n\u0002.%\n\n\u0007\n\u0005!\u0003\u0006\"\n\u0016!(/\"0\r \" (1\u001c\u0017\r\u0010\u000f\u0011\u000f\u0010\u000f\u001b\r\u0019+2\u001f\u001e\u000f\n\n(1)\n\n(2)\n\ntop-ranked items.\n\nAn ordered partition of\nto\nin the \ufb01rst position,\n\n\u0007\u0012>\n\n\u00148\u001f\nis equivalent to a partial ranking, where there is a full ordering of the +\nThe set of all such partial rankings forms the quotient space %\n\u0014 .\n\n\u00143\u00054(5\u001d3\t\f\u00056\u00077\t\n\n\t:9\u0017%\n\n\u0002'%\n\nis a sequence ;<(\n\n\u0007 \u000b\u000e\r\u0010\u000f\u0011\u000f\u0010\u000f\u0002\n\n\u0007\u0002= of positive integers that sum\n\nBO\u001c\u0017\r\u0010\u000f\u0011\u000f\u0011\u000f\f\n\nfor which\n\n\u0007\u0012\u0014IHD\u000bE(\n(*\u001d\n\n(*\u001dG\u001cG\r\u0011\u000f\u0010\u000f\u0011\u000f\f\n\nitems\nitems in the second position and so on. No further information\nitems is\n. More formally, let\n\n. Such an ordered partition corresponds to a partial ranking of type ; with\nis conveyed about orderings within each position. A partial ranking of the top +\na special case with ?@(A+CB*\u001c\u0017\r\n\u0007D\u000bE(F\u000f\u0011\u000f\u0011\u000f (\n\u0007 \u001f .\n>M(N\u001d\n\u001fG\r\nBO\u001cG\r\u0011\u000f\u0011\u000f\u0010\u000f\f\r\nThen the subgroup %RQOS\n\tGTVUW&\u0011&\u0011&2UX%\nLR[ holds for each \" ; that is, all permutations that only permute\nLR[\nthe set equality \u0005!\u0003\n\u0016\\(\nwithin each LR[ . A partial ranking of type ;\n\u0005 and the set of such\npartial rankings forms the quotient space %\n\n\u0007.JK+\n\u0007\u0012\u0014@(F\u001cG\r\n\u0007\u0002>\u000e\u001fG\r\u0010&\u0011&\u0010&\u001b\r\nBO&\u0011&\u0010&PB\n\tZY contains all permutations \u0005\nis equivalent to a coset %\n\t:9\u0017%]Q .\n\nWe now describe a convenient notation for permutations and cosets. In the following, we\nlist items separated by vertical lines, indicating that the items on the left side of the line are\npreferred to (ranked higher than) the items on the right side of the line. For example, the\n#\fh\nmay thus be\nranked in\n\nb . A partial ranking %gf\n\nis denoted by ^:\u0007e\u001c\u001e\u0007\nQ where ;l(mb:\r\u0019^ with items \u001c\u0017\r\u0019b\u0004\r\u0019k\n\nj\u0015\r\u0019k . A classi\ufb01cation\n\npermutation \u0005!\u00037\u001c3\u0016M(*^\u0004\r7\u0005!\u0003_^G\u0016`(a\u001c\u0017\r\u0013\u0005!\u0003cb\u001e\u0016d(*b\nis denoted by b2\u0007\nwhere the top 3 items are b\u0004\r\u0019^\u0004\r\u0011\u001c\n^\u0015\u0007i\u001c8\u0007\ndenoted by b\u0015\u0007e\u001c\u0017\r\u0019^\u0004\r7j\u0015\r\u0019k . A partial ranking %\nthe \ufb01rst position is denoted by \u001cG\r\u0013b:\r\u0019k:\u0007\n^n\r\u0013j .\nA distance function o on %\nis a function oNpq%\nproperties: o2\u0003\u0006\u0005\u0018\r7\u0005\u0012\u0016g(wv , o2\u0003\u0006\u0005\u0018\r\u0013\t\u0002\u0016\bxyv when \u0005az\n(A\t\n\nthat satis\ufb01es the usual\n\n\tNsut\n\n\trUl%\n, o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016C(wo2\u0003c\t\f\r\u0013\u0005\u0012\u0016 , and the triangle\n\n\u0002W%\n\n\u0004\n\u0002\n\b\n\t\n\u0001\n\u000b\n\u0001\n\t\n\u0007\n\u000b\n\u0007\n\u0007\n\t\n\t\n%\n\t\n#\n%\n\t\n#\n\t\n%\n\t\n#\n\t\n#\n\t\n#\n\u0007\n\u0007\n\u0007\n\u000b\nL\n\u000b\n\u0007\n\u000b\nL\n\u0007\n\u000b\n\u0007\n\u000b\nB\nL\n=\n\u0007\n\u000b\n\u0007\n=\n#\n\u000b\n(\n%\n\t\nQ\n\u0005\n\u0001\nh\n\t\n\fPSfrag replacements\u0002\u0001\n\u0007\u0001\n\n\u0003\u0004\u0001\n\u0003\u0006\b\n\n\u0002\u0001\n\u0007\b\n\n\u0005\u0006\u0001\n\u0003\u0006\u0001\n\n\u0003\u0006\u0001\n\u0007\b\n\n\u0002\u0001\n\u0005\u0006\u0001\n\n\u0005\u0004\u0001\n\u0003\u0006\u0001\n\n\u0002\u0001\n\u0007\b\n\n\u0007\b\n\n\u0003\u0002\b\n\n\u0003\u0006\u0001\nPSfrag replacementsT\n\t\n\u0005\u0004\u0001\nT\n\t\n\n\u0005\u0004\u0001\n\u0007\b\n\n\u0003\u0004\u0001\n\u0005\u0004\u0001\n\n\u0005\u0006\u0001\n\u0003\u0002\b\n\nT\u0002\f\nT\u0002\f\nT\u0007\t\n\n\u000b\u0004\f\n\u000b\u0004\f\n\u000b\u0006\t\n\n\u0006\f\n\r\u0002\t\n\r\u0006\f\n\n\u000b\u0004\f\n\u000b\u0004\f\n\n\u0006\f\n\r\u0002\t\n\nT\u0002\f\n\n\u000b\u0002\t\n\n\u0004\f\n\nT\u0007\t\n\n\u000b\u0002\t\n\n\u0006\t\n\nFigure 1: The Hasse diagram of\nof the lines are dotted for easier visualization.\n\n(left) and a partial Hasse diagram of\n\nT\u0007\f\nT\n\t\nT\u0007\f\n\n\u0004\f\n\r\u0004\f\n\u000b\u0006\t\n\n\u000b\u0004\f\n\u000b\u0004\f\n\r\u0006\t\n\nT\u0007\f\nT\n\t\n\n\u0004\f\n\r\u0004\f\n\n\u000b\u0002\t\n\u000b\u0002\t\n\n\b\u0010\u000f (right). Some\n\nof the items\n\no2\u0003\u0006\u0005\u0018\r\u0006\u0012:\u0016\u001bB\n\n. In addition, since the indexing\n.\n\nis arbitrary, it is appropriate to require invariance to relabeling of\n\no2\u0003\u0013\u0012\u0004\r\u0019\t\u0002\u0016 for all \u0005\u0018\r\u0019\t\f\r\u0006\u0012\ninequality o\u0015\u0003c\u0005\u0018\r\u0013\t\u0002\u0016\n\u0002.%\n(/o2\u0003\u0006\u0005\u0014\u0012:\r\u0013\t\u0015\u0012:\u0016 , for all \u0005\u0018\r\u0019\t\f\r\u0006\u0012\n\u00018\u000b\u000e\r\u0010\u000f\u0011\u000f\u0011\u000f\u001b\r\nFormally, this amounts to right invariance o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016\nis Kendall\u2019s Tau\u0016]\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016 , given by\nA popular right invariant distance on %\n\u001a\u001c\u001b\n[\u0019\u0018\n[\u001e\u001d\n\u00062\u0016E(4v otherwise [8]. Kendall\u2019s Tau \u0016]\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016 can be\n, or the minimum\n\u000b . An adjacent transposition\n\ufb02ips a pair of items that have adjacent ranks. Critchlow [2] derives extensions of Kendall\u2019s\n\n\u0003 \u001f_\u00167\u0016\nwhere \u001d\ninterpreted as the number of discordant pairs of items between \u0005 and \t\nnumber of adjacent transpositions needed to bring \u0005$#\nTau and other distances on %\n3 The Ranking Poset\n\n\u0016]\u0003c\u0005\u0018\r\u0013\t\u0002\u0016\n\u0006,xAv and \u001d\n\nto distances on partial rankings.\n\n\u0003c\"\n\u0016 JX\u0005\u0002\t\n\nto \tD#\n\n\u00062\u0016E(\n\n\u0003\u0006\u0005\u0002\t\n\n\u0002'%\n\nfor\n\n\t\f\u000f\n\n(3)\n\nand\n\nand\n\n, (2) if\n\n\u0002,!\n\nwhen\n\n\u00012\r+'\n\n\u0001&#('\n\n\u0006%#\n\u00060-\n\nis a binary relation\n\nfor all\nand write\n\nWe \ufb01rst de\ufb01ne partially ordered sets and then proceed to de\ufb01ne the ranking poset. Some of\nthe de\ufb01nitions below are taken from [12], where a thorough introduction to posets can be\nfound.\n\nand\n. We write\n\n\u0006%#\n\u0006/.\n\u0006\u0012\n\nthe covering relation. In addition, we require that if\n\nthat satis\ufb01es (1)\nthen\n. A\ncovers\n\ufb01nite poset is completely described by the covering relation. The planar Hasse diagram of\nare the nodes and the edges are given by\n\nis a set and#\nA partially ordered set or poset is a pair \u0003\"!d\r$#q\u0016 , where!\n\u0006%#\n\u0001&#\n, and (3) if\nthen\n\u0006,#\n\u0006,-\n\u0001\n\u0006)#*'\n\u0006)(\nwhen\nand there is no'\n\u00060-)' and'2-\n\u00021!\n\u0006,z\nsuch that\nis the graph for which the elements of !\n\u00064.\n\n\u0003\"!d\r3#q\u0016\nis the poset in which the elements are all possible cosets %65\u0017\u0005\nwhere7\nre\ufb01nement; that is, \u0005&-l\t\nif we can get from \u0005\nto \t by adding vertical lines. Note that\n\b\n\t\n\u0007 \u001f ordered by partition re\ufb01nement\nis different from the poset of all set partitions of \u001dG\u001c\u0017\r\u0010\u000f\u0011\u000f\u0010\u000f2\r\nthe order of the partition elements matters. Figure 1 shows the Hasse diagram\n\b\u0010\u000f .\nA subposet \u0003 8E\r$#:9!\u0016 of \u0003\"!d\r3#<;$\u0016\n\u0006@#:;\n\u0006&#:9\nis de\ufb01ned by8>=?!\nA chain is a poset in which every two elements are comparable. A saturated chain A\n\nand a portion of the Hasse diagram of\n\nis an ordered partition of\n\n,\nis de\ufb01ned by\n\n. The partial order of\n\nis drawn higher than\n\nThe ranking poset\n\nsince in\nof\n\n. We say that\n\nif and only if\n\n.\n\n.\nof\n\nand \u0005\n\nthen\n\nand\n\n\u0005\n\u0003\n\u0005\n\u0003\n\u0005\n\u0003\n\u0005\n\u0003\n\n\u0005\n\n\n\u0005\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\u000e\n\b\nh\n\u0011\n\t\n\u0001\n\t\n\u0004\n\t\n(\n\t\n#\n\u000b\n\u0017\n\u000b\n\u0017\n#\n\u000b\n#\n\u000b\n\u0003\n\u001c\n\u0003\n\u000b\n\t\n\u0006\n\u0001\n\u0006\n\u0001\n\u0001\n\u0001\n(\n\u0001\n\u0001\n\u0006\n\u0001\n\u0001\n\u0001\n\u0001\n\u0001\n\u0006\n\b\n\t\n\u0007\n\u0002\n%\n\t\n\b\n\t\n\b\n\t\n\b\nh\n\u0001\n\u0001\n\fgraded poset of rank\n\n\u0002\n\b\n\n\t\b\u0007\n\nhowever, remains.\n\non\n\n, given by\n\n\u000b\u001e\u0007I&\u0010&\u0011&\u0010\u0007i\u001d\n\n\u0014 .\n\nif\n\nif\n\n.\n\n\u00062\u0016\u0012B\n\nis a graded poset of rank\n\nto denote the subposet of\n\nis a sequence of elements\n\nthat satisfy\n\n\t\b\u0007\n\n\t\u0002\u0007\n\n\u001dG\u001cG\r\u0011\u000f\u0010\u000f\u0011\u000f\u001b\n\n\u0007 \u001f .\n\n\u0006\f\u0016\n\u001d\u001e\u001c\u0017\r\u0010\u000f\u0011\u000f\u0011\u000f\f\n\nis a right invariant function on\n\n&$\u0012y(5\u001d3\u0012\u0002\u0003\n\nconsisting of \u001d\n\n\u000b . Classi\ufb01cations \"\n\n\u000b . Other elements of\n\n4 Conditional Models on the Ranking Poset\n\nWe now present a family of conditional models de\ufb01ned in terms of the ranking poset. To\n\nall of which are incomparable, are denoted by\ngrade\n\n.m&\u0010&\u0011&:.\nthat contains it. A\n. In a graded\nis a\n(mv\n\nIt is easy to see that\nis the number of vertical lines in its denotation. We use\n\nis a poset in which every maximal chain has length\n\u0006\f\u0016\n\n\u0007XJ,\u001c and the rank of every element\n\u0014 . Full orderings occupy the topmost\n\u000b are\n\n\u0006\u0002\n\u0006\u0001\u001e\r\u0011\u000f\u0010\u000f\u0011\u000f\u001b\r\nlength +\nis a maximal chain if there is no other saturated chain of!\nA chain of!\n\u001d3v:\r\u0011\u000f\u0010\u000f\u0011\u000f\f\r\u0005\u0004\u001e\u001f such that\u0003\f\u0003\nposet, there is a rank or grade function\u0003)p\u001e!As\nminimal element and\u0003\f\u0003\n\u0001\u0004\u0016!(\u0006\u00032\u0003\n\u00064.\n\u0002\f\u000b\\\u001f . In particular, the elements in the + th grade,\n\u0007\n\u0003\f\u0003\n\t\b\u0007\n\u0007 \u001f\u000e\r\u0010\" reside in\n\t\b\u0007\nmultilabel classi\ufb01cations\u000fl\u0007i\u001dG\u001c\u0017\r\u0010\u000f\u0011\u000f\u0010\u000f2\r\n\u0007 \u001f\u000e\r\u0010\u000f where\u000f\no\u0015\u0003c\u0005\u0014\u0012\u0004\r\u0013\t\u0015\u0012:\u0016 for\nbegin, suppose that o\n. That is, o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016\nand\u0012\nall \u0005\u0018\r\u0013\t\naction of %\n[\u0015\u0013\nT\u0010\u00160\u001f\u0012\u0011\nT\u001a\u0007\n\u0016I\u001f\u0010\u0011\n\u000b\u000e\u001f\u0012\u0011\nT3\u001f\u0010\u0011\nTZ\u0007i\u001d\n&\u0011&\u0010&3\u0007i\u001d3\u0012\u0002\u0003\n\u0012\u0002\u0003\nThe function o may or may not be a metric; its interpretation as a measure of dissimilarity,\nWe will examine several distances that are based on the covering relation . of\nand up moves on the Hasse diagram will be denoted by\u0016\no de\ufb01ned in terms of\u0016\ngroup action of %\ncommutes with\u0016\nis, the group action of %\n\u001b\u001d\u001c\n\n[\u0014\u0013\n\u001f\u0010\u0011\nand\u0017 moves is easily shown to be right invariant because the\n\u0018\u001a\u0019\n\n\u000b\u000e\u00160\u001f\u0010\u0011\n\u000b\u0017\u0007\nand\u0017\nand\u0017 moves:\n\u0018\u001a\u0019\n\ndoes not change the covering relation between any two elements; that\n\n. Here right invariance is de\ufb01ned with respect to the natural\n\n. Down\nrespectively. A distance\n\n\u0013\t\u0015>\u0017\r\u0010\u000f\u0011\u000f\u0010\u000f\u001b\r\u0013\t\n\nin the same manner.\n\nrankings \t\n\n, which will be the \u201ccarrier density\u201d\n\nis essential since we want to treat all\n\nWe are now ready to give the general form of a conditional model on\n\nWhile the metric properties of o are not required in our model, the right invariance property\n. Let o be an\ninvariant function, as above. The model takes as input +\n\u0014 contained\n\nin some subset \n\t\u0002\u0007\nor default model. Then o and\u0004\" specify an exponential model \n\u0003\u0006\u0005\b\u0007$#M\u0016 given by\n!\u0002o2\u0003\u0006\u0005\u0018\r\u0019\t+!\u000e\u0016$3\n\u000b21\n. The term%\n\r\u0005#M\u0016\n\u0003'&\nwhere&\n\u0016\n3\n\nof the ranking poset. For example, each \t\b! could be an element of\n\u0003c\u00056\u0007$#M\u0016\n\u0002:9\n\u0003'&D\r(#d\u0016\n\n\u000b . Let\u0004\" be a probability mass function on\n\u0004)8\u0003\u0006\u0005\u0012\u0016+*\",.-\n\r(#M\u0016\n, and \t.!\n\u0002; \n\u00026587\n\u0003\u0006\u0005\u0012\u0016+*\",.-\n\nnormalizing constant\n\no\u0015\u0003c\u0005\u0018\r\u0013\t\n\nis the\n\n, \u0005\n\n(6)\n\n(7)\n\nJ\u0011J:J\u0004J0s\n\nJ\u0011J:J\u0004J0s\n\non\n\nJ\u0010J\u0004J\u0004J0s\n\nJ\u0010J\u0004J\u0004J0s\n\n(4)\n\n(5)\n\n\b\n\t\n\n\u0003'&\n<.=?>\n\n\u0006\n\u0014\n\u0002\n!\n.\n\u0006\n\u000b\n\u0006\n\u0007\n\u0007\n\u0006\n\u001c\n\u0001\n\b\n\t\n\b\n\t\n\b\n\t\n\u0006\n\t\n\b\n\b\n\t\n#\n\u0007\n\b\n\b\n=\n\b\n\t\n(\n\u0002\n\b\n\t\n\u0002\n%\n\t\n\t\n\b\n\t\n\u001d\n\u0001\n[\n\u0001\n[\n\u0001\n\u0013\n\u0001\n[\n\u001d\n\u0001\n[\n\u0001\n\u0013\n\u000f\n\t\n\t\n\b\n\t\n\b\n\t\n\b\n\t\n\u001c\n\u001e\n\u001c\n\u001c\n\u001e\n\u001b\n\b\n\t\n\u0018\n\u0019\n\b\n\t\n\b\n\t\n\b\n\t\n\u001f\n\u001c\n\u001c\n\u001e\n\u001c\n\u001c\n\u001e\n\u001f\n\b\n\t\n\u0018\n\u0019\n\b\n\t\n\u0001\n[\n\b\n\t\n\u000b\n=\n\b\n\t\n\b\n\t\n#\n\b\n\t\n\u0001\n\n\u0001\n(\n\u001c\n%\n/\n0\n\u0014\n\u0017\n!\n\u0018\n4\nt\n\u0014\n=\n\b\n\t\n=\n\b\n\t\n%\n(\n\u0017\n\u0004\n\n/\n0\n\u0014\n\u0017\n!\n\u0018\n\u000b\n1\n!\n!\n4\n\u000f\n\f\u0014 ,\n\nThus, conditional on#\n#M\u0016 forms a probability distribution over \u0005\nGiven a data set \u0001\n\u0016\t\b , the parameters1\n&\f\u0016\nmizing the conditional loglikelihood \n\u0017\u0003\n(\f\u000b\nor posterior. Under mild regularity conditions, \n\u0017\u0003\n\n9(=\n5 will typically be selected by maxi-\n\n[\u000e\r\u0007\u000f\u0011\u0010\n\u0016 , a marginal likelihood\n\u0016 will be convex and have a unique global\n\n\u0005#\u0006\u0002\n\nmaximum.\n\n2\u0001\u0004\u0003\u0006\u0005\n\n\u0003c\u0005\u0003\u0002\n\n\u0007$#\n\n\u00037&3\u0007\n\n[\u0005\u0004\n\n[\u0005\u0004\n\n[\u0005\u0004\n\n[\u0007\u0004\n\n.\n\n4.1 A Bayesian interpretation\n\nWe now derive a Bayesian interpretation for the model given by (6). Our result parallels\nthe interpretation of multistage ranking models given by Fligner and Verducci [6]. The key\nfact is that, under appropriate assumptions, the normalizing term does not depend on the\npartial ordering in the one-dimensional case.\n\n.\n\n[\u0005!\n\n\u0001\u0018\u0017\n\n\u0001\u0018\u0017\n\n(8)\n\nthen\n\nt .\n\n&\u0011&\u0011&3\u0007\n\nof%\n\n. If%\n\n<\u001d\u001c\u001e\u0019\n\n, it follows that\n\nProposition 4.1. Supposethato isrightinvariantandthat isinvariantundertheaction\nactstransitivelyon9\n<\u001a\u0019\n=\u0014\u0013\u0016\u0015\n=\u0014\u0013\u001b\u0015\n9 and1\nforall\u0005\u0018\r7\u0005 \u001f\nProof. First, note that since \n. Indeed, \n by the invariance assumption, and :7\nfor each\u0012\n we have \t\"\u001f!(<\u001d3\u0012\f#\nsince for \tl(<\u001d\nthat \t\"\u001f\u0019\u0012\b(/\t\nT\u000e\u001f8\u0007\nacts transitively on9\nsuch that \u0005\"#E(\n, for all \u0005\u0018\r\u0013\u0005 \u001f\nNow, since %\n<\u001a\u0019\nWe thus have that%\n\u0001\u0018\u0017\n=\u0014\u0013\n<\u001d%\u000e\u0019\n\u0001\u0018\u0017\n=\u0014\u0013\n<\u001d\u001c'\u0019\n\u0001\u0018\u0017\n=\u0014\u0013\n<\u001d\u001c'\u0019\n\u0001\u0018\u0017\n=\u0014\u0013\n\u0016 since the normalizing constant for \u0005\n\nis invariant under the action of %\nT\u0011\u00160\u001fn\u0007I&\u0011&\u0010&\u0010\u0007i\u001d3\u0012\f#\n\u0002.%\n\n(by right invariance of& )\n\nfact depend on \u0005\nThe underlying generative model is given as follows. Assume that \u0005\n\n(by invariance of(\n\nthere is #\n\nis drawn from\n\ndoes not in\n\n\u0012\u000e%\n\u0012\u000e%\n\n\u0005$\u001f .\n\n\u0013\u0005\u0012\u0016\n\nsuch\n\n(10)\n\n(12)\n\n(11)\n\n(9)\n\n\u00160\u001f\n\n[\u001e!\n\n)\n\n\u0014 are independently drawn from generalized Mallows\n\u0005\u0012\u0016\nis given by\n\n(13)\n\n. Then under the conditions of Proposition 4.1, we have from Bayes\u2019 rule\n\n.\n\nmodels\n\n7\u0005\u0012\u0016M(\n\nThus, we can write%\nthe prior\u0004\"\u001e\u0003c\u0005\u0012\u0016 and that \t\nwhere \t.!\n\u0002\u001d \n\u0004)8\u0003\u0006\u0005\u0012\u0016,+\n\u0003_\t+!\n\u000b\u001d<+=?>\n\u0004)8\u0003\u0006\u0005\u0012\u0016,+\n\nthat the posterior distribution over \u0005\n\u0005\u0012\u0016\n\n\u0010\u000f\u0011\u000f\u0010\u000f\u0002\r\u0013\t\n\u0001*)\n\n\u0003_\t+!\n\n\u0003c\t.!\n\n\u0005\u0012\u0016\n\n\u0001*)\n\n\u0001*)\n\n<\u001a\u0019\n<+=?>\n\n<\u001a\u0019\n\n!3\u0016\n\u0004)8\u0003\u0006\u0005\u0012\u0016\n\u0003\u0006\u0005\b\u0007$#M\u0016\n\n!3\u00160#\n\n\u0001.)/\u0017\n\n\u0001*)\n\n<0\u0019\n\n(14)\n\n(15)\n\n\u0003c\u0005\u0012\u0016\n\n\u0002\n \n\u0001\n\u0002\n\b\n\t\n\u0002\n\u0002\n\u0002\n1\n\t\n\t\n\u0017\n\u0012\n\u0002\n\u0012\n\u0004\n(\n\u0017\n\u0012\n\u0002\n\u0012\n\u0004\n\u0002\n\u0002\n=\n\b\n\t\n\t\n \n\u0012\n(\n \n\u0002\n%\n\t\n\u0012\n7\n \n\u0012\n\u0001\n[\n\u001d\n\u0001\n\u001f\n\u0002\n\u000b\n\u0003\n\u0001\n[\n\u000b\n\u0003\n\u0001\n\u0002\n \n\t\n\u0002\n9\n\t\n\u0003\n1\n(\n\u0017\n\u0012\n\u0015\n\u0002\n\u0012\n\u0004\n(\n\u0017\n\u0012\n\u0015\n\u0002\n\u0004\n(\n\u0017\n\u0012\n\u0015\n\u0002\n\u0004\n(\n\u0017\n\u0012\n\u0015\n\u0002\n\u0012\n\u0004\n\u0003\n1\n%\n\u0003\n1\n\u0002\n9\n\u0002\n9\n\u000b\n\n\u0007\n(\n\u001c\n%\n\u0003\n1\n\u0015\n\u0001\n)\n\u0017\n\u0002\n\u0012\n)\n\u0004\n!\n\n\u0007\n!\n\n\u0007\n(\n\u0015\n-\n)\n\u0001\n)\n\u0017\n\u0002\n\u0001\n)\n\u0004\n+\n!\n%\n\u0003\n1\n\u000b\n+\n!\n%\n\u0003\n1\n!\n\u0016\n#\n\u000b\n\u000b\n\u0004\n\n\u0015\n-\n)\n\u0002\n\u0004\n(\n\n\u0001\n\f\u0003c\u0005\u0012\u0016\n\nto \t\n\n5\u0002\u0001\n\n\t:9G%\n\n^:\u0007\n\nb\u0004\r\u0013b2\u0007\n\n^\u0015\u0007i\u001c\u000e\u0016\n\nacts\n\n(16)\n\n\u0001*)\n\n\u0003\n&\u000e\u0007\n\n\t:9G%\n\n\u0003\u0006\u0005\u0012\u0016\n\n /(\n\n5 Special Cases\n\n) moves on the Hasse diagram of\n\nas is assumed in the special cases of the next section.\n\nmodels may be easily derived, corresponding to the exponential loss used in boosting.\n\nThis section derives several special cases of model (6), corresponding to existing ensem-\nin the\nis taken to be uniform, though the extension\nis immediate. Following [9], the unnormalized versions of all the\n\nand o\n\n,and%\n\u0005\u0012\u0016 ,withprior\u0005\n\n\u0003\n&\u000e\u0007$#M\u0016 de\ufb01ned in equation (6) is the posterior under\n\u0004\" .\n5 and \n\nWe thus have the following characterization of \n\n\u00037&3\u0007$#M\u0016 .\nProposition 4.2. Ifo isrightinvariant, isinvariantundertheactionof%\ntransitively on9 , then the model\nindependentsamplingofgeneralizedMallowsmodels,\t2!\nThe conditions of this proposition are satis\ufb01ed, for example, when91(*%\nble methods. The special cases correspond to different choices of9]\r\nde\ufb01nition of the model. In each case\u0004\nto non-uniform\u0004\n\r(9\nLet5N(\nup (\u0016\n(N%\nmove over the Hasse diagram, o\u0015\u0003c\u0005\u0018\r\u0013\t\u0002\u0016\n\u0016]\u00037\u001c\u001e\u0007\n^\u0015\u0007i\u001c8\u0007\n\nb\u0015\u0007e\u001c\n^:\u0007\nIn this case model (6) becomes the cranking model [10]\n\nb and the corresponding path in Figure 1 is\n\n5.1 Cranking and Mallows  model\n\nmodel is independent sampling of \t\n\n. Since adjacent\ntranspositions of permutations may be identi\ufb01ed with a down move followed by an up\n\n, and let o\u0015\u0003c\u0005\u0018\r\u0013\t\u0002\u0016 be the minimum number of down-\nis equal to Kendall\u2019s Tau \u0016]\u0003c\u0005\u0018\r\u0013\t\u0002\u0016 . For example,\n^\u0015\u0007i\u001cG\r\u0013b\n<\u001a\u0019\n)\u0006\u0005\n)\u0004\u0003\nThe Bayesian interpretation in this case is well known, and is derived in [6]. The generative\nfrom a Mallows  model whose location parameter\n! . Other special cases that fall into this category are the\n\n#V\r\n&2\u0016\nis \u0005 and whose scale parameter is1\nLet5a(\n,9\n\u000b , and let o\u0015\u0003c\u0005\u0018\r\u0013\t\u0002\u0016 be the minimum number of up-down\n) moves in the Hasse diagram. Since9\n(\u0017\n\u0012:\r0%\nremoved, the model becomes discrete AdaBoost.M2; that is, o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002!8\u0003\n\n(discrete) multiclass weak learner\nin the usual boosting notation. See [9]\nfor details on the correspondence between exponential models and the unnormalized mod-\nels that correspond to AdaBoost.\n\nIn this case model (6) becomes equivalent to the multiclass generalization of logistic re-\ngression. If the normalization constraints in the corresponding convex primal problem are\n\nmodels of Feigin [5] and Critchlow and Verducci [3].\n\n\t:9\u0017%\nif\u0012\f#\n\u0003\n\u001c\u000e\u0016!(\notherwise\u000f\n\n\u00062\u0016\u0013\u0016\u00139Z^ becomes the\n\n(6 *(a%\n\nneeded to bring \u0005\n\n\u001c\u001e\u0007\n\n^:\u0007\n\n\u0003\u0006\u00056\u0007\u0005&2\u0016\n\n(\u001d \n\n#\u0004\u0016!(\n\n^n\r\u0019b\u0015\u0007e\u001c\n\nb\u0015\u0007\n\n^:\u0007e\u001c\u0017\u000f\n\no2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016\n\n(/o2\u0003\n\n\t\b\u0007\n\n\u001c\u0017\r\u0019^\u0015\u0007\n\n\u0005\u0018\r\u0013\t+!\n\n\u0002'%\n\n\t\f\u000f\n\n#2#\n\n\u0003\n\u001c\u000e\u0016\n\n(17)\n\n5.2 Logistic models\n\n9\u0017%\n\n\u0001\u0004\u0016\n\n\u001d3v:\r\u0011\u001c\u0017\u001f\n\n5.3 Error correcting output codes\n\nA more interesting special case of the algebraic structure described in Sections 3 and 4\nis where the ensemble method is error correcting output coding ( ECOC) [4]. Here we set\n\n\u0001\n\t\n\t\n\u0001\nS\n\nS\n(\n\n[\n%\n \n\n5\n\n\nt\n\u0014\n(\n\b\n\t\n#\n\u000b\n\t\n\u0017\n\b\n\t\n(\nb\n\u0016\nb\n\u0017\nb\n\u0016\n\u0017\n\u0016\n\u0017\n\n\u0001\n(\n\u001c\n%\n\u0003\n\u0015\n-\n\u0013\nT\n\u0001\n\u0002\n\u0012\n)\n\u0004\n\n&\n\u0002\nt\n\u0014\n!\nt\n\u0014\n\t\n\t\n#\n\u0016\n(\n%\n\t\n#\n\u000b\n%\n\t\n#\n\u000b\n\t\n#\n\u000b\n\u0007\nv\n\u000b\n\u000b\n^\n\b\n[\n\u0003\n\u0006\n\n\u0002\n\f(18)\n\n) moves in the Hasse diagram\n\nput\n, which corresponds to one of the binary\nclassi\ufb01ers in ECOC for the appropriate column of the binary coding matrix. For example,\n,\n\n. On an input\n\n\u00062\u0016\n\n.\n\n\u0002Ot\n\nv\u0004\u001fG\u000f\n\n\t:9G%\n\nto \t\n\n>q(*&\u0010&\u0011&G(\n\n and1\n\n, the model computes probabilities of classi\ufb01cations\n\n, and take the parameter space to be\n[\u0001\n\n(,%\nAs before, o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016\nneeded to bring \u0005\nSince \u0005\nconsider a binary classi\ufb01er trained on the coding column \u0003\n\u001c\u0017\r\u0019v\u0004\r\u0010\u001c\u0017\r\u0013v:\r\u0013v:\r\u0013vG\u0016\nthe classi\ufb01er outputs 0 or 1, corresponding to the partial rankings \t\n^n\r\u0013j:\r0kn\r\b\u0007\n\t@(*\u001cG\r\u0013b\u0015\u0007\nSince \u0005\n\u0002@%\n9G%\n\n\u000b , l(\n\r\u000e9\n\t\u0002\u0007\n55(-\u001d\u0012&\nis the minimal number of up-down (\u0017\n, the base rankers output \t\b!G\u0003\n\t\u0002\u0007\n\n\u000b and \t\n\n, respectively.\n\no2\u0003\u0006\u0005\u0018\r\u0013\t\u0002\u0016\n\n\t\b\u0007\n\n\u00129\n\no2\u0003\n\n\u0007i\u001d\n\n?9\n\u0012\u0004\r\u0011\u001d\nif\u0012\f#\n\n\u0003\n\u001c\u000e\u0016\notherwise.\n\n=\n\t\n\n\u0004 . On in-\n\n\u0001\u0003\u0002\u0005\u0004\n\n^n\r7j\u0015\r\u0019k\u0004\r\b\u0007\u0015\u0007e\u001c\u0017\r\u0019b and\n\n(19)\n\n(20)\n\n\u001c , as can be seen\n\n(21)\n\nand \tX(N^n\r\u0013j:\r\u0019k\u0004\r\b\u00072\u0007i\u001cG\r\u0013b , then o2\u0003\u0006\u0005\u0018\r\u0013\t\u0002\u0016`(\n^n\r\u0013j:\r\u0019k\u0004\r\b\u00072\u0007i\u001cG\r\u0013b`\u000f\n\nj:\r0kn\r\b\u00072\u0007i\u001cG\r\u0013b\n\n^:\u0007\n\nFor example, if \u0005\n\nfrom the sequence of moves\n\n(*^:\u0007e\u001c\u0017\r\u0019b\u0004\r\u0013j:\r\u0019k\u0004\r\b\u0007\n^:\u0007e\u001c\u0017\r\u0019b\u0004\r\u0013j:\r\u0019k\u0004\r\b\u0007\nand \t'(\n\u001c\u001e\u0007\n^n\r7j\u0015\r\u0019k\u0004\r\b\u0007\u0015\u0007\n\nIf \u0005'(*\u001c8\u0007\n\n^\u0004\r\u0013b:\r7j:\r0kn\r\u000b\u0007\n\n\u001c\u001e\u0007\n\n^n\r\u0019b\u0004\r7j\u0015\r\u0019k\u0004\r\b\u0007\n\n^n\r7j\u0015\r\u0019kn\r\u000b\u0007\u0015\u0007e\u001c\u0017\r\u0019b , then o2\u0003\u0006\u0005\u0018\r\u0019\t\u0002\u0016\n\n^ , with the sequence of moves\n\n(22)\n\n\u001c\u0017\r\u0019^\u0004\r7j\u0015\r\u0019kn\r\u000b\u0007\u0015\u0007\n\n^\u0004\r7j:\r0kn\r\u000b\u0007\u0015\u0007e\u001c\u001e\u0007\n\n^\u0004\r7j:\r0kn\r\u000b\u0007\u0015\u0007e\u001c\u0017\r\u0013bM\u000f\n\narg\n\n\u000b`(\n\n\u0003\u0006\u0005\b\u0007\n\n(*&\u0011&\u0011&8(\n\nis strictly negative,\n\nis the Hamming distance between the appropriate row of the coding matrix and the con-\ncatenation of the bits returned from the binary classi\ufb01ers.\n\no2\u0003\u0006\u0005\u0018\r\u0019\t+!\u000e\u0016 . At test time, the model\n\f\u0001n\u0003\u0006\u00056\u0007$#M\u0016 . Now,\no2\u0003\u0006\u0005\u0018\r\u0019\t+!\u000e\u0016 .\no\u0015\u0003c\u0005\u0018\r\u0013\t+!\u001a\u0016nJ\b+\n\n! , the exponent of the model becomes1\n\nSince1\n\r\u000f\u000e\u000e,\nthus selects the label corresponding to the partial ranking \u0005\u0001\f\n#M\u0016 is a monotonically decreasing function in \u000b\nsince1\nEquivalence with the ECOC decision rule thus follows from the fact that \u000b\nThus, with the appropriate de\ufb01nitions of9\\\r\nstraint1\nfrom using a nonuniform carrier density\u0004\n\nand o , the conditional model on the ranking\n\u0014 results in a more general model that corresponds to ECOC with a\n\nposet is a probabilistic formulation of ECOC that yields the same classi\ufb01cation decisions.\nThis suggests ways in which ECOC might be naturally extended. First, relaxing the con-\n\nweighted Hamming distance, or index sensitive \u201cchannel,\u201d where the learned weights may\nadapt to the precision of the various base classi\ufb01ers. Another simple generalization results\n\ntrained classi\ufb01er for a given column outputs either \u001d\n\n. Allowing the output of the classi\ufb01er instead to belong to other\ndepending on the input\ngrades of\nresults in a model that corresponds to error correcting output codes with non-\nbinary codes. While this is somewhat antithetic to the original spirit of ECOC\u2014reducing\nmulticlass to binary\u2014the base classi\ufb01ers in ECOC are often multiclass classi\ufb01ers such as\ndecision trees in [4]. For such classi\ufb01ers, the task instead can be viewed as reducing mul-\nticlass to partial ranking. Moreover, there need not be an explicit coding matrix. Instead,\nthe input rankers may output different partial rankings for different inputs, which are then\ncombined according to model (6). In this way, a different coding matrix is built for each\nexample in a dynamic manner. Such a scheme may be attractive in bypassing the problem\nof designing the coding matrix.\n\nA further generalization is achieved by considering that for a given coding matrix, the\n\n=\n\t\n\u0011 or \u001d\n\n\u0003\u0006\u0005\u0012\u0016 .\n\n=\u0010\t\n\n\u0007i\u001d\n\n9\n\t\n#\n\b\n\u000b\n\u0014\n\u0007\n1\n\u000b\n(\n1\n1\n\u0014\n\u0016\n(\n%\n\t\n#\n\u000b\n\u0012\nT\n\u0002\n\u000b\n\u0006\n\u0002\n\b\n\u000b\n\u0006\n\u0006\n(\n\t\n\t\n#\n\u0002\n\b\n\u000b\n(\n%\n\t\n#\n\u000b\n\u0001\n[\n\u001f\n[\n=\n\u0011\n\u0001\n[\n\u001f\n[\n\u0011\n\u0016\n(\n\u0007\n\u001c\n\u000b\n\u0002\n\u001d\n^\n\u0017\n\u0016\n(\n\u0017\nb\n\u0016\nb\n\u0017\nb\n\u0016\n[\n(\n1\n\u000b\n!\n(\n<\n\u0001\n!\n\u0014\n!\n\u0018\n\u000b\n \n1\n>\n1\n\n\u0001\n[\n\u001f\n[\n=\n\u0011\n\u0001\n[\n\u001f\n[\n\u0001\n[\n\u001f\n[\n\u0011\n\u0007\n\u001d\n\u0001\n[\n\u001f\n[\n=\n\u0011\n\u0006\n\b\n\t\n\f6 Summary\n\nAn algebraic framework has been presented for classi\ufb01cation and ranking, leading to con-\nditional models on the ranking poset that are de\ufb01ned in terms of an invariant distance or\ndissimilarity function. Using the invariance properties of the distances, we derived a gen-\nerative interpretation of the probabilistic model, which may prove to be useful in model\n\nselection and validation. Through different choices of the components\u0004\u0012\u001e\r\u00059]\n\nand o , the\nfamily of models was shown to include as special cases the Mallows  model, and the\nclassi\ufb01er combination methods used by logistic models, boosting, cranking, and error cor-\nrecting output codes. In the case of ECOC, the poset framework shows how probabilities\nmay be assigned to partial rankings in a way that is consistent with the usual de\ufb01nitions of\nECOC, and suggests several natural extensions.\n\nAcknowledgments\n\nWe thank D. Critchlow, G. Hulten and J. Verducci for helpful input on the paper. This work\nwas supported in part by NSF grant CCR-0122581.\n\nReferences\n\n[1] M. Collins, R. E. Schapire, and Y. Singer. Logistic regression, AdaBoost and Breg-\n\nman distances. Machine Learning, 48, 2002.\n\n[2] D. E. Critchlow. Metric Methods for Analyzing Partially Ranked Data. Lecture Notes\n\nin Statistics, volume 34, Springer, 1985.\n\n[3] D. E. Critchlow and J. S. Verducci. Detecting a trend in paired rankings. Journal of\n\nthe Royal Statistical Society C, 41(1):17\u201329, 1992.\n\n[4] T. G. Dietterich and G. Bakiri. Solving multiclass learning problems via error-\n\ncorrecting codes. Journal of Arti\ufb01cial Intelligence Research, 2:263\u2013286, 1995.\n\n[5] P. D. Feigin. Modeling and analyzing paired ranking data.\n\nIn M. A. Fligner and\nJ. S. Verducci, editors, Probability Models and Statistical Analyses for Ranking Data.\nSpringer, 1992.\n\n[6] M. A. Fligner and J. S. Verducci. Posterior probabilities for a concensus ordering.\n\nPsychometrika, 55:53\u201363, 1990.\n\n[7] Y. Freund and R. E. Schapire. Experiments with a new boosting algorithm. In Inter-\n\nnational Conference on Machine Learning, 1996.\n\n[8] M. G. Kendall. A new measure of rank correlation. Biometrika, 30, 1938.\n[9] G. Lebanon and J. Lafferty. Boosting and maximum likelihood for exponential mod-\n\nels. In Advances in Neural Information Processing Systems, 15, 2001.\n\n[10] G. Lebanon and J. Lafferty. Cranking: Combining rankings using conditional prob-\nability models on permutations. In International Conference on Machine Learning,\n2002.\n\n[11] C. L. Mallows. Non-null ranking models. Biometrika, 44:114\u2013130, 1957.\n[12] R. P. Stanley. Enumerative Combinatorics, volume 1. Wadsworth & Brooks/Cole\n\nMathematics Series, 1986.\n\n \n\f", "award": [], "sourceid": 2146, "authors": [{"given_name": "Guy", "family_name": "Lebanon", "institution": null}, {"given_name": "John", "family_name": "Lafferty", "institution": null}]}