{"title": "Learning Networks of Heterogeneous Influence", "book": "Advances in Neural Information Processing Systems", "page_first": 2780, "page_last": 2788, "abstract": "Information, disease, and influence diffuse over networks of entities in both natural systems and human society. Analyzing these transmission networks plays an important role in understanding the diffusion processes and predicting events in the future. However, the underlying transmission networks are often hidden and incomplete, and we observe only the time stamps when cascades of events happen. In this paper, we attempt to address the challenging problem of uncovering the hidden network only from the cascades.  The structure discovery problem is complicated by the fact that the influence among different entities in a network are heterogeneous, which can not be described by a simple parametric model. Therefore, we propose a kernel-based method which can capture a diverse range of different types of influence without any prior assumption. In both synthetic and real cascade data, we show that our model can better recover the underlying diffusion network and drastically improve the estimation of the influence functions between networked entities.", "full_text": "Learning Networks of Heterogeneous In\ufb02uence\n\nNan Du\u2217 Le Song\u2217 Alex Smola\u2020 Ming Yuan\u2217\nGeorgia Institute of Technology\u2217, Google Research\u2020\ndunan@gatech.edu lsong@cc.gatech.edu\nalex@smola.org myuan@isye.gatech.edu\n\nAbstract\n\nInformation, disease, and in\ufb02uence diffuse over networks of entities in both nat-\nural systems and human society. Analyzing these transmission networks plays\nan important role in understanding the diffusion processes and predicting future\nevents. However, the underlying transmission networks are often hidden and in-\ncomplete, and we observe only the time stamps when cascades of events happen.\nIn this paper, we address the challenging problem of uncovering the hidden net-\nwork only from the cascades. The structure discovery problem is complicated by\nthe fact that the in\ufb02uence between networked entities is heterogeneous, which can\nnot be described by a simple parametric model. Therefore, we propose a kernel-\nbased method which can capture a diverse range of different types of in\ufb02uence\nwithout any prior assumption. In both synthetic and real cascade data, we show\nthat our model can better recover the underlying diffusion network and drastically\nimprove the estimation of the transmission functions among networked entities.\n\nIntroduction\n\n1\nNetworks have been powerful abstractions for modeling a variety of natural and arti\ufb01cial systems\nthat consist of a large collection of interacting entities. Due to the recent increasing availability\nof large-scale networks, network modeling and analysis have been extensively applied to study\nthe spreading and diffusion of information, ideas, and even virus in social and information net-\nworks (see e.g., [17, 5, 18, 1, 2]). However, the process of in\ufb02uence and diffusion often occurs in\na hidden network that might not be easily observed and identi\ufb01ed directly. For instance, when a\ndisease spreads among people, epidemiologists can know only when a person gets sick, but they\ncan hardly ever know where and from whom he (she) gets infected. Similarly, when consumers\nrush to buy some particular products, marketers can know when purchases occurred, but they cannot\ntrack in further where the recommendations originally came from [12]. In all such cases, we could\nobserve only the time stamp when a piece of information has been received by a particular entity,\nbut the exact path of diffusion is missing. Therefore, it is an interesting and challenging question\nwhether we can uncover the diffusion paths based just on the time stamps of the events.\nThere are many recent studies on estimating correlation or causal structures from multivariate time-\nseries data (see e.g., [2, 6, 13]). However, in these models, time is treated as discrete index and not\nmodeled as a random variable. In the diffusion network discovery problem, time is treated explic-\nitly as a continuous variable, and one is interested in capturing how the occurrence of event at one\nnode affects the time for its occurence at other nodes. This problem recently has been explored by\na number of studies in the literature. Speci\ufb01cally, Meyers and Leskovec inferred the diffusion net-\nwork by learning the infection probability between two nodes using a convex programming, called\nCONNIE [14]. Gomez-Rodriguez et al. inferred the network connectivity using a submodular opti-\nmization, called NETINF [4]. However, both CONNIE and NETINF assume that the transmission\nmodel for each pair of nodes is \ufb01xed with prede\ufb01ned transmission rate. Recently, Gomez-Rodriguez\net al. proposed an elegant method, called NETRATE [3], using continuous temporal dynamics\nmodel to allow variable diffusion rates across network edges. NETRATE makes fewer number of\nassumptions and achieves better performance in various aspects than the previous two approaches.\nHowever, the limitation of NETRATE is that it requires the in\ufb02uence model on each edge to have a\n\n1\n\n\f(a) Pair 1\n\n(b) Pair 2\n\n(c) Pair 3\n\nFigure 1: The histograms of the interval between the time when a post appeared in one site and the time when\na new post in another site links to it. Dotted and dash lines are density \ufb01tted by NETRATE. The solid lines are\ngiven by KernelCascade.\n\ufb01xed parametric form, such as exponential, power-law, or Rayleigh distribution, although the model\nparameters learned from cascades could be different.\nIn practice, the patterns of information diffusion (or a spreading disease) among entities can be quite\ncomplicated and different from each other, going far beyond what a single family of parametric\nmodels can capture. For example, in twitter, an active user can be online for more than 12 hours a\nday, and he may instantly respond to any interesting message. However, an inactive user may just\nlog in and respond once a day. As a result, the spreading pattern of the messages between the active\nuser and his friends can be quite different from that of the inactive user.\nAnother example is from the information diffusion in a blogsphere: the hyperlinks between posts can\nbe viewed as some kind of information \ufb02ow from one media site to another, and the time difference\nbetween two linked posts reveal the pattern of diffusion. In Figure 1, we examined three pairs of\nmedia sites from the MemeTracker dataset [3, 9], and plotted the histograms of the intervals between\nthe the moment when a post \ufb01rst appeared in one site and the moment when it was linked by a new\npost in another site. We can observe that information can have very different transmission patterns\nfor these pairs. Parametric models \ufb01tted by NETRATE may capture the simple pattern in Figure 1(a),\nbut they might miss the multimodal patterns in Figure 1(b) and Figure 1(c). In contrast, our method,\ncalled KernelCascade, is able to \ufb01t both data accurately and thus can handle the heterogeneity.\nIn the reminder of this paper, we present the details of our approach KernelCascade. Our key idea\nis to model the continuous information diffusion process using survival analysis by kernelizing the\nhazard function. We obtain a convex optimization problem with grouped lasso type of regularization\nand develop a fast block-coordinate descent algorithm for solving the problem. The sparsity patterns\nof the coef\ufb01cients provide us the structure of the diffusion network. In both synthetic and real world\ndata, our method can better recover the underlying diffusion networks and drastically improve the\nestimation of the transmission functions among networked entities.\n2 Preliminary\nIn this section, we will present some basic concepts from survival analysis [7, 8], which are essential\nfor our later modeling. Given a nonnegative random variable T corresponding to the time when an\n0 f (x)dx\nbe its cumulative distribution function. The probability that an event does not happen up to time t\nis thus given by the survival function S(t) = P r(T \u2265 t) = 1 \u2212 F (t). The survival function is a\ncontinuous and monotonically decreasing function with S(0) = 1 and S(\u221e) = limt\u2192\u221e S(t) = 0.\nGiven f (t) and S(t), we can de\ufb01ne the instantaneous risk (or rate) that an event has not happened\nyet up to time t but happens at time t by the hazard function\n\nevent happens, let f (t) be the probability density function of T and F (t) = P r(T \u2264 t) =(cid:82) t\n\nP r(t \u2264 T \u2264 t + \u2206t|T \u2265 t)\n\nh(t) = lim\n\u2206t\u21920\n\n\u2206t\n\n=\n\nf (t)\nS(t)\n\n.\n\n(1)\n\nWith this de\ufb01nition, h(t)\u2206t will be the approximate probability that an event happens in [t, t + \u2206t)\ngiven that the event has not happened yet up to t. Furthermore, the hazard function h(t) is also\nrelated to the survival function S(t) via the differential equation h(t) = \u2212 d\ndt log S(t), where we\nhave used f (t) = \u2212S(cid:48)(t). Solving the differential equation with boundary condition S(0) = 1, we\ncan recover the survival function S(t) and the density function f (t) based on the hazard function\nh(t), i.e.,\n\nS(t) = exp\n\nh(x) dx\n\nand\n\nf (t) = h(t) exp\n\nh(x) dx\n\n.\n\n(2)\n\n(cid:90) t\n\n(cid:18)\n\n\u2212\n\n(cid:19)\n\n(cid:90) t\n\n(cid:18)\n\n\u2212\n\n(cid:19)\n\n0\n\n0\n\n2\n\n0102030405000.10.20.3t(hours)pdf  histogramexprayleighKernelCascade0204000.020.040.060.080.1t(hours)pdf  histogramexprayleighKernelCascade05010000.020.040.060.080.1t(hours)pdf  histogramexprayleighKernelCascade\f(a) Hidden network\n\n(b) Node e gets infected at time t4\n\n(c) Node e survives\n\nFigure 2: Cascades over a hidden network. Solid lines in panel(a) represent connections in a hidden network.\nIn panel (b) and (c), \ufb01lled circles indicate infected nodes while empty circles represent uninfected ones. Node\na, b, c and d are the parents of node e which got infected at t0 < t1 < t2 < t3 respectively and tended to infect\nnode e. In panel (b), node e survives given node a, b and c shown in green dash lines. However, it was infected\nby node d. In panel (c), node e survives even though all its parents got infected.\n3 Modeling Cascades using Survival Analysis\nWe use survival analysis to model information diffusion for networked entities. We will largely\nfollow the presentation of Gomez-Rodriguez et al. [3], but add clari\ufb01cation when necessary. We\nassume that there is a \ufb01xed population of N nodes connected in a directed network G = (V,E).\nNeighboring nodes are allowed to directly in\ufb02uence each other. Nodes along a directed path may\nin\ufb02uence each other only through a diffusion process. Because the true underlying network is un-\nknown, our observations are only the time stamps when events occur to each node in the network.\nThe time stamps are then organized as cascades, each of which corresponds to a particular event.\nFor instance, a piece of news posted on CNN website about \u201cFacebook went public\u201d can be treated\nas an event. It can spread across the blogsphere and trigger a sequence of posts from other sites\nreferring to it. Each site will have a time stamp when this particular piece of news is being discussed\nand cited. The goal of the model is to capture the interplay between the hidden diffusion network\nand the cascades of observed event time stamps.\nMore formally, a directed edge, j \u2192 i, is associated with an transmission function fji(ti|tj), which\nis the conditional likelihood of an event happening to node i at time ti given that the same event has\nalready happened to node j at time tj. The transmission function attempts to capture the temporal\ndependency between the two successive events for node i and j. In addition, we focus on shift-\ninvariant transmission functions whose value only depends on the time difference, i.e., fji(ti|tj) =\nfji(ti \u2212 tj) = fji(\u2206ji) where \u2206ji := ti \u2212 tj. Given the likelihood function, we can compute the\ncorresponding survival function Sji(\u2206ji) and hazard function hji(\u2206ji). When there is no directed\nedge j \u2192 i, the transmission function and hazard function are both identically zeros, i.e., fji(\u2206ji) =\n0 and hji(\u2206ji) = 0, but the survival function is identically one, i.e., Sji(\u2206ji) = 1. Therefore, the\nstructure of the diffusion network is re\ufb02ected in the non-zero patterns of a collection of transmission\nfunctions (or hazard functions).\nN )(cid:62) with i-th dimension recording the time\nA cascade is an N-dimensional vector tc := (tc\n1, . . . , tc\ni \u2208 [0, T c] \u222a {\u221e}, and the symbol \u221e labels\nstamp when event c occurs to node i. Furthermore, tc\nnodes that have not been in\ufb02uenced during observation window [0, T c] \u2014 it does not imply that\nnodes are never in\ufb02uenced. The \u2018clock\u2019 is set to 0 at the start of each cascade. A dataset can\n\ncontain a collection, C, of cascades(cid:8)t1, . . . , t|C|(cid:9). The time stamps assigned to nodes by a cascade\n\ninduce a directed acyclic graph (DAG) by de\ufb01ning node j as the parent of i if tj < ti. Thus, it\nis meaningful to refer to parents and children within a cascade [3], which is different from the\nparent-child structural relation on the true underlying diffusion network. Since the true network is\ninferred from many cascades (each of which imposes its own DAG structure), the inferred network\nis typically not a DAG.\nThe likelihood (cid:96)(tc) of a cascade induced by event c is then simply a product of all individual\nlikelihood (cid:96)i(tc) that event c occurs to each node i. Depending on whether event c actually occurs\nto node i in the data, we can compute this individual likelihood as:\nEvent c did occur at node i. We assume that once an event occurs at node i under the in\ufb02uence of\na particular parent j in a cascade, the same event will not happen again. In Figure 2(b), node e is\nsusceptible given its parent a, b, c and d. However, only node d is the \ufb01rst parent who infects node e.\nBecause each parent could be equally likely to \ufb01rst in\ufb02uence node i, the likelihood is just a simple\nsum over the likelihoods of the mutually disjoint events that node i has survived from the in\ufb02uence\nof all the other parents except the \ufb01rst parent j, i.e.,\nSki(\u2206c\n\n(cid:88)\n\n(cid:88)\n\n(cid:89)\n\n(cid:89)\n\nSki(\u2206c\n\nki).\n\nhji(\u2206c\n\nji)\n\nki) =\n\n(cid:96)+\ni (tc) =\n\nfji(\u2206c\n\nji)\n\nj:tc\n\nj <tc\ni\n\nk:k(cid:54)=j,tc\n\nk<tc\ni\n\n(3)\n\nj:tc\n\nj <tc\ni\n\nk:tc\n\nk<tc\ni\n\n3\n\na b c d e t0 t1 t3 t2 t4 a b c d e t0 t1 t3 t2 a b c d e \fEvent c did not occur at node i. In other words, node i survives from the in\ufb02uence of all par-\nents (see Figure 2(c) for illustration). The likelihood is a product of survival functions, i.e.,\n\nCombining the above two scenarios together, we can obtain the overall likelihood of a cascade tc by\nmultiplying together all individual likelihoods, i.e.,\n(cid:96)\u2212\ni (tc)\n\n(cid:96)+\ni (tc)\n\n(cid:96)(tc) =\n\n(5)\n\n.\n\n(cid:89)\n\ntj\u2264T c\n\n(cid:96)\u2212\ni (tc) =\n\nSji(T c \u2212 tj).\n\n(cid:89)\n\n(cid:124)\n\ntc\ni >T c\n\n(cid:123)(cid:122)\n\nuninfected nodes\n\n\u00d7 (cid:89)\n(cid:124)\n\n(cid:125)\n\n(cid:123)(cid:122)\n\ni\u2264T c\ntc\ninfected nodes\n\n(cid:125)\n\n(4)\n\n\uf8f6\uf8f7\uf8f8 (6)\n\nthe likelihood of all cascades is a product of the these individual cascade likeli-\nc=1,...,|C| (cid:96)(tc). In the end, we take the negative log of this likeli-\n\nhood function and regroup all terms associated with edges pointing to node i together to derive\n\nTherefore,\n\nhoods, i.e. (cid:96)({t1, . . . , t|C|}) =(cid:81)\n\uf8eb\uf8ec\uf8ed(cid:88)\nL({t1, . . . , t|C|}) = \u2212(cid:88)\n\n(cid:88)\n\n(cid:88)\n\n(cid:88)\n\ni\n\nj\n\n{c|tc\n\ni}\nj <tc\n\nlog S(\u2206c\n\nji) +\n\nlog\n\n{c|tc\n\ni\n\n(cid:54)T c}\n\ni}\n{tc\nj <tc\n\nh(\u2206c\n\nji)\n\nThere are two interesting implications from this negative log likelihood function. First, the function\ncan be expressed using only the hazard and the survival function. Second, the function is decom-\nposed into additive contribution from each node i. We can therefore estimate the hazard and survival\nfunction for each node separately. Previously, Gomez-Rodriguez et al. [3] used parametric hazard\nand survival functions, and they estimated the model parameters using the convex programming. In\ncontrast, we will instead formulate an algorithm using kernels and grouped parameter regularization,\nwhich allows us to estimate complicated hazard and survival functions without over\ufb01tting.\n4 KernelCascade for Learning Diffusion Networks\nThis section presents our kernel method for uncovering diffusion networks from cascades. Our key\nidea is to kernelize the hazard function used in the negative log-likelihood in (6), and then estimate\nthe parameters using grouped lasso type of optimization.\n4.1 Kernelizing survival analysis\nKernel methods are powerful tools for generalizing classical linear learning approaches to analyze\nnonlinear relations. A kernel function, k : X \u00d7 X \u2192 R, is a real-valued positive de\ufb01nite symmetric\nfunction iff. for any set of points {\u03c41, \u03c42, . . . , \u03c4m} \u2208 X the kernel matrix K with entris Kls :=\nk(\u03c4l, \u03c4s) is positive de\ufb01nite. We want to model heterogeneous transmission functions, fji(\u2206ji),\nfrom j to i. Rather than directly kernelizing the transmission function, we kernelize the hazard\nfunction instead, by assuming that it is a linear combination of m kernel functions, i.e.,\n\n\u03b1l\n\nhji(\u2206ji) =\n\nji, . . . , \u03b1m\n\njik(\u03c4l, \u2206ji),\n\n(7)\nwhere we \ufb01x one argument of each kernel function, k(\u03c4l,\u00b7), to a point \u03c4l in a uniform grid of m\nlocations in the range of (0, maxc T c]. To achieve fully nonparametric modeling of the hazard\nfunction, we can let m grow as we see more cascades. Alternatively, we can also place a non-\nlinear basis function on each time point in the observed cascades. For ef\ufb01ciency consideration,\nwe will use a \ufb01xed uniform grid in our later experiments. Since the hazard function is always\npositive, we use positive kernel functions and require the weights to be positive, i.e., k(\u00b7,\u00b7) \u2265 0\nji \u2265 0 to capture such constraint. For simplicity of notation, we will de\ufb01ne vectors \u03b1ji :=\nand \u03b1l\nji )(cid:62), and k(\u2206ji) := (k(\u03c41, \u2206ji), . . . , k(\u03c4m, \u2206ji))(cid:62). Hence, the hazard function can be\n(\u03b11\nwritten as hji(\u2206ji) = \u03b1(cid:62)\nIn addition, the survival function and likelihood function can also be kernelized based on their\nk(\u03c4l, x)dx\n\nrespective relation with the hazard function in (2). More speci\ufb01cally, let gl(\u2206ji) :=(cid:82) \u2206ji\nSji(\u2206ji) = exp(cid:0)\u2212\u03b1(cid:62)\njik(\u2206ji)(cid:1) exp(cid:0)\u2212\u03b1(cid:62)\n\nand the corresponding vector g(\u2206ji) := (g1(\u2206ji), . . . , gm(\u2206ji))(cid:62). We then can derive\n\n(8)\nIn the formulation, we need to perform integration over the kernel function to compute gl(\u2206ji). This\ncan be done ef\ufb01ciently for many kernels, such as the Gaussian RBF kernel, the Laplacian kernel,\nthe Quartic kernel, and the Triweight kernel. In later experiments, we mainly focus on the Gaussian\n\nfji(\u2206ji) =(cid:0)\u03b1(cid:62)\n\njig(\u2206ji)(cid:1) .\n\njig(\u2206ji)(cid:1)\n\njik(\u2206ji).\n\nand\n\n0\n\nm(cid:88)\n\nl=1\n\n4\n\n\f(cid:18)\n\n(cid:19)(cid:19)\n\n(cid:18) \u03c4l\u221a\n\n(cid:18) \u03c4l \u2212 \u2206ji\u221a\n\n(cid:19)\n\n(cid:90) \u2206ji\n(cid:82) \u221e\nt e\u2212x2\n\n0\n\n\u03c0\n\n2\u03c3\n\n\u221a\n\nerfc\n\n2\u03c0\u03c3\n2\n\n\u2212 erfc\n\nk(\u03c4l, x) dx =\n\nRBF kernel, k(\u03c4l, \u03c4s) = exp(\u2212(cid:107)\u03c4l \u2212 \u03c4s(cid:107)2/(2\u03c32)), and derive a closed form solution for gl(\u2206ji) as\n(9)\n\ngl(\u2206ji) =\nwhere erfc(t) := 2\u221a\ndx is the error function. Yet, our method is not limited to the particu-\nlar RBF kernel. If there is no closed form solution for the one-dimensional integration, we can use a\nlarge number of available numerical integration methods for this purpose [15]. We note that given a\ndataset, both the vector k(\u2206ji) and g(\u2206ji) need to be computed only once as a preprocessing, and\nthen can be reused in the algorithm.\n4.2 Estimating sparse diffusion networks\nNext we plug in the kernelized hazard function and survival function into the likelihood of cascades\nin (6). Since the negative log likelihood is separable for each node i, we can optimize the set of\nvariables {\u03b1ji}N\nj=1 separately. As a result, the negative log likelihood for the data associated with\nnode i can be estimated as\n\n2\u03c3\n\n,\n\n(cid:0){\u03b1ji}N\n\nj=1\n\n(cid:88)\n\n(cid:1) =\n\nLi\n\n(cid:88)\n\nji) \u2212 (cid:88)\n\n\u03b1(cid:62)\njig(\u2206c\n\n\u03b1(cid:62)\njik(\u2206c\n\nji).\n\n(10)\n\n(cid:88)\n\nj\n\n{c|tc\n\ni}\nj <tc\n\nlog\n\n{c|tc\n\ni\n\n(cid:54)T c}\n\ni}\n{tc\nj <tc\n\nA desirable feature of this function is that it is convex in its arguments, {\u03b1ji}N\nto bring various convex optimization tools to solve the problem ef\ufb01ciently.\nIn addition, we want to induce a sparse network structure from the data and avoid over\ufb01tting.\nBasically, if the coef\ufb01cients \u03b1ji = 0, then there is no edge (or direct in\ufb02uence) from node j\nto i. For this purpose, we will impose grouped lasso type of regularization on the coef\ufb01cients\nj (cid:107)\u03b1ji(cid:107))2 [16, 19]. Grouped lasso type of regularization has the tendency to select a\nsmall number of groups of non-zero coef\ufb01cients but push other groups of coef\ufb01cients to be zero.\nOverall, the optimization problem trades off between the data likelihood term and the group sparsity\nof the coef\ufb01cients\n\n\u03b1ji, i.e., ((cid:80)\n\nj=1, which allows us\n\n(cid:0){\u03b1ji}N\n\nj=1\n\n(cid:16)(cid:88)\n\n(cid:1) + \u03bb\n\n(cid:107)\u03b1ji(cid:107)(cid:17)2\n\n,\n\nLi\n\ns.t. \u03b1ji \u2265 0, \u2200j,\n\n(11)\n\nmin\n{\u03b1ji}N\n\nj=1\n\nj\n\nwhere \u03bb is the regularization parameter. After we obtain a sparse solution from the above opti-\nmization, we obtain partial network structures, each of which centers around a particular node i.\nWe can then join all the partial structures together and obtain the overall diffusion network. The\ncorresponding hazard function along each edge can also be obtained from (8).\n4.3 Optimization\nWe note that (11) is a nonsmooth optimization problem because of the regularizer. There are\nmany ways to solve the optimization problem, and we will illustrate this using a simple algo-\nrithm originating from multiple kernel learning [16, 19].\nIn this approach, an additional set of\nvariables are introduced to turn the optimization problem into a smooth optimization problem.\nj \u03b3j = 1. Then using Cauchy-Schwartz inequality, we have\nj (cid:107)\u03b1ji(cid:107)2 /\u03b3j, where\n\nMore speci\ufb01cally, let \u03b3i \u2265 0 and(cid:80)\n((cid:80)\nj (cid:107)\u03b1ji(cid:107))2 = ((cid:80)\n\nj \u03b3j) = (cid:80)\n\n)2 \u2264 ((cid:80)\n\nj((cid:107)\u03b1ji(cid:107) /\u03b31/2\n\n)\u03b31/2\n\nj\n\nj\n\nthe equality holds when\n\nj (cid:107)\u03b1ji(cid:107)2 /\u03b3j)((cid:80)\n(cid:88)\n\n(cid:107)\u03b1ji(cid:107) .\n\n(12)\nWith these additional variables, \u03b3j, we can solve an alternative smooth optimization problem, which\nis jointly convex in both \u03b1ji and \u03b3j\n\nj\n\n\u03b3j = (cid:107)\u03b1ji(cid:107) /\n\n(cid:0){\u03b1ji}N\n\nj=1\n\n(cid:1) + \u03bb\n\n(cid:88)\n\nLi\n\n(cid:107)\u03b1ji(cid:107)2\n\n\u03b3j\n\nmin\n\n{\u03b1ji,\u03b3j}N\n\n, s.t. \u03b1ji \u2265 0, \u03b3j \u2265 0,\n\n\u03b3j = 1, \u2200j.\n\n(13)\n\nj=1\n\nj\n\nj\nThere are many ways to solve the convex optimization problem in (13).\nIn this paper, we used\na block coordinate descent approach alternating between the optimization of \u03b1ji and \u03b3j. More\nspeci\ufb01cally, when we \ufb01xed \u03b1ji, we can obtain the best \u03b3j using the closed form formula in (12);\nwhen we \ufb01xed \u03b3j , we can optimize over \u03b1ji using, e.g., a projected gradient method. The overall\nalgorithm pseudocodes are given in Algorithm 1. Moreover, we can speed up the optimization in\nthree ways. First, because the optimization is independent for each node i, the overall process can\nbe easily parallelized into N separate sub-problems. Second, we can prune the possible nodes that\nwere never infected before node i in any cascade where i was infected. Third, if we further assume\nthat all the edges from the same node belong to the same type of models, especially when the sample\n\n(cid:88)\n\n5\n\n\fsize is small, the N edges could share a common set of m parameters, and thus we can only estimate\nN \u00d7 m parameters in total.\nAlgorithm 1: KernelCascade\nInitialize the diffusion network G to be empty;\nfor i = 1 to N do\n\nIntialize {\u03b1ji}N\nrepeat\n\nj=1 and {\u03b3j}N\n\nj=1;\n\nj=1 using projected gradient method with {\u03b3j}N\nj=1 using formula (12) with {\u03b1ji}N\n\nUpdate {\u03b1ji}N\nUpdate {\u03b3j}N\nuntil convergence;\nExtract the sparse neighborhood N (i) of node i from nonzero \u03b1ji;\nJoin N (i) to the diffusion network G\n\nj=1 from last update;\n\nj=1 from last update;\n\nexp\n\nbi\n\nbi\n\n(cid:17)2\n\n(cid:17)2(cid:19)\n\n(cid:16) t\u2212ai\n\n(cid:18)\n\u2212(cid:16) t\u2212ai\n\n5 Experimental Results\nWe will evaluate KernelCascade on both realistic synthetic networks and real world networks. We\ncompare it to NETINF [4] and NETRATE [3], and we show that KernelCascade can perform signif-\nicantly better in terms of both recovering the network structures and the transmission functions.\n5.1 Synthetic Networks\nNetwork generation. We \ufb01rst generate synthetic networks that mimic the structural properties of\nreal networks. These synthetic networks can then be used for simulation of information diffusion.\nSince the latent networks for generating cascades are known in advance, we can perform detailed\ncomparisons between various methods. We use Kronecker generator [10] to examine two types of\nnetworks with directed edges: (i) the core-periphery structure [11], which mimics the information\ndiffusion process in real world networks, and (ii) the Erd\u02ddos-R\u00b4enyi random networks.\nIn\ufb02uence function. For each edge j \u2192 i in a network G, we will assign it a mixture of two Rayleigh\ndistributions: fji(t|\u03b8, a1, b1, a2, b2) = \u03b8R1(t|a1, b1) + (1 \u2212 \u03b8)R2(t|a2, b2) where Ri(t|ai, bi) =\n, t (cid:62) ai, and \u03b8 \u2208 (0, 1) is a mixing proportion. We examine\n2\nt\u2212ai\nthree different parameter settings for the transmission function: (1) all edges in network G have\nthe same transmission function p(t) = f (t|0.5, 10, 1, 20, 1); (2) all edges in network G have the\nsame transmission function q(t) = f (t|0.5, 0, 1, 20, 1); and (3) all edges in network G are uniformly\nrandomly assigned to either p(t) or q(t).\nCascade generation. Given a network G and the collection of transmission functions fji for each\nedge, we generate a cascade from G by randomly choosing a node of G as the root of the cascade.\nThe root node j is then assigned to time stamp tj = 0. For each neighbor node i pointed by j,\nits event time ti is sampled from fji(t). The diffusion process will continue by further infecting\nthe neighbors pointed by node i in a breadth-\ufb01rst fashion until either the overall time exceed the\nprede\ufb01ned observation time window T c or there is no new node being infected. If a node is infected\nmore than once by multiple parents, only the \ufb01rst infection time stamp will be recorded.\nExperiment setting and evaluation metric. We consider a combination of two network topolo-\ngies (i)-(ii) with three different transmission function settings (1)-(3), which results in six different\nexperimental settings. For each setting, we randomly instantiate the network topologies and trans-\nmission functions for 10 times and then vary the number of cascades from 50, 100, 200, 400, 800 to\n1000. For KernelCascade, we use a Gaussian RBF kernel. The kernel bandwidth \u03c3 is chosen using\nmedian pairwise distance between grid time points. The regularization parameter is chosen using\ntwo fold cross-validation. NETINF requires the desired number of edges as input, and we give it an\nadvantage and supply the true number of edges to it. For NETRATE, we experimented with both\nexponential and Rayleigh transmission function.\nWe compare different methods in terms of (1) F 1 score for the network recovery. F 1 :=\n2\u00b7precision\u00b7recall\nprecision+recall , where precision is the fraction of edges in the inferred network that also present in\nthe true network and recall is the fraction of edges in the true network that also present in the inferred\nnetwork; (2) KL divergence between the estimated transmission function and the true transmission\nfunction, averaged over all edges in a network; (3) the shape of the \ufb01tted transmission function\ncompared to the true transmission function.\n\n6\n\n\f(a) Core-Periphery, p(t)\n\n(b) Core-Periphery, q(t)\n\n(c) Core-Periphery, mix p(t), q(t)\n\n(d) Random, p(t)\n\n(e) Random, q(t)\n\nFigure 3: F1 Scores for network recovery.\n\n(f) Random, mix p(t), q(t)\n\n(a) Core-Periphery, p(t)\n\n(b) Core-Periphery, q(t)\n\n(c) Core-Periphery, mix p(t), q(t)\n\n(d)Random, p(t)\nFigure 4: KL Divergence between the estimated and the true transmission function.\n\n(e) Random, q(t)\n\n(f) Random, mix p(t), q(t)\n\nF1 score for network recovery. From Figure 3, we can see that in all cases, KernelCascade per-\nforms consistently and signi\ufb01cantly better than NETINF and NETRATE. Furthermore, its perfor-\nmance also steadily increases as we increase the number of cascades, and \ufb01nally KernelCascade re-\ncovers the entire network with around 1000 cascades. In contrast, the competitor methods seldom\nfully recover the entire network given the same number of cascades. We also note that the perfor-\nmance of NETRATE is very sensitive to the choice of the transmission function (exponential vs.\nRayleigh). For instance, depending on the actual data generating process, the performance of NE-\nTRATE with Rayleigh model can vary from the second best to the worst.\nKL divergence for transmission function. Besides better network recovery, KernelCascade also\nestimates the transmission function better. In all cases we experimented, KernelCascade leads to\ndrastic improvement in recovering the transmission function (Figure 4). We also observe that as we\nincrease the number of cascades, KernelCascade adapts better to the actual transmission function. In\ncontrast, the performance of NETRATE with exponential model does not improve with increasing\nnumber of cascades, since the parametric model assumption is incorrect. We note that NETINF does\nnot recover the transmission function, and hence there is no corresponding curve in the plot.\nVisualization of the transmission function. We also visualize the estimated transmission function\nfor an edge from different methods in Figure 5. We can see that KernelCascade captures the essential\n\n7\n\n0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100000.20.40.60.81num of cascadesaverage F1  netinfnetrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade0500100002468num of cascadeslog(KL Distance)  netrate(rayleigh)netrate(exp)KernelCascade\f(a) An edge with transmission function p(t)\n\n(b) An edge with transmission function q(t)\n\nFigure 5: Estimated transmission function of a single edge based on 1000 cascades against the true\ntransmission function (blue curve).\n\n(a) KernelCascade\n\n(b) NETINF\n\n(c) NETRATE\n\nFigure 6: Estimated network of top 32 sites. Edges in grey are correctly uncovered, while edges\nhighlighted in red are either missed or estimated falsely.\nfeatures of the true transmission function, i.e., bi-modal behavior, while the competitor methods miss\nout the important statistical feature completely.\n5.2 Real world dataset\nFinally, we use the MemeTracker dataset [3] to compare NETINF, NETRATE and KernelCascade.\nIn this dataset, the hyperlinks between articles and posts can be used to represent the \ufb02ow of in-\nformation from one site to another site. When a site publishes a new post, it will put hyperlinks to\nrelated posts in some other sites published earlier as its sources. Later as it also becomes \u201colder\u201d, it\nwill be cited by other newer posts as well. As a consequence, all the time-stamped hyperlinks form\na cascade for particular piece of information (or event) \ufb02owing among different sites. The networks\nformed by these hyperlinks are used to be the ground truth. We have extracted a network consisting\nof top 500 sites with 6,466 edges and 11,530 cascades from 7,181,406 posts in a month, and we\nwant to recover the the underlying networks. From Table 1, we can see that KernelCascade achieves\na much better F 1 score for network recovery compared to other methods. Finally, we visualize the\nestimated sub-network structure for the top 32 sites in Figure 6. By comparison, KernelCascade has\na relatively better performance with fewer misses and false predictions.\n\nTable 1: Network recovery results from MemeTracker dataset.\npredicted edges\n\nprecision\n\nmethods\nNETINF\n\nNETRATE(exp)\nKernelCascade\n\n0.62\n0.93\n0.79\n\nrecall\n0.62\n0.23\n0.66\n\nF1\n0.62\n0.37\n0.72\n\n6466\n1600\n5368\n\n6 Conclusion\nIn this paper, we developed a \ufb02exible kernel method, called KernelCascade, to model the latent dif-\nfusion processes and to infer the hidden network with heterogeneous in\ufb02uence between each pair\nof nodes. In contrast to previous state-of-the-art, such as NETRATE, NETINF and CONNIE, Ker-\nnelCascade makes no restricted assumption on the speci\ufb01c form of the transmission function over\nnetwork edges. Instead, it can infer it automatically from the data, which allows each pair of nodes to\nhave a different type of transmission model and better captures the heterogeneous in\ufb02uence among\nentities. We obtain an ef\ufb01cient algorithm and demonstrate experimentally that KernelCascade can\nsigni\ufb01cantly outperforms previous state-of-the-art in both synthetic and real data. In future, we will\nexplore the combination of kernel methods, sparsity inducing norms and other point processes to\naddress a diverse range of social network problems.\nAcknowledgement: L.S. is supported by NSF IIS-1218749 and startup funds from Gatech.\n\n8\n\n010203000.10.20.30.40.5tpdf  KernelCascadeexprayleighdata010203000.10.20.30.40.5tpdf  KernelCascadeexprayleighdata\fReferences\n[1] M. De Choudhury, W. A. Mason, J. M. Hofman, and D. J. Watts.\n\nnetworks from interpersonal communication. In WWW, pages 301\u2013310, 2010.\n\nInferring relevant social\n\n[2] N. Eagle, A. S. Pentland, and D. Lazer. From the cover:\n\nInferring friendship network\nstructure by using mobile phone data. Proceedings of the National Academy of Sciences,\n106(36):15274\u201315278, Sept. 2009.\n\n[3] M. Gomez-Rodriguez, D. Balduzzi, and B. Sch\u00a8olkopf. Uncovering the temporal dynamics of\n\ndiffusion networks. In ICML, pages 561\u2013568, 2011.\n\n[4] M. Gomez-Rodriguez, J. Leskovec, and A. Krause. Inferring networks of diffusion and in\ufb02u-\n\nence. In KDD, pages 1019\u20131028, 2010.\n\n[5] D. Kempe, J. M. Kleinberg, and \u00b4E. Tardos. Maximizing the spread of in\ufb02uence through a\n\nsocial network. In KDD, pages 137\u2013146, 2003.\n\n[6] M. Kolar, L. Song, A. Ahmed, and E. Xing. Estimating time-varying networks. The Annals of\n\nApplied Statistics, 4(1):94\u2013123, 2010.\n\n[7] J. F. Lawless. Statistical Models and Methods for Lifetime Data. Wiley-Interscience, 2002.\n[8] E. T. Lee and J. Wang. Statistical Methods for Survival Data Analysis. Wiley-Interscience,\n\nApr. 2003.\n\n[9] J. Leskovec, L. Backstrom, and J. M. Kleinberg. Meme-tracking and the dynamics of the news\n\ncycle. In KDD, pages 497\u2013506, 2009.\n\n[10] J. Leskovec, D. Chakrabarti, J. M. Kleinberg, C. Faloutsos, and Z. Ghahramani. Kronecker\ngraphs: An approach to modeling networks. Journal of Machine Learning Research, 11:985\u2013\n1042, 2010.\n\n[11] J. Leskovec, K. J. Lang, and M. W. Mahoney. Empirical comparison of algorithms for network\n\ncommunity detection. In WWW, pages 631\u2013640, 2010.\n\n[12] J. Leskovec, A. Singh, and J. M. Kleinberg. Patterns of in\ufb02uence in a recommendation network.\n\nIn PAKDD, pages 380\u2013389, 2006.\n\n[13] A. C. Lozano and V. Sindhwani. Block variable selection in multivariate regression and high-\n\ndimensional causal inference. In NIPS, pages 1486\u20131494, 2010.\n\n[14] S. A. Myers and J. Leskovec. On the convexity of latent social network inference. In NIPS,\n\npages 1741\u20131749, 2010.\n\n[15] W. Press, S. Teukolsky, W. Vetterling, and B. Flannery. Numerical recipes in C: the art of\n\nscienti\ufb01c computing. Cambridge, 1992.\n\n[16] A. Rakotomamonjy, F. Bach, S. Canu, Y. Grandvalet, et al. Simplemkl. Journal of Machine\n\nLearning Research, 9:2491\u20132521, 2008.\n\n[17] D. Watts and S. Strogatz.\n\n393(6684):440\u2013442, June 1998.\n\nCollective dynamics of small-world networks.\n\nNature,\n\n[18] D. J. Watts and P. S. Dodds. In\ufb02uentials, networks, and public opinion formation. Journal of\n\nConsumer Research, 34(4):441\u2013458, 2007.\n\n[19] Z. Xu, R. Jin, H. Yang, I. King, and M. R. Lyu. Simple and ef\ufb01cient multiple kernel learning\n\nby group lasso. In ICML, pages 1175\u20131182, 2010.\n\n9\n\n\f", "award": [], "sourceid": 1277, "authors": [{"given_name": "Nan", "family_name": "Du", "institution": null}, {"given_name": "Le", "family_name": "Song", "institution": null}, {"given_name": "Ming", "family_name": "Yuan", "institution": null}, {"given_name": "Alex", "family_name": "Smola", "institution": null}]}