{"title": "Efficient Algorithms for Smooth Minimax Optimization", "book": "Advances in Neural Information Processing Systems", "page_first": 12680, "page_last": 12691, "abstract": "This paper studies first order methods for solving smooth minimax optimization problems $\\min_x \\max_y g(x,y)$ where $g(\\cdot,\\cdot)$ is smooth and $g(x,\\cdot)$ is concave for each $x$. In terms of $g(\\cdot,y)$, we consider two settings -- strongly convex and nonconvex -- and improve upon the best known rates in both. For strongly-convex $g(\\cdot, y),\\ \\forall y$, we propose a new direct optimal algorithm combining Mirror-Prox and Nesterov's AGD, and show that it can find global optimum in $\\widetilde{O}\\left(1/k^2 \\right)$ iterations, improving over current state-of-the-art rate of $O(1/k)$. We use this result along with an inexact proximal point method to provide $\\widetilde{O}\\left(1/k^{1/3} \\right)$ rate for finding stationary points in the nonconvex setting where $g(\\cdot, y)$ can be nonconvex. This improves over current best-known rate of $O(1/k^{1/5})$. Finally, we instantiate our result for finite nonconvex minimax problems, i.e., $\\min_x \\max_{1\\leq i\\leq m} f_i(x)$, with nonconvex $f_i(\\cdot)$, to obtain convergence rate of $O(m^{1/3}\\sqrt{\\log m}/k^{1/3})$.", "full_text": "Ef\ufb01cientAlgorithmsforSmoothMinimaxOptimizationKiranKoshyThekumprampilUniversityofIllinoisatUrbana-Champaignthekump2@illinois.eduPrateekJainMicrosoftResearch,Indiaprajain@microsoft.comPraneethNetrapalliMicrosoftResearch,Indiapraneeth@microsoft.comSewoongOhUniversityofWashington,Seattlesewoong@cs.washington.eduAbstractThispaperstudies\ufb01rstordermethodsforsolvingsmoothminimaxoptimizationproblemsminxmaxyg(x,y)whereg(\u00b7,\u00b7)issmoothandg(x,\u00b7)isconcaveforeachx.Intermsofg(\u00b7,y),weconsidertwosettings\u2013stronglyconvexandnonconvex\u2013andimproveuponthebestknownratesinboth.Forstrongly-convexg(\u00b7,y),\u2200y,weproposeanewdirectoptimalalgorithmcombiningMirror-ProxandNesterov\u2019sAGD,andshowthatitcan\ufb01ndglobaloptimumineO(cid:0)1/k2(cid:1)iterations,improvingovercurrentstate-of-the-artrateofO(1/k).WeusethisresultalongwithaninexactproximalpointmethodtoprovideeO(cid:0)1/k1/3(cid:1)ratefor\ufb01ndingstationarypointsinthenonconvexsettingwhereg(\u00b7,y)canbenonconvex.Thisimprovesovercurrentbest-knownrateofO(1/k1/5).Finally,weinstantiateourresultfor\ufb01nitenonconvexminimaxproblems,i.e.,minxmax1\u2264i\u2264mfi(x),withnonconvexfi(\u00b7),toobtainconvergencerateofO(m1/3\u221alogm/k1/3).1IntroductionInthispaperwestudysmoothminimaxproblemsoftheform:minx\u2208Xmaxy\u2208Yg(x,y),g:X\u00d7Y\u2192R,gissmoothi.e.,gradientLipschitz.(1)Theproblemhasapplicationsinseveraldomainssuchasmachinelearning[15,29],optimization[5],statistics[3],mathematics[23],andgametheory[31].Giventheimportanceoftheseproblems,thereisanextensivebodyofworkthatstudiesvariousalgorithmsandtheirconvergenceproperties.Thevastmajorityofexistingresultsforthisproblemfocusontheconvex-concavesetting,whereg(\u00b7,y)isconvexforeveryyandg(x,\u00b7)isconcaveforeveryx.ThebestknownconvergencerateinthissettingisO(1/k)fortheprimal-dualgap,achievedforexamplebyMirror-Prox[34].Thisrateisalsoknowntobeoptimalfortheclassofsmoothconvex-concaveproblems[41].Anaturalquestioniswhetherwecanachieveafasterconvergenceifwehavestrongconvexity(asopposedtojustconvexity)ofg(\u00b7,y).Weanswerthisintheaf\ufb01rmative,byintroducinganalgorithmthatachievesaconvergencerateofeO(cid:0)1/k2(cid:1)forthegeneralsmooth,strongly-convex\u2013concaveminimaxproblem.ThealgorithmweproposeisanovelcombinationofMirror-ProxandNesterov\u2019sacceleratedgradientdescent.Thismatchestheknownlowerboundof\u2126(1/k2)from[41],closingthegapuptoapoly-logarithmicfactor.Therealsoexistsaconceptuallysimplesmoothingtechniquebasedindirectalgorithm,whichpre\ufb01xesthetoleranceof\u03b5.However,ourgoalisto\ufb01ndadirectalgorithmwhichdoesnotpre\ufb01xthetolerance.OtherknownmethodsthatobtainarateofO(1/k2)inthiscontextareforveryspecialcases,wherexandyareconnectedthroughabi-lineartermorg(x,\u00b7)islineariny[35,20,14,8,49,16,48].33rdConferenceonNeuralInformationProcessingSystems(NeurIPS2019),Vancouver,Canada.\fSettingOptimalitynotionPreviousstate-of-the-artOurresultsSmoothingschemesLowerboundConvexPrimal-dualgapO(cid:0)k\u22121(cid:1)[34]--\u2126(k\u22121)[41]StronglyconvexPrimal-dualgapO(cid:0)k\u22121(cid:1)[34]eO(cid:0)k\u22122(cid:1)eO(cid:0)k\u22122(cid:1)\u2126(k\u22122)[41]NonconvexApprox.stat.pointO(cid:0)k\u22121/5(cid:1)[18]eO(cid:0)k\u22121/3(cid:1)eO(cid:0)k\u22121/3(cid:1)[26]-Table1:Comparisonofourresultswithpreviousstate-of-the-art.Weassumethatg(\u00b7,\u00b7)issmooth(i.e.,hasLipschitzgradients)andg(x,\u00b7)isconcave\u2200x\u2208X.Convexity,strongconvexityandnonconvexityinthe\ufb01rstcolumnreferstog(\u00b7,y)for\ufb01xedy.Smoothingschemesareindirectmethodsusingthesmoothingtechnique[36].Whilemosttheoreticalresultsfocusontheconvex-concavesetting,severalrealworldproblemsfalloutsidethisclass.Aslightlylargerclass,whichcapturesseveralmoreapplications,istheclassofsmoothnonconvex\u2013concaveminimaxproblems,whereg(x,\u00b7)isconcaveforeveryxbutg(\u00b7,y)canbenonconvex.Forexample,\ufb01niteminimaxproblems,i.e.,minxmaxmi=1fi(x)=minxmax0(cid:22)y(cid:22)1,Pmi=1yi=1Piyi\u00b7fi(x):=g(x,y)belongtothisclass,andsodosmoothnon-convexconstrainedoptimizationproblems[25].Inaddition,severalmachinelearningproblemswithnon-decomposablelossfunctions[22]alsobelongtothisclass.Inthisgeneralnonconvexconcavesettinghowever,wecannothopeto\ufb01ndglobaloptimumef\ufb01cientlyaseventhespecialcaseofnonconvexoptimizationisNP-hard.Similartononconvexoptimization,wemighthopeto\ufb01ndanapproximatestationarypoint[37].Oursecondcontributionisanewalgorithmandafasterrateforthegeneralsmoothnonconvex\u2013concaveminimaxproblem.Ouralgorithmisaninexactproximalpointmethodforthenonconvexfunctionf(x):=maxy\u2208Yg(x,y).Thekeyinsightisthattheproximalpointproblemineachiterationresultsinastrongly-convexconcaveminimaxproblem,forwhichweuseourimprovedalgorithmtoobtaintheoverallcomputation/iterationcomplexityofeO(cid:0)1/k1/3(cid:1)thusimprovingoverthepreviousbestknownrateofO(1/k1/5)[18]1.Morerecently,independenttoourwork,asmoothingbasedalgorithmhasalsobeenproposedtoachievethesameO(cid:0)k\u22121/3(cid:1)rate[26].Finally,wespecializeourresultto\ufb01niteminimaxproblems,i.e.,minxmax1\u2264i\u2264mfi(x)wherefi(x)canbenonconvexfunctionbuteachfiisasmoothfunction;nonconvexconstrainedopti-mizationproblemscanbereducedtosuch\ufb01niteminimaxproblems.Forthese,weobtainarateofeO(cid:0)m1/3\u221alogm/k1/3(cid:1)totalgradientcomputationswhichimprovesuponthestate-of-the-artrateO(m1/4/k1/4)[11]inthissettingaswell.Summaryofcontributions:SeealsoTable1.1.OptimaleO(cid:0)1/k2(cid:1)convergencerateforsmooth,strongly-convex\u2013concaveproblems,improvinguponthepreviousbestknownrateofO(1/k)foradirectalgorithmand,2.eO(cid:0)1/k1/3(cid:1)convergencerateforsmooth,nonconvex\u2013concaveproblems,improvinguponthepreviousbestknownrateofO(cid:0)1/k1/5(cid:1).Relatedworks:Forstrongly-convex-concaveminimaxproblemswithspecialstructures,severalalgorithmshavebeenproposed.Inanincreasingorderofgenerality,[14,49,50]studyoptimizingastronglyconvexfunctionwithlinearconstraints,whichcanbeposedasaspecialcaseofminimaxoptimization,[35,8]studyaminimaxproblemwherexandyareconnectedonlythroughabi-lineartermyTAx,and[16,20]studyacasewhereg(x,\u00b7)islineariny.Inallthesecases,itisshownthatO(1/k2)convergencerateisachievableifg(\u00b7,y)isstrongly-convex\u2200y.Recently,[12]showedlinearconvergenceofgradientdescentascentforstrongly-convex\u2013concaveproblemwithbilinearcouplingwhenAhasfullrowrank.However,ithasremainedanopenquestionifthefastrateofO(1/k2)canbeachievedforgeneralstrongly-convex-concaveminimaxproblems.See[32,9,7,17,51,1]1While[18]givesarateofO(cid:0)1/k1/4(cid:1)withanapproximatemaximizationoracleformaxy\u2208Yg(x,y),takingintoaccountthecostofimplementingsuchamaximizationoraclegivesarateofO(cid:0)1/k1/5(cid:1).2\ffordetailedsurveysontheresultsfortheconvex-concaveminimaxproblems.Forexamplesandapplicationofsaddlepointproblemsrefer[36,19,20,7,43].Fornonconvex-concaveminimaxproblems,[42]considersbothdeterministicandstochasticsettings,andproposesinexactproximalpointmethodsforsolvingsmoothnonconvex\u2013concaveproblems.Inthedeterministicsetting,theirresultguaranteesanerrorofO(1/k1/6).Wenotethattherehavealsobeenothernotionsofstationarityproposedinliteraturefornonconvex-concaveminimaxproblems[28,40].Thesenotionshoweverareweakerthantheoneconsideredinthispaper,inthesensethat,ournotionofstationarityimpliestheseothernotions(withoutlossinparameters).Foronesuchweakernotion,[40]proposesanalgorithmwithaconvergencerateofO(cid:0)k\u22121/3.5(cid:1).Sincethenotiontheyconsiderisweaker,itdoesnotimplythesameconvergencerateinoursetting.Wewouldalsoliketohighlighttheworks[6,13,33,46,34,47,10]designingef\ufb01cientalgorithmsforsolvingmonotonevariationalinequalitieswhichgeneralizestheconvex-concaveminimaxproblems.Notations:Risthereallineandforanynaturalnumberp,Rpistherealvectorspaceofdimensionp.k\u00b7kisanormonsomemetricspacewhichwouldbeevidentfromthecontext.ForaconvexsetX\u2286Rpandx\u2208Rp,PX(x)=argminx0\u2208Xkx\u2212x0kistheprojectionofxontoX.Foradifferentiablefunctiong(x,y),\u2207xg(x,y)isitsgradientwithrespecttoxat(x,y).Weusethestandardbig-Onotations.ForfunctionsT,S:R\u2192Rsuchthat0<liminfx\u2192\u221eT(x),0<liminfx\u2192\u221eS(x),(a)T(x)=O(S(x))meanslimsupx\u2192\u221eT(x)/S(x)<\u221e;(b)T(x)=\u0398(S(x))meansT(x)=O(S(x))andS(x)=O(T(x));and(c)T(x)=eO(S(x))meansthatT(x)=O(S(x)R(x))forsomepoly-logarithmicfunctionR:R\u2192R.Paperorganization:InSection2,wepresentpreliminariesandallrelevantbackground.InSection3,wepresentourresultsforstrongly-convex\u2013concavesettingandinsection4,resultsfornonconvex\u2013concavesetting.InSection5,wepresentempiricalevaluationofouralgorithmfornonconvex-concavesettingandcompareittoastate-of-the-artalgorithm.WeconcludeinSection6.Severaltechnicaldetailsarepresentedintheappendix.2PreliminariesandbackgroundmaterialInthissection,wewillpresentsomepreliminaries,describingthesetupandreviewingsomeback-groundmaterialthatwillbeusefulinthesequel.2.1MinimaxproblemsWeareinterestedintheminimaxproblemsoftheform(1)whereg(x,y)isasmoothfunction.De\ufb01nition1.Afunctiong(x,y)issaidtobeL-smoothif:max{k\u2207xg(x,y)\u2212\u2207xg(x0,y0)k,k\u2207yg(x,y)\u2212\u2207yg(x0,y0)k}\u2264L(kx\u2212x0k+ky\u2212y0k).Throughout,weassumethatg(x,.)isconcaveforeveryx\u2208X.Forg(\u00b7,y)behaviorintermsofx,therearebroadlytwosettings:2.1.1Convex-concavesettingInthissetting,g(\u00b7,y)isconvex\u2200y\u2208Y.Givenanygand\u2200(bx,by),thefollowingholdstrivially:minx\u2208Xg(x,by)\u2264g(bx,by)\u2264maxy\u2208Yg(bx,y),whichthenimpliesthatmaxy\u2208Yminx\u2208Xg(x,y)\u2264minx\u2208Xmaxy\u2208Yg(x,y).Thecelebratedmin-imaxtheoremfortheconvex-concavesetting[44]saysthatifYisacompactsetthentheaboveinequalityisinfactanequality,i.e.,maxy\u2208Yminx\u2208Xg(x,y)=minx\u2208Xmaxy\u2208Yg(x,y).Further-more,anypoint(x\u2217,y\u2217)isanoptimalsolutionto(1)ifandonlyif:minx\u2208Xg(x,y\u2217)=g(x\u2217,y\u2217)=maxy\u2208Yg(x\u2217,y).(2)Hence,ourgoalisto\ufb01nd\u03b5-primal-dualpair(bx,by)withsmallprimal-dualgap:maxy\u2208Yg(bx,y)\u2212minx\u2208Xg(x,by).De\ufb01nition2.Foraconvex-concavefunctiong:X\u00d7Y\u2192R,(\u02c6x,\u02c6y)isan\u03b5-primal-dual-pairofgiftheprimal-dualgapislessthan\u03b5:maxy\u2208Yg(bx,y)\u2212minx\u2208Xg(x,by)\u2264\u03b5.3\f2.1.2Nonconvex-concavesettingInthissettingthefunctiong(\u00b7,y)neednotbeconvex.Onecannothopetosolvesuchproblemsingeneral,sincethespecialcaseofnonconvexoptimizationisalreadyNP-hard[39].Furthermore,theminimaxtheoremnolongerholds,i.e.,maxy\u2208Yminx\u2208Xg(x,y)canbestrictlysmallerthanminx\u2208Xmaxy\u2208Yg(x,y),andthereforetheorderofminandmaxmightbeimportantforagivenapplicationi.e.,wemightbeinterestedonlyinminimaxbutnotmaximin(orviceversa).So,theprimal-dualgapmaynotbeameaningfulquantitytomeasureconvergence.Inthispaperwewillfocusontheminimaxproblem:minx\u2208Xmaxy\u2208Yg(x,y).Oneapproach,inspiredbynonconvexoptimization,tomeasureconvergenceistoconsiderthefunctionf(x)=maxy\u2208Yg(x,y)andconsidertheconvergenceratetoapproximate\ufb01rstorderstationarypoints(i.e.,\u2207f(x)issmall)[42,18].Butasf(x)couldbenon-smooth,\u2207f(x)mightnotevenbede\ufb01ned.Itturnsoutthatwheneverg(x,y)issmooth,f(x)isweaklyconvex(De\ufb01nition4)forwhich\ufb01rstorderstationaritynotionsarewell-studiedandarediscussedbelow.Approximate\ufb01rst-orderstationarypointforweaklyconvexfunctions:We\ufb01rstneedtogeneral-izethenotionofgradientforanon-smoothfunction.De\ufb01nition3.TheFr\u00e9chetsub-differentialofafunctionf(\u00b7)atxisde\ufb01nedastheset,\u2202f(x)={u|liminfx0\u2192xf(x0)\u2212f(x)\u2212hu,x0\u2212xi/kx0\u2212xk\u22650}.Inordertode\ufb01neapproximatestationarypoints,wealsoneedthenotionofweaklyconvexfunctionandMoreauenvelope.De\ufb01nition4.Afunctionf:X\u2192R\u222a{\u221e}isL-weaklyconvexif,f(x)+hux,x0\u2212xi\u2212L2kx0\u2212xk2\u2264f(x0),(3)forallFr\u00e9chetsubgradientsux\u2208\u2202f(x),forallx,x0\u2208X.De\ufb01nition5.Foraproperlowersemi-continuous(l.s.c.)functionf:X\u2192R\u222a{\u221e}and\u03bb>0(X\u2286Rp),theMoreauenvelopefunctionisgivenbyf\u03bb(x)=minx0\u2208Xf(x0)+12\u03bbkx\u2212x0k2.(4)Lemma4(inAppendixB.2)providessomeusefulpropertiesoftheMoreauenvelopeforweaklyconvexfunctions.Now,\ufb01rstorderstationarypointofanon-smoothnonconvexfunctioniswell-de\ufb01ned,i.e.,x\u2217isa\ufb01rstorderstationarypoint(FOSP)ofafunctionf(x)if,0\u2208\u2202f(x\u2217)(seeDe\ufb01nition3).However,unlikesmoothfunctions,itisnontrivialtode\ufb01neanapproximateFOSP.Forexample,ifwede\ufb01nean\u03b5-FOSPasthepointxwithminu\u2208\u2202f(x)kuk\u2264\u03b5,theremayneverexistsuchapointforsuf\ufb01cientlysmall\u03b5,unlessxisexactlyaFOSP.Incontrast,byusingabovepropertiesoftheMoreauenvelopeofaweaklyconvexfunction,it\u2019sapproximateFOSPcanbede\ufb01nedas[11]:De\ufb01nition6.GivenanL-weaklyconvexfunctionf,wesaythatx\u2217isan\u03b5-\ufb01rstorderstationarypoint(\u03b5-FOSP)if,k\u2207f12L(x\u2217)k\u2264\u03b5,wheref12ListheMoreauenvelopewithparameter1/2L.UsingLemma4,wecanshowthatforany\u03b5-FOSPx\u2217,thereexists\u02c6xsuchthatk\u02c6x\u2212x\u2217k\u2264\u03b5/2Landminu\u2208\u2202f(\u02c6x)kuk\u2264\u03b5.Inotherwords,an\u03b5-FOSPisO(\u03b5)closetoapoint\u02c6xwhichhasasubgradientsmallerthan\u03b5.WenotethatothernotionsofFOSPhavealsobeenproposedrecentlysuchasin[40].However,itcanbeshownthatan\u03b5-FOSPaccordingtotheabovede\ufb01nitionisalsoan\u0001-FOSPwith[40]\u2019sde\ufb01nitionaswell,butthereverseisnotnecessarilytrue.2.2Mirror-ProxMirror-Prox[34]isapopularalgorithmproposedforsolvingconvex-concaveminimaxproblems(1).ItachievesaconvergencerateofO(1/k)fortheprimaldualgap.TheoriginalMirror-Proxpaper[34]motivatesthealgorithmthroughaconceptualMirror-Prox(CMP)method,whichbringsoutthemainideabehinditsconvergencerateofO(1/k).CMPforEuclideannorm(afterignoringprojectionstoXandY)doesthefollowingupdate:(xk+1,yk+1)=(xk,yk)+1\u03b2(\u2212\u2207xg(xk+1,yk+1),\u2207yg(xk+1,yk+1)).(5)4\fThemaindifferencebetweenCMPandstandardgradientdescentascent(GDA)isthatinthekthstep,whileGDAusesgradientsat(xk,yk),CMPusesgradientsat(xk+1,yk+1).Thekeyobservationof[34]isthatifg(\u00b7,\u00b7)issmooth,itcanbeimplementedef\ufb01ciently.CMPisanalyzedasfollows:ImplementabilityofCMP:Let(x(0)k,y(0)k)=(xk,yk).For\u03b2<1L,theiteration(x(i+1)k,y(i+1)k)=(xk,yk)+1\u03b2(cid:16)\u2212\u2207xg(cid:16)x(i)k,y(i)k(cid:17),\u2207yg(cid:16)x(i)k,y(i)k(cid:17)(cid:17).(6)canbeshowntobe1\u221a2-contraction(wheng(\u00b7,\u00b7)issmooth)andthatits\ufb01xedpointis(xk+1,yk+1).So,inlog1\u0001iterationsof(6),wecanobtainanaccurateversionoftheupdaterequiredbyCMP.Infact,[34]showedthatjusttwoiterationsof(6)suf\ufb01ce[30].ConvergencerateofCMP:UsingCMPupdatewithsimplemanipulationsleadstothefollowing:g(xk+1,y)\u2212g(x,yk+1)\u2264\u03b22(cid:0)kx\u2212xkk2\u2212kx\u2212xk+1k2+ky\u2212ykk2\u2212ky\u2212yk+1k2(cid:1),\u2200x\u2208X,y\u2208Y.O(1/k)convergenceratefollowseasilyusingtheaboveresult.Finally,ourmethodandanalysisalsorequiresNesterov\u2019sacceleratedgradientdescentmethod(seeAlgorithm3inAppendixA)andit\u2019sper-stepanalysisby[2](Lemma2inAppendixA).3Strongly-convexconcavesaddlepointproblemWe\ufb01rststudytheminimaxproblemoftheform:minx\u2208X[f(x)=maxy\u2208Yg(x,y)],(P1)whereg(x,\u00b7)isconcave,g(\u00b7,y)is\u03c3-strongly-convex,g(\u00b7,\u00b7)isL-smooth,i.e.,0<\u03c3\u2264L.X=RpandY\u2282Rqisaconvexcompactsub-setofRqandletthefunctionftakeaminimumvaluef\u2217(>\u2212\u221e).LetDY=maxy,y0\u2208Yky\u2212y0kbethediameterofY.Ourobjectivehereisto\ufb01ndan\u0001-primal-dualpair(bx,by)(seeDe\ufb01nition2).Nowthefactthatf(\u02c6x)\u2212f\u2217\u2264maxy\u2208Yg(bx,y)\u2212minx\u2208Xg(x,by)impliesthatif(\u02c6x,\u02c6y)isan\u03b5-primal-dual-pair,then\u02c6xisalsoan\u03b5-approximateminimaoff.Furthermore,bySion\u2019sminimaxtheorem[24],strong-convexity\u2013concavityofg(\u00b7,\u00b7)ensuresthat:minx[f(x):=maxyg(x,y)]=maxy[h(y):=minxg(x,y)].Hence,oneapproachtoef\ufb01cientlysolvingtheproblemisbyoptimizingthedualproblemmaxyh(y).ByLemma6(inAppendixB.6),h(y)isan(L+L2\u03c3)-smoothfunction.SowecanuseAGDtoensurethath(yk)\u2212h(y\u2217)=O(1/k2).Now,eachstepofAGDrequirescomputingargminxg(x,yk)whichcanbedoneef\ufb01ciently(i.e.,logarithmicnumberofsteps)asg(\u00b7,yk)isstrongly-convexandsmooth.So,theoverall\ufb01rst-orderoraclecomplexityish(yk)\u2212h(y\u2217)=eO(cid:0)1/k2(cid:1).Sodoesthissimpleapproachgiveusourdesiredresult?Unfortunatelythatisnotthecase,astheaboveboundonthedualfunctionhdoesnottranslatetothesameerrorrateforprimalfunctionf,i.e.,thesolutionneednotbeeO(cid:0)1/k2(cid:1)-primal-dualpair.E.g.,considerminx\u2208Rmaxy\u2208[\u22121,1][g(x,y)=xy+x2/2],whereminxmaxyg(x,y)=0,f(x)=x2/2+|x|andh(y)=\u2212y2/2.Ifh(yk)=\u0398(k\u22122),thenxk\u2208argminxg(x,yk)=\u0398(1/k)andsof(xk)is\u0398(k\u22121).Thisisduetothenon-smoothnessofargmaxy\u2208Yg(x,y)w.r.t.x.InsteadofusingAGD,weintroduceanewmethodtosolvethedualproblemthatwerefertoasDIAG,whichstandsforDualImplicitAcceleratedGradient.DIAGcombinesideasfromAGD[38]andNemirovski\u2019soriginalderivationoftheMirror-Proxalgorithm[34],andcanensureafastconvergencerateof\u02dcO(k\u22122)fortheprimal-dualgap.Wenotethattherealsoexistsaconceptuallysimplersmoothingtechniquebasedindirectalgorithm,whichpre\ufb01xesthetoleranceof\u03b5(AppendixD).However,ourgoalisto\ufb01ndadirectalgorithmwhichdoesnoterequirepre\ufb01xingthetoleranceat\u03b5.Forbetterexposition,we\ufb01rstpresentaconceptualversionofDIAG(C-DIAG),whichisnotimplementableexactly,butbringsoutthemainnewideasinouralgorithm.Wethenpresentadetailederroranalysisfortheinexactversionofthisalgorithm,whichisimplementable.3.1Conceptualversion:C-DIAGConsiderthefollowingupdateswhichisamodi\ufb01edversionofAGD(seeAlgorithm3inAppendixA):5\f(a)wk=(1\u2212\u03c4k)yk+\u03c4kzk(b)Choosexk+1,yk+1ensuring:xk+1\u2208argminxg(x,yk+1),andyk+1=PY(wk+1\u03b2\u2207yg(xk+1,wk))(c)zk+1=PY(zk+\u03b7k\u2207yg(xk+1,wk))CompletepseudocodeforC-DIAGalgorithmispresentedinAlgorithm4inAppendixB.4.ThemainideaofthealgorithmisinStep(b)above(i.e.,Step4ofAlgorithm4inAppendixB.4),wherewesimultaneously\ufb01ndxk+1andyk+1satisfyingthefollowingrequirements:\u2022xk+1istheminimizerofg(\u00b7,yk+1),and\u2022yk+1correspondstoanAGDstep(seeAlgorithm3inAppendixA)forg(xk+1,\u00b7)Implementability:The\ufb01rstquestioniswhetheritiseasyenoughtoimplementsuchastep?Itturnsoutthatitisindeedpossibletoquickly\ufb01ndpointsxk+1andyk+1thatapproximatelysatisfytheaboverequirements.Thereasonisthat:\u2022Sinceg(\u00b7,y)issmoothandstronglyconvexforeveryy\u2208Y,wecan\ufb01nd\u0001-approximateminimizerforagivenyinO(cid:0)log1\u0001(cid:1)iterations.\u2022Letx\u2217(y):=argminx\u2208Xg(x,y).Theiterationyi+1=PY(cid:16)wk+1\u03b2\u2207yg(x\u2217(yi),wk)(cid:17)isa1/2-contractionwithaunique\ufb01xedpointsatisfyingtheupdatesteprequirements(i.e.,Step4ofAlgorithm4inAppendixB.4).SeeLemma5inAppendixB.5foraproof.ThismeansthatonlyO(cid:0)log1\u0001(cid:1)iterationsagainsuf\ufb01ceto\ufb01ndanupdatethatapproximatelysatis\ufb01estherequirements.Convergencerate:Sinceyk+1andzk+1correspondtoanAGDupdateforg(xk+1,\u00b7),wecanusethepotentialfunctiondecreaseargumentforAGD(Lemma2inAppendixA)toconcludethat\u2200y\u2208Y,(k+1)(k+2)(g(xk+1,y)\u2212g(xk+1,yk+1))+2\u03b2\u00b7ky\u2212zk+1k2\u2264k(k+1)(g(xk+1,y)\u2212g(xk+1,yk))+2\u03b2\u00b7ky\u2212zkk2\u2264k(k+1)(g(xk+1,y)\u2212g(xk,y))+k(k+1)(g(xk,y)\u2212g(xk,yk))+2\u03b2\u00b7ky\u2212zkk2,wherethelaststepfollowsfromthefactthatxk=argminxg(x,yk)andsog(xk,yk)\u2264g(xk+1,yk).Notingthatwecanfurtherrecursivelyboundk(k+1)(g(xk,y)\u2212g(xk,yk))+2\u03b2\u00b7ky\u2212zkk2asabove,weobtain(k+1)(k+2)(g(xk+1,y)\u2212g(xk+1,yk+1))+2\u03b2\u00b7ky\u2212zk+1k2\u2264k(k+1)g(xk+1,y)\u2212kXi=1(2i)\u00b7g(xi,y)+2\u03b2\u00b7ky\u2212z0k2\u21d2k+1Xi=1(2i)\u00b7g(xi,y)\u2212(k+1)(k+2)g(xk+1,yk+1)\u22642\u03b2\u00b7ky\u2212z0k2.Sinceg(xk+1,yk+1)\u2264g(x,yk+1)foreveryx\u2208X,wehavek+1Xi=1(2i)\u00b7g(xi,y)\u2212(k+1)(k+2)g(x,yk+1)\u22642\u03b2\u00b7ky\u2212z0k2\u21d2g(\u00afxk+1,y)\u2212g(x,yk+1)\u22642\u03b2\u00b7ky\u2212z0k2(k+1)(k+2),where\u00afxk+1:=1(k+1)(k+2)Pk+1i=1(2i)\u00b7xi.Sincexandyarearbitraryabove,thisgivesaO(cid:0)1/k2(cid:1)convergenceratefortheprimaldualgap.3.2ErroranalysisThemainissuewithAlgorithm4isthattheupdatestepisnotexactlyimplementable.However,aswenotedintheprevioussection,wecanquickly\ufb01ndupdatesthatalmostsatisfytherequirements.Algorithm1presentsthisinexactversion.ThefollowingtheoremstatesourformalresultandadetailedproofisprovidedinAppendixB.5.6\fAlgorithm1:DualImplicitAcceleratedGradient(DIAG)forstrongly-convex\u2013concaveprogrammingInput:g,L,\u03c3,DY,x0,y0,K,{\u03b5(k)step}Kk=1Output:\u00afxK,yK1Set\u03b2\u21902L2\u03c3,z0\u2190y02fork=0,1,...,K\u22121do3\u03c4k\u21902(k+2),\u03b7k\u2190(k+1)2\u03b2,wk\u2190(1\u2212\u03c4k)yk+\u03c4kzk4xk+1,yk+1\u2190Imp-STEP(g,L,\u03c3,x0,wk,\u03b2,\u03b5(k+1)step),ensuring:g(xk+1,yk+1)\u2264minxg(x,yk+1)+\u03b5(k+1)step,yk+1=PY(cid:18)wk+1\u03b2\u2207yg(xk+1,wk)(cid:19)5zk+1\u2190PY(zk+\u03b7k\u2207yg(xk+1,wk)),\u00afxk+1\u21902(k+1)(k+2)Pk+1i=1i\u00b7xi6return\u00afxK,yK7Imp-STEP(g,L,\u03c3,x0,w,\u03b2,\u03b5step):8Set\u03b5mp\u21902\u03c35Lq2\u03b5stepL,R\u2190dlog22DY\u03b5mpe,\u03b5agd\u2190\u03c3\u03b22\u03b52mp32L2,y0\u2190w9forr=0,1,...,Rdo10Startingatx0useAGD[38]forstrongly-convexg(\u00b7,yr),tocompute\u02c6xrsuchthat:g(\u02c6xr,yr)\u2264minxg(x,yr)+\u03b5agd,(7)11yr+1\u2190PY(cid:0)w+1\u03b2\u2207yg(\u02c6xr,w)(cid:1)12return\u02c6xR,yR+1Theorem1(ConvergencerateofDIAG).Letg:X\u00d7Y\u2192RbeaL-smooth,\u03c3-strongly-convex\u2013concavefunctiononX=Rpandaconvexcompactsub-setY\u2282Rq(withdiameterDY).Then,afterKiterations,DIAG(Algorithm1)withatolerancescheduleof{\u03b5(k)step}Kk=1foritsImp-STEPsub-routine,\ufb01nds(\u00afxK,yK)s.t.:max\u02dcy\u2208Yg(\u00afxK,\u02dcy)\u2212min\u02dcx\u2208Xg(\u02dcx,yK)\u22644L2\u03c3D2Y+PKk=1k(k+1)\u03b5(k)stepK(K+1).(8)Inparticular,setting\u03b5(k)step=L2D2Y\u03c3k3(k+1)wehave:max\u02dcy\u2208Yg(\u00afxK,\u02dcy)\u2212min\u02dcx\u2208Xg(\u02dcx,yK)\u22646L2\u03c3D2YK(K+1).Furthermore,forthissettingthetotal\ufb01rstorderoraclecomplexityisgivenby:O(qL\u03c3Klog2(K)).Remark1:Theorem1showsthatDIAGneeds\u02dcO((DYL/\u221a\u03c3\u03b5)\u00b7(pL/\u03c3))gradientqueriesfor\ufb01ndinga\u03b5-primal-dual-pair,whilecurrentbest-knownrateisO(1/\u03b5)achievedbyMirror-Prox.Thisdependencein\u03b5andDYisoptimal,asitisshownin[41,Theorem10]that\u2126(DY(L\u2212\u03c3)/\u221a\u03c3\u03b5)gradientqueriesarenecessarytoachieve\u03b5errorintheprimal-dualgap.Remark2:UnlikestandardAGDforh(y),whichonlyupdatesykintheouter-loop,DIAG\u2019souter-stepupdatesbothxkandykthusallowingustobettertracktheprimal-dualgap.However,DIAG\u2019sdependenceontheconditionnumberL/\u03c3seemssub-optimalandcanperhapsbeimprovedifwedonotcomputeImp-STEPnearlyoptimallyallowingforinexactupdates;weleavefurtherinvestigationintoimproveddependenceontheconditionnumberforfuturework.4NonconvexconcavesaddlepointproblemWestudythenonconvexconcaveminimaxproblem(1)whereg(x,\u00b7)isconcave,g(\u00b7,y)isnonconvex,andg(\u00b7,\u00b7)isL-smooth,X=Rp(suchthatProjX(x)=x)andYisaconvexcompactsub-setofRq.AsmentionedinSection2,wemeasuretheconvergencetoanapproximateFOSPofthisproblem(seeDe\ufb01nition6)butitrequiresweak-convexityoff(x):=maxy\u2208Yg(x,y).Thefollowinglemmaguaranteesweakconvexityoffgivensmoothnessofg.7\fLemma1.Letg(\u00b7,y)becontinuousandYbecompact.Thenf(x)=maxy\u2208Yg(x,y)isL-weaklyconvex,ifgisL-weaklyconvexinx(De\ufb01nition1),orifgisL-smoothinx.SeeAppendixB.3fortheproof.Theargumentsof[18]easilyextendtoshowthatapplyingsubgradientmethodonf(x),[11]givesaconvergencerateofO(cid:0)1/k1/5(cid:1).Instead,weexploitthesmoothminimaxformoff(\u00b7)todesignafasterconvergingscheme.Themainintuitioncomesfromtheproximalviewpointthatgradientdescentcanbeviewedasiterativelyformingandoptimizinglocalquadraticupperbounds.Asfisweaklyconvex,addingenoughquadraticregularizationshouldensurethattheresultingsequenceofproblemsareallstrongly-convex\u2013concave.WethenexploitDIAGtoef\ufb01cientlysolvesuchlocalquadraticproblemstoobtainimprovedconvergencerates.Concretely,letbf(x;xk)=maxyg(x,y)+Lkx\u2212xkk2.(9)ByL-weak-convexityoff,bf(x;xk)isstrongly-convex\u2013concave(Lemma3)thatcanbesolvedusingDIAGuptocertainaccuracytoobtainxk+1.WerefertothisalgorithmasProx-DIAGandprovideapseudo-codeforthesameinAlgorithm2.ThefollowingtheoremgivesconvergenceguaranteesforAlgorithm2:ProximalDualImplicitAcceleratedGradient(Prox-DIAG)fornonconvexconcaveprogrammingInput:g,L,\u03b5,x0,y0Output:xk1Set\u02dc\u03b5\u2190\u03b5264L2fork=0,1,...,Kdo3UsingDIAGforstronglyconvexconcaveminimaxproblem,minxmaxy\u2208Y[bg(x,y;xk)=g(x,y)+Lkx\u2212xkk2](10)\ufb01ndxk+1suchthat,maxy\u2208Yg(xk+1,y)+Lkxk+1\u2212xkk2\u2264minxmaxy\u2208Yg(x,y)+Lkx\u2212xkk2+\u02dc\u03b54(11)ifmaxy\u2208Yg(xk,y)\u22123\u02dc\u03b54\u2264maxy\u2208Yg(xk+1,y)+Lkxk+1\u2212xkk2then4returnxkProx-DIAG.Theorem2(ConvergencerateofProx-DIAG).Letg(x,y)beL-smooth,g(x,\u00b7)beconcave,XbeRp,YbeaconvexcompactsubsetofRq,andtheminimumvalueoffunctionf(x)=maxy\u2208Yg(x,y)beboundedbelow,i.e.f(x)\u2265f\u2217>\u2212\u221e.ThenProx-DIAG(Algorithm2)after,K=(cid:24)44L(f(x0)\u2212f\u2217)3\u03b52(cid:25)stepsoutputsan\u03b5-FOSP.Thetotal\ufb01rst-orderoraclecomplexitytooutput\u03b5-FOSPis:O(cid:0)L2DY(f(x0)\u2212f\u2217)\u03b53log2(cid:0)1\u03b5(cid:1)(cid:1).AproofisprovidedinAppendixB.7.NotethatProx-DIAGsolvesthequadraticapproximationproblemtohigheraccuracyofO(\u00012)whichthenhelpsboundingthegradientoftheMoreauenvelope.Alsoduetothemodularstructureoftheargument,afasterinnerloopforspecialsettings,e.g.,wheng(x,y)isa\ufb01nite-sum,canensuremoreef\ufb01cientalgorithm.Whileouralgorithmisabletosigni\ufb01cantlyimproveuponexistingstate-of-the-artrateofO(1/\u03b55)ingeneralnonconvex-concavesetting[18],itisuncleariftheratecanbefurtherimproved.Infact,preciselower-boundsforthissettingaremostlyunexploredandweleavefurtherinvestigationintolower-boundsasatopicoffutureresearch.WealsospecializetheProx-DIAGalgorithm,asProx-FDIAG(Algorithm5inAppendixC),forthecaseofminimizingaweaklyconvexf(x),withthespecialstructureof\ufb01nitemax-typefunction:minxhf(x)=max1\u2264i\u2264mfi(x)i,(P3)8\fwherefi\u2019scouldbenonconvexbutareL-smooth,G-Lipschitzandboundedfrombelow.Forthiscase,weimprovethecurrentknownbestrateofO(cid:0)m/\u03b54(cid:1)andobtainafasterrateofO(mlog3/2m/\u03b53)usingtheProx-FDIAGalgorithm.PleaserefertoAppendixCformoredetails.5ExperimentsWeempiricallyverifytheperformanceofProx-FDIAG(Algorithm5inAppendixC)onasyn-thetic\ufb01nitemax-typenonconvexminimizationproblem(P3).Weconsiderthefollowingproblem.minx\u2208R2(cid:2)f(x)=max1\u2264i\u2264m=9fi(x)(cid:3)wherefi(x)=q(\u22121,(X(1)i,X(2)i),Ci)(x)forall1\u2264i\u22648,whereq(a,b,c)(x)=akx\u2212bk22+c,X(1)i,X(2)i,andciarerandomlygenerated.ThuseachfiissmoothwithparameterL=1,whichimpliesthatfisL-weaklyconvex.Weimplementthreealgorithms:Prox-FDIAG(Algorithm5,redcircles),AdaptiveProx-FDIAG(Algorithm6,blackdots),andsubgradientmethod[11](bluetriangles).AdaptiveProx-FDIAGisapracticallyfastervariantofProx-FDIAG,withthesame\ufb01rst-orderoraclecomplexityguarantee(uptoanO(log(1/\u03b5))factor).InFigure1,weplotthenormofgradientofMoreauenvelopek\u2207f12L(xk)k2againstthenumber100101102103104105106107108106104102100102Sub-gradient methodProx-FDIAG (ours)Adaptive Prox-FDIAG (ours)NormofthegradientofMoreauenvelopek\u2207f12L(xk)k2numberofgradientoracleaccesseskFigure1:Forsmalltargetaccuracy\u03b5regime,AdaptiveProx-FDIAG(ours)hasthefastestconvergenceratefollowedbyProx-FDIAG(ours)andsubgradientmethod.ofiterationskinlog-logscale.Weseethat,Prox-FDIAGandAdaptiveProx-FDIAGhaveafasterconvergenceratethansubgradientmethod,andAdaptiveProx-FDIAGisalmostalwaysfasterthanProx-FDIAG.WeprovidemoredetailsaboutthealgorithmsandtheexperimentsinAppendixE.6ConclusionInthispaper,westudysmoothminimaxproblems,wherethemaximizationisconcavebuttheminimizationiseitherstronglyconvexornonconvex.Inbothofthesesettings,wepresentnewalgorithmsimprovingstate-of-the-art.Thekeyideasarei)anovelwaytocombineMirror-ProxandNesterov\u2019sAGDforstronglyconvexcasethatcantightlyboundprimal-dualgapandii)aninexactproxmethodwithgoodconvergenceratetostationarypointsforthenonconvexcase.WhileweonlypresentourresultsfortheEuclideansetting,generalizingittonon-EuclideansettingswiththeframeworkofBregmandivergencesshouldbestraightforward.Finally,weshowcasetheempiricalsuperiorityofournonconvexalgorithmoverstate-of-the-artsubgradientmethodforacaseof\ufb01nitemax-typenonconvexminimizationproblems.Someofthemoreinterestingquestionswouldbetounderstandtheoptimalityoftheratesthatweobtainanddependenceonthestrongconvexityparameter.Furtherextensionsoftheseresultstothestochasticsettingwouldalsobequiteinteresting.AcknowledgementThisworkispartiallysupportedbyNSFawardsCCF-1927712andRI-1929955.9\fReferences[1]MohammadAlkousa,DarinaDvinskikh,FedorStonyakin,andAlexanderGasnikov.Acceler-atedmethodsforcompositenon-bilinearsaddlepointproblem.arXivpreprintarXiv:1906.03620,2019.[2]NikhilBansalandAnupamGupta.Potential-functionproofsfor\ufb01rst-ordermethods.arXivpreprintarXiv:1712.04581,2017.[3]JamesOBerger.StatisticaldecisiontheoryandBayesiananalysis.SpringerScience&BusinessMedia,2013.[4]DimitriPBertsekas.Convexoptimizationtheory.AthenaScienti\ufb01cBelmont,2009.[5]DimitriPBertsekas.ConstrainedoptimizationandLagrangemultipliermethods.Academicpress,2014.[6]RonaldEBruckJr.Ontheweakconvergenceofanergodiciterationforthesolutionofvariationalinequalitiesformonotoneoperatorsinhilbertspace.JournalofMathematicalAnalysisandApplications,61(1):159\u2013164,1977.[7]AntoninChambolleandThomasPock.Anintroductiontocontinuousoptimizationforimaging.ActaNumerica,25:161\u2013319,2016.[8]AntoninChambolleandThomasPock.Ontheergodicconvergenceratesofa\ufb01rst-orderprimal\u2013dualalgorithm.MathematicalProgramming,159(1-2):253\u2013287,2016.[9]YunmeiChen,GuanghuiLan,andYuyuanOuyang.Optimalprimal-dualmethodsforaclassofsaddlepointproblems.SIAMJournalonOptimization,24(4):1779\u20131814,2014.[10]YunmeiChen,GuanghuiLan,andYuyuanOuyang.Acceleratedschemesforaclassofvariationalinequalities.MathematicalProgramming,165(1):113\u2013149,2017.[11]DamekDavisandDmitriyDrusvyatskiy.StochasticsubgradientmethodconvergesattherateO(k\u22121/4)onweaklyconvexfunctions.arXivpreprintarXiv:1802.02988,2018.[12]SimonSDuandWeiHu.Linearconvergenceoftheprimal-dualgradientmethodforconvex-concavesaddlepointproblemswithoutstrongconvexity.InThe22ndInternationalConferenceonArti\ufb01cialIntelligenceandStatistics,pages196\u2013205,2019.[13]JonathanEcksteinandDimitriPBertsekas.Onthedouglas\u2014rachfordsplittingmethodandtheproximalpointalgorithmformaximalmonotoneoperators.MathematicalProgramming,55(1-3):293\u2013318,1992.[14]TomGoldstein,BrendanO\u2019Donoghue,SimonSetzer,andRichardBaraniuk.Fastalternatingdirectionoptimizationmethods.SIAMJournalonImagingSciences,7(3):1588\u20131623,2014.[15]IanGoodfellow,JeanPouget-Abadie,MehdiMirza,BingXu,DavidWarde-Farley,SherjilOzair,AaronCourville,andYoshuaBengio.Generativeadversarialnets.InAdvancesinneuralinformationprocessingsystems,pages2672\u20132680,2014.[16]ErfanYazdandoostHamedaniandNecdetSerhatAybat.Aprimal-dualalgorithmforgeneralconvex-concavesaddlepointproblems.arXivpreprintarXiv:1803.01401,2018.[17]YunlongHeandRenatoDCMonteiro.Anacceleratedhpe-typealgorithmforaclassofcompositeconvex-concavesaddle-pointproblems.SIAMJournalonOptimization,26(1):29\u201356,2016.[18]ChiJin,PraneethNetrapalli,andMichaelIJordan.Minmaxoptimization:Stablelimitpointsofgradientdescentascentarelocallyoptimal.arXivpreprintarXiv:1902.00618,2019.[19]AnatoliJuditskyandArkadiNemirovski.Firstordermethodsfornonsmoothconvexlarge-scaleoptimization,i:generalpurposemethods.OptimizationforMachineLearning,pages121\u2013148,2011.10\f[20]AnatoliJuditskyandArkadiNemirovski.Firstordermethodsfornonsmoothconvexlarge-scaleoptimization,II:utilizingproblemsstructure.OptimizationforMachineLearning,30(9):149\u2013183,2011.[21]ShamKakade,ShaiShalev-Shwartz,andAmbujTewari.Onthedualityofstrongconvexityandstrongsmoothness:Learningapplicationsandmatrixregularization.UnpublishedManuscript,2009.[22]PurushottamKar,HarikrishnaNarasimhan,andPrateekJain.Surrogatefunctionsformaximiz-ingprecisionatthetop.arXivpreprintarXiv:1505.06813,2015.[23]DavidKinderlehrerandGuidoStampacchia.Anintroductiontovariationalinequalitiesandtheirapplications,volume31.Siam,1980.[24]HidetoshiKomiya.Elementaryproofforsion\u2019sminimaxtheorem.Kodaimathematicaljournal,11(1):5\u20137,1988.[25]JunpeiKomiyama,AkikoTakeda,JunyaHonda,andHajimeShimao.Nonconvexoptimizationforregressionwithfairnessconstraints.InICML,pages2742\u20132751,2018.[26]WeiweiKongandRenatoDCMonteiro.Anacceleratedinexactproximalpointmethodforsolvingnonconvex-concavemin-maxproblems.arXivpreprintarXiv:1905.13433,2019.[27]AYaKruger.Onfr\u00e9chetsubdifferentials.JournalofMathematicalSciences,116(3):3325\u20133358,2003.[28]SongtaoLu,IoannisTsaknakis,MingyiHong,andYongxinChen.Hybridblocksuccessiveapproximationforone-sidednon-convexmin-maxproblems:Algorithmsandapplications.arXivpreprintarXiv:1902.08294,2019.[29]AleksanderMadry,AleksandarMakelov,LudwigSchmidt,DimitrisTsipras,andAdrianVladu.Towardsdeeplearningmodelsresistanttoadversarialattacks.arXivpreprintarXiv:1706.06083,2017.[30]AryanMokhtari,AsumanOzdaglar,andSarathPattathil.Auni\ufb01edanalysisofextra-gradientandoptimisticgradientmethodsforsaddlepointproblems:Proximalpointapproach.arXivpreprintarXiv:1901.08511,2019.[31]RogerBMyerson.Gametheory.Harvarduniversitypress,2013.[32]AngeliaNedi\u00b4candAsumanOzdaglar.Subgradientmethodsforsaddle-pointproblems.Journalofoptimizationtheoryandapplications,142(1):205\u2013228,2009.[33]ArkadiNemirovski.Ef\ufb01cientmethodsforsolvingvariationalinequalities.EkonomikaiMatem.Metody,17:344\u2013359,1981.[34]ArkadiNemirovski.Prox-methodwithrateofconvergenceO(1/t)forvariationalinequali-tieswithlipschitzcontinuousmonotoneoperatorsandsmoothconvex-concavesaddlepointproblems.SIAMJournalonOptimization,15(1):229\u2013251,2004.[35]YuNesterov.Excessivegaptechniqueinnonsmoothconvexminimization.SIAMJournalonOptimization,16(1):235\u2013249,2005.[36]YuNesterov.Smoothminimizationofnon-smoothfunctions.Mathematicalprogramming,103(1):127\u2013152,2005.[37]YuriiNesterov.Introductorylecturesonconvexprogrammingvolumei:Basiccourse.1998.[38]YuriiENesterov.AmethodforsolvingtheconvexprogrammingproblemwithconvergencerateO(1/k2).InDokl.akad.naukSssr,volume269,pages543\u2013547,1983.[39]MaherNouiehed,JasonDLee,andMeisamRazaviyayn.Convergencetosecond-orderstation-arityforconstrainednon-convexoptimization.arXivpreprintarXiv:1810.02024,2018.11\f[40]MaherNouiehed,MaziarSanjabi,JasonDLee,andMeisamRazaviyayn.Solvingaclassofnon-convexmin-maxgamesusingiterative\ufb01rstordermethods.arXivpreprintarXiv:1902.08297,2019.[41]YuyuanOuyangandYangyangXu.Lowercomplexityboundsof\ufb01rst-ordermethodsforconvex-concavebilinearsaddle-pointproblems.arXivpreprintarXiv:1808.02901,2018.[42]HassanRa\ufb01que,MingruiLiu,QihangLin,andTianbaoYang.Non-convexmin-maxop-timization:Provablealgorithmsandapplicationsinmachinelearning.arXivpreprintarXiv:1810.02060,2018.[43]MaziarSanjabi,JimmyBa,MeisamRazaviyayn,andJasonDLee.Ontheconvergenceandrobustnessoftrainingganswithregularizedoptimaltransport.InAdvancesinNeuralInformationProcessingSystems,pages7091\u20137101,2018.[44]MauriceSion.Ongeneralminimaxtheorems.Paci\ufb01cJournalofmathematics,8(1):171\u2013176,1958.[45]SuvritSra,SebastianNowozin,andStephenJWright.Optimizationformachinelearning.MitPress,2012.[46]PaulTseng.Onlinearconvergenceofiterativemethodsforthevariationalinequalityproblem.JournalofComputationalandAppliedMathematics,60(1-2):237\u2013252,1995.[47]PaulTseng.Onacceleratedproximalgradientmethodsforconvex-concaveoptimization.2008.[48]ZhipengXieandJianwenShi.Acceleratedprimaldualmethodforaclassofsaddlepointproblemwithstronglyconvexcomponent.arXivpreprintarXiv:1906.07691,2019.[49]YangyangXu.Iterationcomplexityofinexactaugmentedlagrangianmethodsforconstrainedconvexprogramming.arXivpreprintarXiv:1711.05812,2017.[50]YangyangXuandShuzhongZhang.Acceleratedprimal\u2013dualproximalblockcoordinateupdatingmethodsforconstrainedconvexoptimization.ComputationalOptimizationandApplications,70(1):91\u2013128,2018.[51]RenboZhao.Optimalalgorithmsforstochasticthree-compositeconvex-concavesaddlepointproblems.arXivpreprintarXiv:1903.01687,2019.12\f", "award": [], "sourceid": 6893, "authors": [{"given_name": "Kiran", "family_name": "Thekumparampil", "institution": "Univ. of Illinois at Urbana-Champaign"}, {"given_name": "Prateek", "family_name": "Jain", "institution": "Microsoft Research"}, {"given_name": "Praneeth", "family_name": "Netrapalli", "institution": "Microsoft Research"}, {"given_name": "Sewoong", "family_name": "Oh", "institution": "University of Washington"}]}