Risk and Optimal Policies in Bandit Experiments.

We provide a decision‐theoretic analysis of bandit experiments under local asymptotics. Working within the framework of diffusion processes, we define suitable notions of asymptotic Bayes and minimax risk for these experiments. For normally distributed rewards, the minimal Bayes risk can be characte...

Descripción completa

Detalles Bibliográficos
Publicado en:Econometrica Vol. 93; no. 3; pp. 1003 - 1030
Autor principal: Adusumilli, Karun
Formato: Artículo
Publicado: Wiley-Blackwell May2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=185839955&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 185839955
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00129682
        ECN
      jtl: Econometrica
      issn: 00129682
      maglogo: Y
    pubinfo:
      dt: May2025
      vid: 93
      iid: 3
      pid: 480
      pub: Wiley-Blackwell
    artinfo:
      ui:
        185839955
        10.3982/ECTA21075
      ppf: 1003
      ppct: 27
      formats:
      tig:
        atl: Risk and Optimal Policies in Bandit Experiments.
      aug:
        au: Adusumilli, Karun
        affil: Department of Economics, University of Pennsylvania
      su:
        Decision making
        Decision theory
        Multi-armed bandit problem (Probability theory)
        Bayes' theorem
        Partial differential equations
        Feature selection
        Asymptotic expansions
      sug:
        subj:
          Decision making
          Decision theory
          Multi-armed bandit problem (Probability theory)
          Bayes' theorem
          Partial differential equations
          Feature selection
          Asymptotic expansions
      keyword:
        Bayes and minimax regret
        diffusion processes
        Multi‐armed bandits
        statistical decision theory
        Bayes and minimax regret
        diffusion processes
        Multi‐armed bandits
        statistical decision theory
      ab: We provide a decision‐theoretic analysis of bandit experiments under local asymptotics. Working within the framework of diffusion processes, we define suitable notions of asymptotic Bayes and minimax risk for these experiments. For normally distributed rewards, the minimal Bayes risk can be characterized as the solution to a second‐order partial differential equation (PDE). Using a limit of experiments approach, we show that this PDE characterization also holds asymptotically under both parametric and non‐parametric distributions of the rewards. The approach further describes the state variables it is asymptotically sufficient to restrict attention to, and thereby suggests a practical strategy for dimension reduction. The PDEs characterizing minimal Bayes risk can be solved efficiently using sparse matrix routines or Monte Carlo methods. We derive the optimal Bayes and minimax policies from their numerical solutions. These optimal policies substantially dominate existing methods such as Thompson sampling; the risk of the latter is often twice as high.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N