Skip to content

Design: evaluation loop — performance data as reward signal #5

Description

@gidutz

Adtech has unusually clean ground truth: ROAS, CTR, fill rate, viewability are measurable within hours. We should exploit this.

Components:

  • Performance data ingest (from ad platforms back into our store)
  • Reward signal definition per agent role (what does "good" mean?)
  • Offline eval harness (replay historical decisions)
  • Online eval / shadow mode (run agent in parallel with a human, compare outcomes)
  • Regression tracking across agent versions

Open questions:

  • Attribution lag — how long do we wait before scoring a decision?
  • Confounders (seasonality, exogenous events, creative fatigue)
  • Per-account vs. global eval

Deliverable: eval harness spec plus initial metric definitions.

Metadata

Metadata

Assignees

No one assigned

    Labels

    pillar:evalsPerformance feedback loop and metricstype:designDesign discussion

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions