Lucas Franco Growth Systems Weekly

Playbook

How to Test a Growth Agent Before Giving It Control

Build a bounded growth agent pilot with explicit metrics, decision history, and a fair comparison against manual analysis before granting execution access.

Format
Playbook
Question answered
Help growth teams design and evaluate an AI diagnostic agent before allowing it to execute changes.
Updated
Direct answer

Start with one recurring growth review and give the agent read-only access to the evidence needed for that decision. Define its objective, guardrails, decision history, and proposal format, then compare its recommendations with manual analysis using the same observation window. Expand its permissions only when the pilot demonstrates useful decisions at an acceptable total cost; evaluate business impact separately through experiments.

An agent that can read a dashboard and suggest a change is easy to imagine. The harder question is whether its suggestion deserves anyone’s time—and eventually permission to act.

A useful starting point comes from Madhu Guru’s discussion of self-improving products: agents need explicit metrics, strategic context, prior decisions, access to operational systems, and a workflow connecting observation to action. That is a useful architecture. It does not establish that adding an agent improves growth.

My recommendation is to test the decision process first. Pick one recurring review, keep execution under human control, and measure whether the agent helps the team find better opportunities for the effort involved.

Choose a decision small enough to evaluate

“Improve growth” leaves too much room for plausible recommendations that nobody can assess.

Choose a bounded surface, such as diagnosing where new users abandon one onboarding path. Define the eligible population, the relevant behavior, the observation window, and the person responsible for deciding what happens next.

A useful pilot question might be: “Can the agent identify worthwhile experiments for this onboarding path more efficiently than our current weekly review?” This is an illustrative question, not a reported experiment.

The surface needs accessible evidence and someone who can judge proposal quality. If the underlying events are unreliable, the first useful output may be an instrumentation repair brief. Treat that as a finding rather than forcing an experiment recommendation from weak data.

Write a one-page agent contract

Before connecting systems, write down what the agent is trying to accomplish and what it may do.

Include these elements:

  • Primary outcome: The customer or business result the team wants to improve, with its population and measurement window.
  • Diagnostics: Intermediate behaviors that help explain changes in that outcome.
  • Guardrails: Outcomes the team will protect, such as retention, margin, or customer experience.
  • Strategic context: Current priorities, known constraints, and work already underway.
  • Evidence access: The datasets, dashboards, and experiment records available to the agent.
  • Permitted actions: For the initial pilot, reading evidence and drafting diagnostic briefs.
  • Decision owner: The person who accepts, rejects, or requests more work on each proposal.

For the onboarding example, a defined first-value behavior might be the primary outcome. Step completion could help diagnose friction. A downstream retention measure could reveal whether an apparent improvement persists.

Keep diagnostics separate from objectives. A faster step completion rate can look attractive while leaving customer value unchanged. The contract should make that distinction explicit before the agent starts ranking opportunities.

Give memory a structure—and an expiry check

A folder of old documents gives an agent material to retrieve. It does not tell the agent which conclusions still apply.

Maintain a decision record containing the question, relevant customer segment, evidence window, hypothesis, decision, experiment configuration, observed result, and remaining uncertainty. Include who made the decision and when the conclusion should be reviewed.

Preserve earlier entries and append corrections or reversals. This lets a reviewer reconstruct why the team acted without silently rewriting history.

The crucial retrieval question is: “Under what conditions was this conclusion valid?” A finding from a different audience, product version, or offer may be useful context without being a rule for today.

Otherwise, memory can make an agent confidently repeat an old mistake. A previous decision should carry its evidence and limits with it.

Require proposals that can survive review

Ask the agent to produce a short decision brief for each opportunity. A useful brief contains:

  1. The observed behavior and the evidence supporting it.
  2. A proposed explanation, clearly separated from the observation.
  3. Alternative explanations and missing information.
  4. A bounded intervention and the mechanism by which it could help.
  5. A measurement plan, guardrails, and reasons to stop.

This format exposes the jump from “something changed” to “we should act.” If activation falls, the agent should consider changes in traffic composition or event collection before prescribing a new onboarding screen.

Require traceable references to the evidence available to reviewers. A recommendation that takes longer to verify than to produce may still be useful, but that verification time belongs in the cost calculation.

Run a fair comparison with manual analysis

A four-to-six-week pilot is a reasonable proposed starting window for recurring reviews. It may be too short to establish downstream business impact.

Run the agent alongside the existing review process. Give both workflows the same observation window and comparable evidence access. Keep their proposals separate until submission so one workflow does not borrow from the other.

Before the pilot starts, define a review rubric: evidence accuracy, relevance to the chosen objective, a plausible mechanism, testability, and implementation effort. Hide the proposal’s origin where practical, while recognizing that writing style may still reveal it.

Score proposals individually, then deduplicate overlapping ideas when counting unique opportunities. Record overlap separately; agreement can be informative even when it adds no new experiment.

The proposed primary measure is approved, unique proposals per analyst hour. Count time spent preparing inputs, operating the agent, checking claims, and reviewing outputs. Track setup and maintenance effort separately so recurring efficiency does not conceal an expensive system.

Also record unsupported claims, duplicate recommendations, reviewer disagreement, and proposals that cannot be tested. An approval is a judgment of usefulness, not evidence that the intervention works.

Make the permission decision explicit

At the end of the pilot, choose among stopping, revising the contract, continuing diagnostic use, or allowing a narrow additional capability.

If the agent produces useful recommendations but requires heavy correction, improve its evidence and proposal process before expanding access. If it reliably finds worthwhile opportunities, a sensible next capability might be drafting an experiment configuration for review.

For any later execution permission, specify the allowed surface, action limits, owner, monitoring window, and recovery procedure. Automatic rollback only makes sense when the change is reversible and the failure signal is reliable. Delayed outcomes cannot protect customers in real time.

Keep two evaluations separate: whether the agent improves the team’s decision process, and whether the resulting interventions create incremental business value. Proposal throughput can improve while launched experiments fail. A short-term metric can rise while retention or margin deteriorates.

Evidence and limitations

This playbook builds on Madhu Guru’s public discussion of the components needed around a self-improving product agent. The available source account used third-party metadata because official embed information was incomplete.

The source provides an architectural checklist, not controlled results. It does not establish better growth outcomes, acceptable error rates, implementation costs, or the reliability of autonomous execution.

The pilot design, proposal format, and permission progression here are operating recommendations. Their value needs to be tested against the team’s existing process. Sparse feedback, delayed outcomes, weak causal measurement, and changes that are difficult to reverse can all make expanded autonomy a poor tradeoff—even when the diagnostic agent is useful.

Source basis

  • Madhu Guru’s public discussion of metrics, strategic context, decision memory, operational access, and workflows for self-improving product agents.
  • A proposed read-only growth diagnostic pilot comparing agent recommendations with manual analysis; no completed experiment results were supplied.
By Lucas Franco

Growth operator focused on lifecycle, experimentation, and practical systems.

Follow Lucas on X

Growth Systems Weekly is coming soon.