Lucas Franco Growth Systems Weekly

Playbook

Build an AI Creative Workflow That Produces Useful Tests

Build an AI-assisted creative workflow that connects customer evidence to testable hypotheses, tracks production effort, and checks acquisition economics.

Format
Playbook
Question answered
How to build and evaluate an AI-assisted workflow that turns customer evidence into useful creative tests.
Updated
Direct answer

Start with one customer dataset, one channel, and a clear definition of a useful creative test. Connect each brief to source evidence, generate variants around an explicit hypothesis, and keep launch and budget decisions with a named operator. Evaluate completed tests and total human effort alongside downstream conversion costs; asset volume alone does not show that the workflow helps.

Generating more creative is easy to count. Whether that creative helps you make a better decision is harder to establish.

An AI workflow can produce dozens of assets while leaving the team with the same unanswered question: which customer problem should we lead with, and why?

The useful unit of output is a completed test with a traceable hypothesis, enough delivery to interpret, and a readout that informs a decision. Build the workflow around that unit.

GREG ISENBERG’s public post on the “marketing engineer” describes operators building AI agents for customer research, content, and creative testing. That is a useful prompt for system design. The practical question is whether a bounded workflow can increase useful testing capacity without adding hidden work or weakening acquisition economics.

Here is a proposed way to find out.

Start with a narrow operating contract

Choose one channel and one bounded set of customer evidence. For example, use a permitted collection of product reviews to develop briefs for a single paid campaign. This is an illustrative starting point, not a reported implementation.

Specify what the workflow receives, what it produces, and who owns the next decision.

Its first deliverable should be a brief containing a customer problem, supporting evidence, a creative hypothesis, and a proposed test. A marketer checks that brief before it enters production. A named campaign owner decides what launches and how much delivery it receives.

Also define the exception path. If evidence is missing, contradictory, or too sensitive to use, the workflow should flag the gap. Filling the gap with a plausible customer story would make the brief look complete while making the test less trustworthy.

Preserve the connection between evidence and interpretation

Customer language becomes useful when you can inspect its context.

For each candidate theme, retain a source reference, the relevant excerpt, the situation it describes, and any contradictory evidence. Remove identifying information that the creative team does not need.

Keep three things separate:

  • Observation: What someone actually said or did.
  • Interpretation: What that may suggest about a problem, objection, or desired outcome.
  • Recommendation: What message the team should test because of that interpretation.

Suppose several reviews describe a setup process as confusing. The observation concerns setup friction. The interpretation might be that uncertainty delays adoption. A proposed creative angle could demonstrate the first successful action.

That evidence does not establish that setup is fast, effortless, or universally easy. Those claims need their own support.

Ask the model to preserve disagreements as well as recurring themes. A polished summary that erases differences between new users and experienced customers can send the team toward the wrong audience or promise.

Turn each theme into a falsifiable brief

Before generating assets, write the decision the test is meant to inform.

A useful hypothesis might be: “For prospects worried about setup, showing the first completed task will produce more qualified conversions than leading with a broad feature list.”

That is an illustrative hypothesis. It still needs evidence that the concern exists and a product demonstration that accurately represents the experience.

Each brief should contain:

  • The audience and situation being addressed.
  • The evidence behind the proposed angle.
  • The message or demonstration being changed.
  • The comparison creative.
  • The primary outcome and business guardrail.
  • The conditions under which the result will be considered inconclusive.

Keep the initial variants few enough that each can receive useful delivery. Several rewrites of one promise should remain grouped under the same concept. Otherwise, the system can inflate its apparent testing output through cosmetic variation.

Plan delivery before producing variants

Production capacity and testing capacity are different constraints.

If the account can only support a small number of interpretable comparisons, generating a large batch creates a queue. It may also spread delivery so thinly that few tests answer anything.

Have the campaign owner define the delivery plan and decision rules before launch. The workflow can check whether a brief contains these fields and flag deviations for investigation.

Avoid a universal click-through-rate rule for declaring winners or stopping ads. Clicks can help diagnose attention and message response, but they do not establish profitable acquisition. A creative may attract curiosity that never becomes customer value.

Record changes to offers, targeting, placements, or measurement during the test. If those changes undermine the comparison, mark the result inconclusive instead of forcing a winner.

Compare the workflow with the existing process

The first evaluation should ask whether the new production process helps the team complete better work.

Consider a four-to-six-week pilot as a planning proposal, not a guaranteed measurement window. Actual duration depends on brief volume, delivery, and conversion lag.

Where practical, randomly assign comparable briefs to the existing process or the AI-assisted process. Keep the channel, offer, budget rules, and quality standards comparable. Avoid assigning only easy briefs to the new workflow.

This process comparison is separate from the creative comparisons run inside campaigns. It estimates whether the workflow changes production and testing capacity; it does not, by itself, establish that advertising caused additional sales.

Count completed tests only when they have a documented hypothesis, meet the predefined delivery requirements, and produce a usable readout. An inconclusive result can still be useful when it clearly explains what the team cannot conclude and why.

Track total human effort, including evidence cleanup, editing, checking, rework, and reporting. Time saved during drafting can disappear during correction.

Alongside throughput, monitor downstream conversion cost, conversion quality, rejected creative, repeated concepts, and relevant margin measures. If you only have an attributed conversion-cost measure, label it that way. Do not call it incremental CAC.

Make the readout change the next brief

The feedback loop needs more detail than “this ad won.”

Record what was tested, where it ran, which audience saw it, whether delivery was adequate, and what decision followed. Keep the original hypothesis attached to the result.

A result from one audience and offer should remain scoped to that context. Before carrying the lesson elsewhere, identify what would need to stay true for it to transfer.

Use corrections to improve the workflow too. Repeated unsupported claims suggest a problem with evidence handling. Repeated concepts suggest weak differentiation between hypotheses. Heavy editing may mean the brief specification is incomplete.

Expand the workflow only when it produces more usable tests at acceptable total effort and business performance. If output rises while useful completions stay flat, investigate the constraint before generating more.

Evidence and limitations

This playbook draws on GREG ISENBERG’s public discussion of marketing engineers and AI agents for growth work. The available source describes workflow possibilities; it does not provide controlled results demonstrating better creative performance or acquisition economics.

The operating contract, evidence structure, and evaluation guidance here are recommendations. The examples are illustrative, and the proposed pilot has no reported results.

A successful process pilot would support a narrower conclusion: the workflow helped a particular team produce and evaluate creative under defined conditions. Claims about incremental business impact require a measurement design that can support them.

Source basis

  • GREG ISENBERG’s public X post describing marketing engineers and AI-assisted growth workflows, as summarized in the supplied source material.
  • Editorial analysis of the proposed creative-testing mechanism and an untested workflow evaluation design; no demonstrated performance outcomes.
By Lucas Franco

Growth operator focused on lifecycle, experimentation, and practical systems.

Follow Lucas on X

Growth Systems Weekly is coming soon.