Lucas Franco Growth Systems Weekly

System

AI Creative Tagging: Build an Ad Library You Can Learn From

Build a retrospective AI creative tagging system with stable IDs, auditable labels, and experiments that separate creative patterns from performance claims.

Format
System
Question answered
How to implement and validate AI creative tagging for useful ad analysis without mistaking spend patterns for creative effectiveness.
Updated
Direct answer

Keep operational ad names simple and store richer AI-generated creative tags separately, linked through stable creative and ad IDs. Validate the tags against human labels before using them in analysis, and version the taxonomy so historical assets can be reprocessed. Use delivery and outcome data to generate hypotheses, then test promising patterns prospectively before treating them as creative advantages.

An ad name has limited room. You can pack it with format, market, offer, hook, audience, creator, and launch date—and still discover next month that the question you need to answer depends on an attribute nobody recorded.

Retrospective AI tagging offers a useful way out: preserve the identifiers needed to run the account, then describe the creative in a separate, updateable layer.

Curtis Howland proposes this hybrid approach in a LinkedIn post about ad naming conventions. The useful idea is that teams can enrich historical creatives as their questions change. My recommendation is to build that capability around two checks: whether the labels are trustworthy and whether the analysis supports the decision being made.

Start with a question worth answering

Before choosing a model, write down the creative decisions this system should improve.

Useful starting questions include:

  • Which concepts have we tested repeatedly through cosmetic variations?
  • Which formats or languages have received little exposure?
  • Which creative patterns deserve a controlled follow-up test?

These questions require different evidence. Finding every video with a product demonstration is a retrieval task. Deciding that demonstrations improve acquisition efficiency is an effectiveness question. Accurate tagging can help with both, but it cannot establish the second answer by itself.

Start with a small set of recurring questions. A taxonomy earns its maintenance cost when it helps someone answer them.

Separate identity, description, and delivery

I would organize the system into three connected records.

The asset record identifies the exact creative version. Preserve a stable creative ID and distinguish materially different edits. If the opening scene or offer changes, the analysis needs to recognize that change.

The tag record describes what appears in that asset. Store each accepted label alongside its taxonomy version, model or processing version, processing date, and review status. Preserve uncertainty rather than forcing every asset into a category.

The delivery record connects the creative to the ads that served it. Keep ad IDs, reporting periods, spend, impressions, outcomes, and relevant delivery context separate from the labels.

That separation matters when one creative runs in several ads. Reusing an asset should not make it look like several independent creative concepts. Conversely, combining all delivery into one lifetime total can hide differences in audience, placement, objective, or creative age.

Keep the operational naming convention your team needs. The enrichment layer should reduce pressure to encode every future analytical question into the name.

Make the first taxonomy easy to audit

Begin with observable attributes: format, duration, language, product visibility, creator presence, transcript, and on-screen text.

Interpretive tags such as hook type, customer problem, and funnel stage need more care. Two analysts can watch the same ad and disagree without either making an obvious mistake.

For each categorical field, define:

  • What the field means and which values are allowed.
  • What evidence qualifies an asset for each value.
  • Whether multiple values are permitted.
  • When to return unknown or send the asset for review.

For example, “product demonstration” needs a boundary. Does holding the product count, or must the viewer see it being used? Choose a definition that serves the question, then apply it consistently.

Use non-identifying descriptions such as “creator present” when identity is unnecessary. More detailed labeling creates additional review work; collect it only when it changes a decision.

Validate labels before building the dashboard

Select a historical sample spanning the formats, languages, and production styles the system will encounter. Have people label it using the proposed definitions, and inspect disagreements before treating those labels as a reference set.

Human disagreement is useful feedback. A field that nobody can apply consistently needs a clearer definition before model tuning will help.

For categorical tags, evaluate precision and recall separately. Precision asks how often an assigned tag is correct. Recall asks how many qualifying assets the system finds. Transcripts and extracted text need their own checks for accuracy and completeness.

Set acceptance criteria according to the decision. A browsing filter may tolerate errors that would be unacceptable in an analysis used to allocate production budget. Model confidence alone is insufficient; check whether confidence corresponds to correctness on your sample.

Keep a separate evaluation sample out of taxonomy and prompt adjustments. Otherwise, repeated improvements may only teach the workflow to handle examples already examined.

When the taxonomy changes, retain previous versions. A changed definition should not silently rewrite the meaning of a historical comparison.

Treat spend as exposure evidence

Once tags pass validation, join them to delivery and outcomes. Build a coverage view showing spend, impressions, attributed outcomes, distinct assets, and distinct concepts by relevant tag combination.

This can reveal where the account has concentrated its exposure. It can also identify questions worth investigating. It does not establish why an attribute received more spend or whether that attribute caused better results.

Budget structure, creative age, audience size, placement, and platform allocation can all affect the pattern. A low-spend category may be underexplored, weak, newly launched, or constrained by the campaign setup.

Semantic similarity can help group near-duplicates, but inspect the groups. Similar-looking assets may test different promises, while visually different assets may repeat the same proposition. Count concepts deliberately before declaring the portfolio diverse.

Apply the same restraint to competitor libraries. Visible creative activity can inspire a brief; it does not establish competitor profitability or spend.

Prove workflow value, then test creative value

The first pilot should answer whether the system improves analysis.

Give analysts the same defined questions using the existing workflow and the tagged library, varying the order across participants. Measure time per correctly answered question and record where errors occur. Account for review effort and processing cost alongside time saved.

The next step is a prospective creative test. Take one promising pattern, state a falsifiable hypothesis, and design a comparison that keeps audience, placement, offer, and landing experience as comparable as possible. Randomize where feasible and choose the outcome and decision rule in advance.

A creative comparison can test relative performance within that setup. Claims about incremental business value need an appropriate incrementality design. Neither follows automatically from a tag-performance correlation.

Keep the pilot small enough to inspect. Expand when the labels are dependable and the answers improve actual decisions. A searchable library can be useful even before it produces a winning creative hypothesis.

Evidence and limitations

This system builds on Curtis Howland’s proposal to combine essential ad naming with retrospective AI enrichment. The record structure, validation sequence, and pilot guidance here are operator recommendations derived from that idea.

The supplied evidence contains no validated tag-accuracy results, implementation costs, measured time savings, or demonstrated acquisition improvements. It supports testing the workflow, not promising a return. Historical creative patterns remain hypotheses until suitable prospective evidence supports stronger conclusions.

Source basis

  • Curtis Howland’s public LinkedIn proposal for minimal operational ad naming and retrospective AI creative tagging.
  • Critical analysis of tag validation, creative coverage, delivery confounding, and prospective testing; no measured performance results were supplied.
By Lucas Franco

Growth operator focused on lifecycle, experimentation, and practical systems.

Follow Lucas on X

Growth Systems Weekly is coming soon.