A saved growth post can feel like progress before it changes a single decision. The idea sounds promising. You file it under distribution or onboarding. Next week, another idea takes its place.
The missing step is translating an interesting claim into a decision you can make with your own evidence.
Billion Dollar Tweets offers a useful starting point. Its observed structure is a hand-picked archive of social posts and articles organized into categories such as Distribution, Career, Startups, and AI. That makes ideas easier to find again. It does not establish that the ideas work, or that the archive improves business outcomes.
The workflow below is a proposal inspired by that structure. It adds evidence checks and experiment design after discovery. Its usefulness still needs to be tested.
Start with a decision that is already live
Before opening an archive, finish this sentence:
“We need to decide whether to change ___ because we are seeing ___.”
The first blank names an action. The second names evidence from your business.
For example, a team might be deciding whether to change its onboarding sequence because new accounts repeatedly stop before completing an important setup step. That gives the research session a job: find plausible ways to address that particular obstacle.
Without a decision, almost any clever post can look relevant. With one, you can reject an interesting idea simply because it addresses a different problem.
Start with a small batch—up to five relevant entries is a reasonable working limit. This is a suggested constraint on research time, not a proven optimum.
Capture the claim and its context
Evaluate individual entries. A category called Distribution can contain ideas for different audiences, products, and stages of growth. The category alone tells you little about fit.
For each candidate, record:
- Claim: What does the author say happened?
- Proposed mechanism: Why might the action have produced that result?
- Context: What audience, product, channel, and starting conditions were involved?
- Evidence: What observations support the claim, and what is missing?
- Provenance: Who made the original claim, and where can the full account be found?
Keep the reported result separate from the author’s explanation. Someone may accurately report a rise in conversions while incorrectly attributing it to a particular change.
If the original account is unavailable, preserve that uncertainty. A short summary can generate a hypothesis, but it cannot supply the missing baseline or comparison group.
Assess four dimensions without hiding uncertainty
Use four questions to make your reasoning explicit:
- Relevance: Does this address the decision we named?
- Evidence: Can we inspect enough context to understand what was observed?
- Novelty: Does it add a mechanism or option we have not already considered?
- Testability: Can we isolate a change and observe a meaningful result?
Mark each dimension low, medium, or high, with a sentence explaining why. Avoid adding the ratings into one total. High novelty should not cancel out weak evidence.
A relevant, testable idea with weak external evidence may still justify a small, reversible experiment. It gives you a reason to investigate. It gives you much less reason to make an expensive commitment.
An idea with stronger evidence may still be unusable because the original conditions differ too much from yours. A tactic built around an established audience may not transfer to a team with little existing reach.
Write the smallest useful experiment brief
Once an idea survives those questions, describe the test before adding it to the backlog.
A useful brief contains:
- Decision: What will the result help us choose?
- Hypothesis: Which change should affect which behavior, and why?
- Eligible audience: Who actually experiences the problem?
- Comparison: What will we compare the change against?
- Primary outcome: Which behavior represents progress?
- Guardrail: What must not deteriorate?
- Decision rule: What would support continuing, stopping, or investigating further?
For a hypothetical onboarding test, the hypothesis might be that showing a worked example at a confusing setup step increases completion. The primary outcome could be completion among eligible accounts, with later activation as a guardrail against encouraging empty setup activity.
That example illustrates the translation process. It is not a reported result from Billion Dollar Tweets.
Choose the comparison and observation window around your traffic and the time users need to act. Where a credible comparison is impractical, label the work exploratory. A change followed by improvement does not, by itself, establish causation.
Test whether the research process earns its time
A more organized backlog is an output. The operating question is whether the process helps the team make useful decisions.
One proposed pilot is to assign newly saved posts randomly to either the structured process above or the team’s usual approach. Use the same source pool and comparable research time. Evaluate the resulting briefs against criteria defined beforehand, with the evaluator unaware of which process produced each brief where practical.
Track the share of posts that produce a falsifiable, decision-relevant brief. Also track time per post, duplicate ideas, unsupported assumptions, and whether the resulting tests actually launch.
A four-week pilot could reveal practical friction. It may be too short to judge downstream business value. Keep following launched tests to see whether their results change decisions; do not treat a higher brief count as proof of better growth.
Watch for predictable failure modes
Borrowed certainty. An impressive outcome becomes a forecast for your business. Preserve the original context and write your expected mechanism as a hypothesis.
Selection bias. A curated collection can make successful tactics look more dependable than they are. Look for failed applications and conditions under which the mechanism would break.
Duplicate ideas in new language. Several posts may describe the same underlying intervention. Group by mechanism before creating separate tests.
Research that outruns execution. If briefs accumulate while few tests launch, inspect the execution constraint before increasing research volume.
Polished briefs winning the evaluation. A structured template can make weak ideas look persuasive. Judge whether the evidence, comparison, and decision are usable, not whether every field is filled.
Evidence and limitations
Billion Dollar Tweets supplies the underlying example of a categorized, hand-picked idea archive. Its observed organization supports a lesson about discovery and retrieval. Short summaries and titles provide limited grounds for evaluating the causal claims behind individual entries.
The assessment rubric, experiment template, and comparison pilot in this article are proposed operating practices. There is no demonstrated result here showing that they outperform ordinary bookmarking or improve revenue, retention, or acquisition efficiency.
Use the workflow to make assumptions visible and select manageable tests. Its value should be judged by the decisions those tests help you make.
Source basis
- The observed categorized, hand-picked archive structure of Billion Dollar Tweets.
- Evidence limitations of growth claims presented through short social-content summaries.
- An original, unvalidated workflow for assessing curated ideas and converting them into experiments.