A generated interface can look convincing before it is useful for an experiment. The first screen may appear quickly, while the work of preserving state, fixing interactions, and making behavior measurable remains unresolved.
The useful question for a growth team is specific: does this workflow reduce the time required to produce a prototype that meets the brief and lets someone complete the intended task?
Runway’s announcement of Solaris provides a reason to investigate that question. Runway describes a system that generates interactive interfaces frame by frame in real time. The announcement alone does not establish that those interfaces improve conversion, support dependable analytics, or work reliably in production.
My recommendation is to evaluate the production workflow first. If it earns a place there, evaluate its suitability for a customer experiment separately.
Choose a task with a clear finish line
Start with one bounded landing-page or onboarding flow. Choose something representative of work your team actually needs, with an interaction that can be assessed from beginning to end.
For example, a hypothetical onboarding brief might ask a user to choose a goal, provide one required input, and reach a confirmation screen that preserves their selection. That is specific enough to reveal whether the interface works beyond its opening frame.
Write down the audience, promise, required content, allowed interactions, completion state, and expected behavior when an input is missing or changed. Include brand and accessibility requirements that matter for this use case.
Define acceptance before either workflow starts. Otherwise, it is easy to accept a visually impressive generated version with fewer requirements than the conventional version.
The operating hypothesis should remain narrow: the generated-interface workflow can reduce prototype lead time without an unacceptable loss of usability or fidelity to the brief. Conversion improvement is a later hypothesis.
Give both workflows the same assignment
Compare the generated-interface approach with the process your team would otherwise use. Supply the same brief, assets, content, and acceptance criteria to each.
Be explicit about the endpoint. A prototype suitable for moderated research may only need to support a defined task under observation. A prototype intended for a live experiment has additional requirements. Comparing one of each would tell you very little.
Record meaningful differences in operator experience and starting materials. A familiar workflow with reusable components has an advantage; so does a generated workflow that receives extensive coaching from its vendor. Those differences belong in the interpretation.
Where practical, repeat the comparison across several matched briefs. One attractive result can reveal a possibility, but it is a weak basis for replacing a workflow. If you can afford only one comparison, keep the decision limited to that kind of task.
Measure the path to acceptance
Use elapsed time from the agreed starting point to an accepted interactive prototype as the primary measure. Across several comparisons, report the median alongside the individual results so difficult briefs remain visible.
Also record hands-on time. Elapsed time tells you how long the team waits; hands-on time tells you how much labor the workflow consumes. A tool could improve one while making the other worse.
Keep a simple work log covering initial production, review, revisions, and rechecking. Record failed attempts and prototypes that never meet the criteria. Do not calculate an apparent speed advantage using only successful generations.
The most revealing interval may be the gap between the first plausible output and acceptance. That is where broken interactions, inconsistent state, and repeated corrections can absorb the initial time saving.
Separate initial setup from recurring production work. Both matter, but they answer different questions about whether the workflow is worth adopting.
Review the task before revealing the method
Have reviewers attempt the same task in both prototypes. Where possible, hide how each was made and vary the order in which reviewers encounter them.
Ask reviewers to complete the task, then inspect what happened. A preference vote about which interface looks more polished cannot establish whether either flow works.
Use a shared review sheet:
- Can the reviewer reach the intended completion state?
- Does the interface preserve required choices and inputs?
- Does it meet the content and behavior requirements?
- Can the reviewer recover from a mistake?
- Do repeated attempts produce consistent behavior?
- What accessibility barriers appear during the task?
Set acceptable quality limits before inspecting the results. Small reviews can expose obvious problems and guide revisions, but they cannot establish broad equivalence between the workflows.
Make two adoption decisions
The first decision is whether the workflow is useful for concept validation. A generated interface might help a team explore language, sequence, or interaction ideas with research participants even when it cannot support production deployment.
The second decision is whether it can support a controlled customer experiment. Before taking that step, verify that the experience can preserve state, expose the required analytics events, and provide a dependable fallback when generation fails. Check accessibility and security against the actual intended use.
For measurement, ask whether you can identify which experience a participant received and reproduce its important behavior. If the interface changes during a session, assignment to a nominal variant may not adequately describe the treatment. Resolve that measurement problem before attributing an outcome to the design.
A workflow that succeeds in research can remain valuable at that scope. There is no need to turn every useful prototype tool into a production delivery system.
Watch for misleading wins
The common failure is stopping the clock at the first impressive screen. Another is comparing unequal briefs or quietly relaxing requirements for the new approach.
A subtler mistake is declaring success because more prototypes were produced. Extra variants create review and prioritization work. They help only when the team can use them to resolve meaningful uncertainty.
Adopt the workflow for tasks where it reaches acceptance faster at an acceptable quality level. If revisions consume the advantage, narrow the use case. If the experience cannot be measured reliably, keep its role in concept exploration until that constraint is resolved.
Evidence and limitations
This playbook draws on Runway’s public Solaris announcement and a proposed comparison of prototype workflows. Runway’s capability and benchmark statements are vendor claims. The supplied evidence does not include benchmark methods, independent evaluations, production results, or an analyzed demonstration.
The evaluation procedure here is an operator recommendation, not a report of a completed experiment. No time savings, usability equivalence, or conversion gains have been established. A successful prototype comparison would support a decision about that workflow and task; it would still leave customer impact to be tested.
Source basis
- Runway’s public announcement describing Solaris as a system for generating interactive interfaces in real time.
- A proposed matched comparison of prototype lead time, usability, specification fidelity, and experiment readiness; no completed results were supplied.