
Autonomous vehicle companies operate under constraints that make standard A/B testing difficult – you cannot freely test safety-adjacent messaging, and your user base is confined to specific geofences. But those same constraints create a structure for rigorous experimentation that most consumer growth teams would envy. Winston Francois helps AV companies design experiments that are ethical, operationally feasible, and actually answer the questions that move the business.
Safety-Adjacent Messaging Cannot Be Freely A/B Tested
Standard growth experimentation assumes you can test any message variation on any segment of your audience. In AV, that assumption breaks down quickly. You cannot run a test where one rider cohort sees 'our vehicles have never had a serious incident' and another sees a more measured safety claim – the ethical and regulatory implications of those tests are not equivalent to testing button colors. Most AV growth teams know they cannot test safety messaging freely but have not built a principled framework for what they can test and how. The result is either no experimentation or experiments that legal and safety teams later require to be pulled.
Geo-Limited User Bases Make Statistical Significance Hard to Reach
A robotaxi operating in a single metro geofence might have tens of thousands of active riders, not millions. That user base is enough to run experiments, but not the kind of experiments a growth team trained on consumer internet products expects to run. Effect sizes need to be larger to detect, sample sizes need to be managed carefully, and the temptation to call experiments early is significant when you have limited runway to wait for results. Growth experimentation in AV requires statistical discipline that many teams lack because they have not worked at this scale before.
Operational Variables Contaminate Experiment Results
In AV, the product experience is partly a function of operational decisions that change week to week – fleet size, route coverage, dispatch algorithm updates, vehicle maintenance cycles. A rider activation experiment running in October might produce different results than the same experiment in December not because the messaging changed but because operational factors shifted. Controlling for those variables when designing experiments, or accounting for them when interpreting results, requires a level of cross-functional coordination that most growth teams do not have a system for.
No Experiment Culture Means No Learning Culture
Many AV companies have the instinct to test but not the infrastructure or the process to run experiments systematically. Individual team members run ad hoc tests that are not documented, reviewed, or built upon. Results that would inform future decisions disappear when the person who ran the test moves on. A growth experimentation program is not a set of individual tests – it is a learning system with documentation standards, review processes, and institutional memory. Most AV growth teams have the first but not the second.
Winston Francois starts every AV growth experimentation engagement by establishing what can and cannot be tested – not as a compliance exercise, but as a strategic one. We work with your legal, safety, and communications teams to define the experiment boundaries clearly, so your growth team can operate confidently within them. This is often the most valuable early output because it replaces a vague sense of 'we probably cannot test that' with a documented framework everyone can reference.
From there we design the experimentation infrastructure. For most AV companies at Series A to growth stage, this means a test registry, a standard experiment brief format, a statistical significance calculator calibrated to your actual user base size, and a review process that gets experiments approved or rejected quickly. We build this infrastructure to match your team's actual capacity, not a best-practice ideal that no one will maintain.
With infrastructure in place, we work with your growth team to prioritize the first experiment backlog. AV growth experiments tend to cluster around a few high-leverage areas: rider onboarding sequence, re-engagement messaging for lapsed riders, pricing and promotion structure, in-app trust signals, and developer onboarding for SDK-exposed platforms. We help your team move from a vague list of 'things we should test' to a ranked backlog with clear hypotheses and success criteria.
Experiment execution support means we are available during the run to flag early problems – sample contamination, operational variable shifts, statistical errors in the design – before they invalidate the results. We run results analysis and write the findings brief that captures what was learned and what the next experiment should be.
The program cadence we establish – typically a two-week experiment cycle with a monthly review – is designed to build institutional momentum. The hardest part of growth experimentation is not running the first test; it is running the twentieth test with the same rigor as the first. We build the habits and the documentation standards that make that possible.
The geo-limited nature of AV launches is not an experimentation handicap – it is a controlled environment. A company launching in two cities simultaneously has a natural comparison group that most consumer growth teams can only approximate through matched-market analysis.
Winston Francois structures growth experimentation engagements around a 90-day sprint. The first 30 days establish the foundation: experiment boundary framework, infrastructure design and build, initial backlog prioritization, and team training on the brief format and review process. We do not run the first experiment until the infrastructure is in place – running experiments without a proper test registry and documentation standard is how teams lose their learning history.
Days 31 through 60 are the first active experimentation cycle. We support the design, execution, and analysis of the first two to three experiments. These are deliberately chosen to be tractable – they have clear hypotheses, adequate sample sizes given your user base, and operational conditions that are stable enough to interpret results cleanly. The goal is to build confidence in the process before running more complex experiments.
The final 30 days are program calibration. We review what the first cycle taught us about your team's capacity, your users' responsiveness to interventions, and the operational variables that most affect experiment validity. We revise the experiment backlog based on those learnings and hand off a running program that your team can sustain without us.
Growth experimentation engagements at Winston Francois are structured as 90-day programs with a defined infrastructure build and a running experiment cadence as the output. We do not run one-off tests and call it a program.
In the first 30 days we establish what can be tested, build the infrastructure, and train the team. You get a documented experiment boundary framework, a functioning test registry, and a prioritized backlog before the first test runs.
Days 31 through 60 are the first experiment cycle. We support execution and analysis. We are available when operational variables shift mid-test and when statistical interpretation questions arise. We write the findings brief and update the backlog based on what was learned.
The final 30 days are calibration and handoff. We document the program, train whoever will own it ongoing, and produce a 90-day experiment roadmap your team can execute independently. We schedule a check-in at the 90-day mark after handoff to review how the program is running.
If your autonomous vehicles company needs growth experimentation leadership, we should talk.

Let us take a custom approach to your growth goals by assembling and leading the best-in-class marketing team to support your next stage.
The safe and productive experimentation territory in AV includes: rider onboarding sequence variations, re-engagement messaging for riders who have not taken a trip in 30 or 60 days, in-app trust signal design (how you present safety information, vehicle information, and trip transparency features), pricing and promotion structure for new market launches, and developer onboarding flows for SDK or API platforms. Safety-adjacent factual claims, incident statistics, and performance comparisons to competitors require legal and communications review before any experiment design is finalized.
Geo-limited launches reduce your available sample size but also reduce the contamination risk from external market variables that plague broader consumer experiments. If you are operating in one metro, your experiment results are not confounded by regional differences in consumer behavior or competitor activity.
This is the most common question we get from AV growth teams, and the answer is: carefully, with legal and communications sign-off, and not in the ways consumer product teams typically mean. Testing whether a specific factual description of your safety record drives higher trip frequency than a different factual description is a legitimate experiment if both descriptions are accurate and the test population is told they are receiving different versions of the app.
The answer depends on the metric you are trying to move and the effect size you expect. For a rider activation experiment targeting a behavior that 20 percent of your riders currently exhibit, you can detect a 5 percentage point improvement with roughly 1,500 riders per variant.
The primary tool is documentation and monitoring. Before any experiment runs, we document the current operational state – fleet size, dispatch algorithm version, average wait time, geofence coverage – and set up monitoring that flags if those variables change materially during the test.
At most Series A to Series B AV companies, a dedicated team is premature. A shared responsibility model works if there is a clear owner – typically one senior growth or product manager – who is accountable for the experiment program and has time allocated to maintain the test registry, run the biweekly review, and write the findings briefs. The failure mode is treating experimentation as everyone's responsibility, which in practice means it becomes no one's responsibility. One owner, supported by a clear process, is the minimum viable setup.
Tuesday, July 28, 2026
Frank Growth – Episode 230 – Growth’s Most Dangerous Trap With Sara Wallace
Tuesday, July 21, 2026
Frank Growth – Episode 229 – Longevity Medicine’s Dirty Secret with Jim Donnelly
Tuesday, June 16, 2026
Frank Growth – Episode 224 – The Bootstrapper’s Revenge with Alex Roy
Tuesday, July 14, 2026
Frank Growth – Episode 228 – Your Bookkeeper Is Failing You with John Zdanowski
Ready to unlock your growth?
Book Free Call