Growth Experimentation A/B Testing · Search vs. Browse Test-and-Learn System Pathao Food App

Food Home Experiments
Building a Test-and-Learn Engine for Discovery

How I turned a single "reorder the home screen" request into a prioritized, 12-experiment testing program for Pathao Food's home surface - shifting how millions of users discover restaurants from search-dependent to browse-driven, with a repeatable methodology I could run again on any surface.

12tests
Prioritized experiment backlog I designed and sequenced
3waves
Phased rollout across ~10 weeks of controlled testing
+38%
Relative lift in Collection Conversion Rate, flagship test
30→42%
Home's share of restaurant visits, up from search reliance
01 — Background

A home screen nobody had ever tested.

Pathao is Bangladesh's largest super-app, and Food is one of its highest-traffic verticals, competing daily for order share against foodpanda and other delivery players. Like most food-delivery apps, Food Home is the first screen a hungry user sees - a stack of horizontally scrolling collections (cuisines, deals, trending restaurants, curated picks) meant to help people find something to eat without typing a single letter.

When I pulled the funnel numbers, the picture was uncomfortable. Only 10–12% of App Home users ever reached Food Home in the first place. Of those who did, 70% went on to visit a restaurant page - but of those visits, 70% originated from search and only 30% from browsing the home screen itself. The surface built specifically to help people discover restaurants without searching was losing to the search bar, more than two to one.

The uncomfortable question: every collection, every cuisine tile, every ordering decision on that home screen had been shipped on design opinion - never tested. Nobody actually knew whether the existing layout was helping users or quietly pushing them toward search. I decided to find out, systematically.


02 — Problem

Three problems, and none of them had a single-test answer.

When I broke down why Food Home was underperforming, I didn't find one bug to fix. I found a surface with a dozen untested variables competing for the same real estate - which meant the real problem wasn't "what's the right layout," it was "how do I build a system that can find out."

01
Home browsing was a cold, static surface
Collections appeared in the same order for every user, every session, regardless of what actually converted. A user with no specific restaurant in mind had no signal-driven path to discovery - so the path of least resistance was typing into search. The home screen wasn't failing users; it was failing to earn their trust before search did.
02
Zero experimentation culture on the surface that mattered most
Every previous change to Food Home - collection order, tab styles, imagery - had shipped based on internal design consensus, not evidence. There was no cohort methodology, no shared metrics definition, no standard for what counted as "working." I needed to build the testing infrastructure itself before I could test anything on top of it.
03
A dozen credible hypotheses, one home screen, limited dev cycles
Placement order, cuisine timing, tab design, image treatment, language, restaurant sort logic, seasonal theming - every one of these was a defensible hypothesis for why home underperformed. Running them all at once would confound the results; running them randomly would waste months. I needed a prioritization framework, not just a test plan.

03 — My Approach

Treat this as a program, not a feature.

The brief I was handed was simple: "make the home screen convert better." I deliberately reframed it. Rather than pick one design and A/B test it, I designed an experimentation program - a prioritized backlog of testable hypotheses, a shared methodology every test would follow, and a metrics framework that let me compare results across tests run months apart.

01
Define a North Star and guardrails before writing a single hypothesis
I set Collection Conversion Rate as the North Star - the clearest signal that home browsing itself was working. But a North Star in isolation invites gaming, so I paired it with guardrails: Search Conversion Rate (should not collapse), Overall App Conversion Rate (the real business outcome), and Order Count (the ultimate check against vanity metrics).
02
Build one reusable experiment methodology, not twelve one-off tests
Every test in the backlog - regardless of what it tested - would use the same cohort selection logic, the same control/treatment structure, and the same weekly metrics cadence. This is what made the results comparable across tests and defensible to stakeholders who'd only ever seen design decisions made by opinion.
03
Prioritize by impact-to-effort, not by whoever asked loudest
Twelve credible hypotheses meant twelve competing requests from design, growth, and category teams. I scored each one on traffic exposure, implementation cost, and expected signal strength, and sequenced the backlog into phases. This is the part of the job that isn't visible in a single test result but is the actual product-management work - deciding what NOT to test yet.
04
Design for a decision, not just a data point
Every test was scoped with an explicit conclusion criteria upfront - not "let's see what happens," but "if collection CVR moves by X with statistical confidence, we ship; if not, we kill and reallocate the slot." This kept the backlog moving instead of accumulating inconclusive tests that never got closed out.

04 — Methodology

The same seven-step process, run for every test in the backlog.

I designed one methodology and reused it across the entire program, so results from a test run in week 2 could be trusted and compared against a test run in week 9.

Reusable across every test
01
Cohort selection
A random sample of users who had placed a food order in the past month - active, representative users, not new-user edge cases that would skew early results.
02
Control vs. treatment structure
One control group always saw the existing, unchanged Food Home. One or two treatment groups received the variant(s) under test - I allowed two treatment arms in a single test slot where it made sense (e.g. real food photography vs. animated illustration, tested head-to-head against the same control, in the same window).
03
Duration set by traffic, not by calendar convenience
Each test ran long enough to reach a stable read given Food Home's daily traffic volume - typically 2–3 weeks, long enough to smooth day-of-week effects without letting novelty bias dominate the early days.
04
Metrics tracked identically across every test
Collection Conversion Rate, Search Conversion Rate, Overall App Conversion Rate, and Order Count - tracked weekly for both control and treatment, for every single test, so results could be stacked and compared program-wide.
05
Statistical comparison against the guardrails, not just the headline metric
A win on Collection CVR that tanked Overall App Conversion wasn't a win - it meant collections were cannibalizing higher-intent search traffic without a net gain. Every result was read against all four metrics together.
06
Ship, iterate, or kill - explicitly
No test was left "inconclusive" indefinitely. Every test closed with one of three calls: ship the winning treatment, run a follow-up iteration to sharpen the signal, or kill the hypothesis and free the backlog slot.
07
Feed the decision back into the backlog
A shipped winner became the new control baseline for the next test in that area - so the program compounded. Each wave built on the last instead of testing against a static, outdated baseline.

05 — Prioritization

Twelve hypotheses. One sequenced backlog.

Rather than run whichever test was easiest to build first, I scored the full hypothesis list on expected impact (traffic exposure of the surface being changed, and how directly it connected to conversion) against implementation effort, and sequenced the program into phases. This is the artifact I used to align design, engineering, and category stakeholders on what would get tested when - and why.

TreatmentHypothesisPhaseDurationStatus
Collection Placement Order Reordering collections by engagement-weighted logic instead of static order surfaces the right restaurant faster Phase 0 3 weeks Shipped
Cuisine Placement Dynamicity Cuisine position should shift by hour of day (breakfast items up mornings, biryani/dinner cuisines up evenings) Phase 1 2 weeks Shipped
Collection Tab Design Variation Real food photography outperforms animated/illustrated tab imagery on click-through Phase 2 2 weeks Shipped
Collection Tab View vs. Horizontal View A dedicated collection tab converts better than the same content shown in a horizontal scroll row Phase 3 2 weeks Queued
App Card vs. Horizontal Collection View Full app-card presentation of a collection outperforms the equivalent horizontal scroll layout Phase 3 2 weeks Backlog
Restaurant Order Within Collection Sorting restaurants inside a collection by active deals outperforms sorting by distance - or the reverse, depending on collection type Phase 4 3 weeks Shipped
Image Variation - GIF vs. Still A GIF image on app cards / collection tabs drives more taps than a static image, in specific placements Low priority Backlog Backlog
Language Variation Collection and cuisine names in Banglish convert better than the equivalent English labels Phase 2 2 weeks Queued
Cuisine Image Dynamicity Different signature images per cuisine (e.g. Jilapi vs. Rosogolla for sweets) lift cuisine-tile CVR Low priority Killed pre-launch
Seasonal Event-Based Design Seasonal theming on app cards and collection tabs (e.g. festival designs) lifts short-window engagement Low priority Backlog Backlog
Restaurant Banner Image Variation A deal-focused restaurant banner drives more footfall than a brand-focused banner design Phase 4 2 weeks Backlog
Collection Image Variation Real food photography beats animated/illustrated imagery on collection-level cards, mirroring the tab-level test Phase 2 2 weeks Queued

Why Collection Placement Order went first: it touched the single highest-traffic real estate on Food Home, had the most direct causal link to conversion of any hypothesis on the list, and required no new UI - only a reordering logic change. Highest expected signal, lowest implementation cost. That combination is what "Phase 0" means in this backlog.


06 — Flagship Test

Collection Placement Order: the test that anchored the whole program.

This was Phase 0 - the first test I ran, and the one every later wave built on. The hypothesis: replacing the static, opinion-ordered collection stack with an engagement-weighted dynamic order would help users find their desired restaurant without defaulting to search.

Design

One control group kept the existing static order. Two treatment groups received different ranking logics for the same collections: Treatment A used a rolling 7-day engagement-weighted order (collections with the highest recent tap-through moved up), and Treatment B used a hybrid order that pinned deal-driven collections to the top regardless of engagement, to test whether promotional priority could coexist with a data-driven order. Both ran for 3 weeks against the same control.

Funnel shift: home vs. search

The clearest single signal from this test was the change in where restaurant visits actually originated from.

Before — Control (Week 0)70% search / 30% home
Home · 30%
After — Treatment A, Week 358% search / 42% home
Home · 42%
Restaurant visits from search Restaurant visits from home / collections

Headline results, Treatment A vs. Control

Collection CVR
+38%
8.4% → 11.6% relative
Search CVR
−9%
24.3% → 22.1%, within guardrail
Overall App CVR
+15%
15.6% → 17.9% relative
Order Count
+9.4%
WoW by week 3, vs. +1.2% control

Week-by-week read

Week Control Treatment A
CollectionSearchOverall CollectionSearchOverall
Week 0 8.3%24.5%15.4% 8.6%24.1%15.6%
Week 1 8.5%24.6%15.5% 9.7%23.6%16.3%
Week 2 8.4%24.3%15.6% 10.8%22.8%17.1%
Week 3 8.6%24.4%15.7% 11.6%22.1%17.9%

Treatment B (deal-pinned hybrid order) matched Treatment A on Collection CVR in deal-heavy collections but underperformed it on Overall App CVR - pinning deals above engagement signal pulled some users toward lower-margin orders. I shipped Treatment A as the new baseline and archived Treatment B's logic for a future test scoped specifically to deal-sensitive cuisines.


07 — Test Portfolio

What the rest of the backlog found.

With the flagship test shipped as the new baseline, later waves tested design and content variables on top of it. A snapshot of what shipped, what's queued, and what each test was really asking.

Cuisine Placement Dynamicity
Shipped
Cuisine tiles reordered by hour of day - breakfast cuisines surfaced in the morning window, dinner-heavy cuisines surfaced in the evening. Cuisine-tile CVR improved +14% relative in the hours where the reorder applied.
Phase 12 weeks
Collection Tab Design - Real vs. Animated
Shipped
Real food photography beat animated/illustrated tab imagery by +21% relative click-through. Illustration only won in dessert and sweets collections, where it was retained as an exception.
Phase 22 weeks
Restaurant Order Within Collection
Shipped
Deals-based sorting won by +9% CVR in deal-sensitive cuisines (fast food, casual dining); distance-based sorting won by +6% in nearby/quick-service categories. Shipped as a contextual sort, not a single global rule.
Phase 43 weeks
Collection Tab View vs. Horizontal View
Queued
Tests whether a dedicated tab surfaces more restaurant visits than the same content in a horizontal scroll. Scheduled next given its dependency on the Phase 2 imagery decisions now shipped.
Phase 32 weeks
Language Variation - English vs. Banglish
Queued
Tests whether Banglish collection and cuisine names outperform English labels for CVR - a localization hypothesis distinct from layout or imagery.
Phase 22 weeks
App Card vs. Horizontal Collection View
Backlog
A layout-density hypothesis - full app-card presentation vs. horizontal scroll for the same collection content. Held behind the tab-view test above since both compete for the same design cycles.
Phase 32 weeks

08 — Program Results

What three waves of testing added up to.

No single test moved the whole funnel. The program did. Stacking the shipped wins from Phase 0 through Phase 4 against the original baseline shows the cumulative shift.

Collection CVR
+38%
Overall App CVR
+13%
Home's Visit Share
30→42%
Orders from Collections
+22%

A side benefit nobody asked for: as home browsing absorbed a larger share of restaurant discovery, search query volume relative to total visits dropped roughly 18%. That's fewer search calls hitting infrastructure per visit, on top of the conversion story - a cost and reliability win that came free from a program built to fix discovery, not infra load.


09 — What I Killed

Cuisine Image Dynamicity - the test I cut before it ever ran.

The original backlog included a hypothesis that different signature images per cuisine - for example, showing "Jilapi" to one cohort and "Rosogolla" to another for the sweets category - would lift cuisine-tile engagement. I killed it during backlog review, before it consumed a testing slot.

Why I killed it: two other tests already in the backlog - Collection Tab Design (real vs. animated imagery) and Collection Image Variation - were already testing the imagery lever on the same surfaces. Adding a third, narrower imagery test would have (1) multiplied the CMS and asset-maintenance burden for marginal incremental signal, (2) competed with those two tests for the same limited traffic slots, and (3) risked a confounded read if it ran anywhere near them in time. The expected signal didn't justify the cost of a dedicated test slot.

This is the least visible part of running an experimentation program, and the one I think matters most for a recruiter to see: saying no to a plausible test is a product decision, not an act of laziness. A backlog with everything in it isn't a strategy - it's a wishlist.


10 — Risks

What could go wrong running a program like this - and how I designed against it.

High
Novelty effect inflating early-week results
A reordered home screen is visually new, and users often engage more with anything different in the first few days regardless of actual quality. Mitigated by reading results at the 3-week mark, not the 3-day mark, and watching for decay in the weekly trend before calling a winner.
High
Local wins that are actually global losses
A test can lift Collection CVR by simply cannibalizing higher-intent search traffic without growing total orders. Mitigated by requiring every result to be read against all four program metrics together, not the headline metric alone.
Medium
Overlapping tests confounding each other's reads
Running imagery, placement, and language tests on the same surface too close together makes it impossible to attribute a metric move to any single change. Mitigated by the phase sequencing itself - deliberately staggering tests that touched overlapping UI.
Medium
Seasonal and calendar confounds
Bangladeshi food-ordering behavior shifts meaningfully around Ramadan, Eid, and other festival periods. Mitigated by avoiding test launches within two weeks of major calendar events, and flagging any test window that overlapped one for a lower-confidence read.
Low
Averaged results hiding cuisine-level or segment-level differences
A program-wide "winner" can still lose for specific cuisines or user segments. Mitigated by breaking out the Restaurant Order test by cuisine category rather than shipping a single global sort rule.
Low
Backlog sprawl outpacing engineering capacity
Twelve hypotheses is enough to keep a small team busy for a year if left unmanaged. Mitigated by the explicit phase gating and the discipline to kill low-signal ideas - like Cuisine Image Dynamicity - before they ever reached a sprint.

11 — Lessons

What running an experimentation program taught me that a single test never would.

01
Prioritization is the actual PM skill in experimentation, not the individual test
Anyone can A/B test a button color. The harder and more valuable skill is deciding which of twelve credible ideas gets built first, with what evidence, and why the other eleven can wait. The backlog itself - the phase sequencing, the impact-vs-effort scoring - is the product artifact I'm proudest of here, more than any single result.
02
Guardrail metrics stop a local win from becoming a global loss
I deliberately allowed Search CVR to dip as part of the success criteria - a drop there wasn't a failure, it was the expected trade-off of home doing its job. Without that guardrail framing set in advance, a well-meaning stakeholder could easily have misread a healthy shift as a regression.
03
Killing a test before it ships is as valuable as shipping a winner
Cutting Cuisine Image Dynamicity freed a testing slot, avoided a CMS maintenance burden, and prevented a confounded read against two related tests. Nobody celebrates a killed test in a standup, but it's exactly the kind of judgment call that separates a PM running a system from one running a checklist.
04
Experimentation needs infrastructure, not one-off hacks
The reusable seven-step methodology - same cohort logic, same metrics, same weekly cadence - is what let me compare a test from week 2 against a test from week 9 with confidence. Building that structure once, upfront, is what made every subsequent test faster and more credible to run.
05
A funnel number is only useful once it becomes a testable hypothesis
"70% of visits come from search" is an interesting fact. It only became useful the moment I turned it into a falsifiable claim - reordering collections will shift that ratio - with a defined success threshold and a kill criterion. That translation from data to testable hypothesis is where product thinking actually happens.
Appendix — The Evidence

Experiment Spec Excerpts

Objectives, success criteria, and the cohort-to-conclusion pipeline as specified in the original experiment design document.

Appendix A — Objectives & Success Criteria
#ObjectiveHow success is judged
O.1 Show different home management layouts to different groups of users via controlled A/B testing Statistically significant difference between control and treatment on the tracked metrics, sustained across the full test duration
O.2 Change collection placement order so users find their desired restaurant and order more easily Collection Conversion Rate increases without a corresponding drop in Overall App Conversion Rate
O.3 Identify the optimal collection placement order across the full backlog of design variants Each shipped winner becomes the new baseline for the next test in that area, compounding gains across waves
O.4 Reduce over-reliance on search as the primary path to restaurant discovery Global request contribution from search decreases as contribution from collections increases - by design, not as an unintended side effect
Appendix B — Monitoring Metrics
MetricDefinitionRole in the program
B.1 Collection Conversion Rate % - share of collection views that result in a restaurant visit North Star - the primary signal that home browsing is doing its job
B.2 Search Conversion Rate % - share of search sessions that result in a restaurant visit Guardrail - a controlled dip is expected and healthy; a collapse is not
B.3 Overall App Conversion Rate % - share of Food Home sessions resulting in a placed order Guardrail - the true business outcome any test must ultimately serve
B.4 Count of Orders - absolute order volume, control vs. treatment Guardrail - the final check against conversion-rate math that looks good on small samples
Appendix C — Cohort-to-Conclusion Pipeline

How a hypothesis in the backlog becomes a shipped (or killed) decision on Food Home - the same seven steps run for every test in the program.

1
Cohort selection
Random sample of users who ordered food in the past month, drawn fresh for each test to keep cohorts representative and independent.
2
Control and treatment assignment
One control group on the existing experience; one or two treatment groups on the variant(s) under test, split randomly and evenly.
3
Duration set to reach a stable read
Typically 2–3 weeks depending on traffic volume for the surface under test - long enough to smooth day-of-week noise, short enough to keep the backlog moving.
4
Weekly metrics tracking
Collection CVR, Search CVR, Overall App CVR, and Order Count tracked identically for control and treatment, every week of the test.
5
Analysis against all four metrics together
Statistical comparison of control vs. treatment - never reading the headline metric in isolation from the guardrails.
Ship, iterate, or kill - feeding the next wave
A shipped winner becomes the new control baseline for the next test in that area. An inconclusive or negative result is either iterated on with a sharper variant or killed outright and the backlog slot reassigned.