How I turned a single "reorder the home screen" request into a prioritized, 12-experiment testing program for Pathao Food's home surface - shifting how millions of users discover restaurants from search-dependent to browse-driven, with a repeatable methodology I could run again on any surface.
Pathao is Bangladesh's largest super-app, and Food is one of its highest-traffic verticals, competing daily for order share against foodpanda and other delivery players. Like most food-delivery apps, Food Home is the first screen a hungry user sees - a stack of horizontally scrolling collections (cuisines, deals, trending restaurants, curated picks) meant to help people find something to eat without typing a single letter.
When I pulled the funnel numbers, the picture was uncomfortable. Only 10–12% of App Home users ever reached Food Home in the first place. Of those who did, 70% went on to visit a restaurant page - but of those visits, 70% originated from search and only 30% from browsing the home screen itself. The surface built specifically to help people discover restaurants without searching was losing to the search bar, more than two to one.
The uncomfortable question: every collection, every cuisine tile, every ordering decision on that home screen had been shipped on design opinion - never tested. Nobody actually knew whether the existing layout was helping users or quietly pushing them toward search. I decided to find out, systematically.
When I broke down why Food Home was underperforming, I didn't find one bug to fix. I found a surface with a dozen untested variables competing for the same real estate - which meant the real problem wasn't "what's the right layout," it was "how do I build a system that can find out."
The brief I was handed was simple: "make the home screen convert better." I deliberately reframed it. Rather than pick one design and A/B test it, I designed an experimentation program - a prioritized backlog of testable hypotheses, a shared methodology every test would follow, and a metrics framework that let me compare results across tests run months apart.
I designed one methodology and reused it across the entire program, so results from a test run in week 2 could be trusted and compared against a test run in week 9.
Rather than run whichever test was easiest to build first, I scored the full hypothesis list on expected impact (traffic exposure of the surface being changed, and how directly it connected to conversion) against implementation effort, and sequenced the program into phases. This is the artifact I used to align design, engineering, and category stakeholders on what would get tested when - and why.
| Treatment | Hypothesis | Phase | Duration | Status |
|---|---|---|---|---|
| Collection Placement Order | Reordering collections by engagement-weighted logic instead of static order surfaces the right restaurant faster | Phase 0 | 3 weeks | Shipped |
| Cuisine Placement Dynamicity | Cuisine position should shift by hour of day (breakfast items up mornings, biryani/dinner cuisines up evenings) | Phase 1 | 2 weeks | Shipped |
| Collection Tab Design Variation | Real food photography outperforms animated/illustrated tab imagery on click-through | Phase 2 | 2 weeks | Shipped |
| Collection Tab View vs. Horizontal View | A dedicated collection tab converts better than the same content shown in a horizontal scroll row | Phase 3 | 2 weeks | Queued |
| App Card vs. Horizontal Collection View | Full app-card presentation of a collection outperforms the equivalent horizontal scroll layout | Phase 3 | 2 weeks | Backlog |
| Restaurant Order Within Collection | Sorting restaurants inside a collection by active deals outperforms sorting by distance - or the reverse, depending on collection type | Phase 4 | 3 weeks | Shipped |
| Image Variation - GIF vs. Still | A GIF image on app cards / collection tabs drives more taps than a static image, in specific placements | Low priority | Backlog | Backlog |
| Language Variation | Collection and cuisine names in Banglish convert better than the equivalent English labels | Phase 2 | 2 weeks | Queued |
| Cuisine Image Dynamicity | Different signature images per cuisine (e.g. Jilapi vs. Rosogolla for sweets) lift cuisine-tile CVR | Low priority | — | Killed pre-launch |
| Seasonal Event-Based Design | Seasonal theming on app cards and collection tabs (e.g. festival designs) lifts short-window engagement | Low priority | Backlog | Backlog |
| Restaurant Banner Image Variation | A deal-focused restaurant banner drives more footfall than a brand-focused banner design | Phase 4 | 2 weeks | Backlog |
| Collection Image Variation | Real food photography beats animated/illustrated imagery on collection-level cards, mirroring the tab-level test | Phase 2 | 2 weeks | Queued |
Why Collection Placement Order went first: it touched the single highest-traffic real estate on Food Home, had the most direct causal link to conversion of any hypothesis on the list, and required no new UI - only a reordering logic change. Highest expected signal, lowest implementation cost. That combination is what "Phase 0" means in this backlog.
This was Phase 0 - the first test I ran, and the one every later wave built on. The hypothesis: replacing the static, opinion-ordered collection stack with an engagement-weighted dynamic order would help users find their desired restaurant without defaulting to search.
One control group kept the existing static order. Two treatment groups received different ranking logics for the same collections: Treatment A used a rolling 7-day engagement-weighted order (collections with the highest recent tap-through moved up), and Treatment B used a hybrid order that pinned deal-driven collections to the top regardless of engagement, to test whether promotional priority could coexist with a data-driven order. Both ran for 3 weeks against the same control.
The clearest single signal from this test was the change in where restaurant visits actually originated from.
| Week | Control | Treatment A | ||||
|---|---|---|---|---|---|---|
| Collection | Search | Overall | Collection | Search | Overall | |
| Week 0 | 8.3% | 24.5% | 15.4% | 8.6% | 24.1% | 15.6% |
| Week 1 | 8.5% | 24.6% | 15.5% | 9.7% | 23.6% | 16.3% |
| Week 2 | 8.4% | 24.3% | 15.6% | 10.8% | 22.8% | 17.1% |
| Week 3 | 8.6% | 24.4% | 15.7% | 11.6% | 22.1% | 17.9% |
Treatment B (deal-pinned hybrid order) matched Treatment A on Collection CVR in deal-heavy collections but underperformed it on Overall App CVR - pinning deals above engagement signal pulled some users toward lower-margin orders. I shipped Treatment A as the new baseline and archived Treatment B's logic for a future test scoped specifically to deal-sensitive cuisines.
With the flagship test shipped as the new baseline, later waves tested design and content variables on top of it. A snapshot of what shipped, what's queued, and what each test was really asking.
No single test moved the whole funnel. The program did. Stacking the shipped wins from Phase 0 through Phase 4 against the original baseline shows the cumulative shift.
A side benefit nobody asked for: as home browsing absorbed a larger share of restaurant discovery, search query volume relative to total visits dropped roughly 18%. That's fewer search calls hitting infrastructure per visit, on top of the conversion story - a cost and reliability win that came free from a program built to fix discovery, not infra load.
The original backlog included a hypothesis that different signature images per cuisine - for example, showing "Jilapi" to one cohort and "Rosogolla" to another for the sweets category - would lift cuisine-tile engagement. I killed it during backlog review, before it consumed a testing slot.
Why I killed it: two other tests already in the backlog - Collection Tab Design (real vs. animated imagery) and Collection Image Variation - were already testing the imagery lever on the same surfaces. Adding a third, narrower imagery test would have (1) multiplied the CMS and asset-maintenance burden for marginal incremental signal, (2) competed with those two tests for the same limited traffic slots, and (3) risked a confounded read if it ran anywhere near them in time. The expected signal didn't justify the cost of a dedicated test slot.
This is the least visible part of running an experimentation program, and the one I think matters most for a recruiter to see: saying no to a plausible test is a product decision, not an act of laziness. A backlog with everything in it isn't a strategy - it's a wishlist.
Objectives, success criteria, and the cohort-to-conclusion pipeline as specified in the original experiment design document.
| # | Objective | How success is judged |
|---|---|---|
| O.1 | Show different home management layouts to different groups of users via controlled A/B testing | Statistically significant difference between control and treatment on the tracked metrics, sustained across the full test duration |
| O.2 | Change collection placement order so users find their desired restaurant and order more easily | Collection Conversion Rate increases without a corresponding drop in Overall App Conversion Rate |
| O.3 | Identify the optimal collection placement order across the full backlog of design variants | Each shipped winner becomes the new baseline for the next test in that area, compounding gains across waves |
| O.4 | Reduce over-reliance on search as the primary path to restaurant discovery | Global request contribution from search decreases as contribution from collections increases - by design, not as an unintended side effect |
| Metric | Definition | Role in the program |
|---|---|---|
| B.1 | Collection Conversion Rate % - share of collection views that result in a restaurant visit | North Star - the primary signal that home browsing is doing its job |
| B.2 | Search Conversion Rate % - share of search sessions that result in a restaurant visit | Guardrail - a controlled dip is expected and healthy; a collapse is not |
| B.3 | Overall App Conversion Rate % - share of Food Home sessions resulting in a placed order | Guardrail - the true business outcome any test must ultimately serve |
| B.4 | Count of Orders - absolute order volume, control vs. treatment | Guardrail - the final check against conversion-rate math that looks good on small samples |
How a hypothesis in the backlog becomes a shipped (or killed) decision on Food Home - the same seven steps run for every test in the program.