01 — Background
A funnel dashboard tells you where users drop off. It doesn't tell you what to do about it.
Pathao Food is the food delivery vertical of Pathao, Bangladesh's leading super-app. By the time I picked up this workstream, the team had a working funnel dashboard - search, restaurant view, menu browse, checkout, and place order were all instrumented, and drop-off at each stage was visible. What didn't exist yet was a structured way to turn those drop-off numbers into a prioritized, shippable set of experiments.
This is the gap most growth teams hit: analytics tells you a story is wrong, but nobody has built the muscle to turn "conversion drops here" into "here are 12 ranked hypotheses, here's how we'll test the top 3, and here's what we'll measure." I took that gap as the brief - not to ship one CRO feature, but to build the operating system for finding, prioritizing, and running conversion experiments across the app, and apply it immediately to the highest-friction parts of the ordering journey and to the search ranking algorithm underneath it.
The reframe: This case study isn't about one feature launch. It's about the process I built to keep finding conversion opportunities - funnel diagnosis, hypothesis writing, prioritization discipline, and a search ranking backlog - so that "improve conversion" stops being a one-off project and becomes a standing program.
02 — Problem
Three structural gaps were blocking conversion gains from ever reaching the roadmap.
01
Friction points were known, but never translated into testable hypotheses
The funnel dashboard could show that the checkout step lost users, or that search-to-order conversion lagged. But a drop-off number is a symptom, not a diagnosis. Nobody had gone stage-by-stage and asked "why here, specifically, and what's the smallest change that would tell us if this theory is right?" Without that translation layer, drop-off data sat in a dashboard and never became a backlog item.
02
Every idea competed for the same engineering queue, regardless of effort or confidence
A one-line color change and a multi-sprint checkout redesign were both pitched the same way - as "improvements." Without a prioritization framework that separated dependency on engineering/data-engineering resourcing from actual impact confidence, the backlog had no mechanism to protect quick, low-risk wins from being crowded out by bigger, uncertain bets.
03
The search ranking algorithm was a static formula nobody was actively experimenting on
Search is the first real decision point in the food ordering journey, and it ran on a fixed scoring formula - name match, featured status, and cuisine tags, with static weights nobody had revisited since launch. A live ranking algorithm that never gets experimented on is a missed lever, especially when small weight changes (how much to reward a featured restaurant, a higher rating, or a higher rating count) can be tested without touching the UI at all.
03 — My Framework
A five-stage lifecycle, so conversion work never stalls at "we noticed a problem."
I structured the program as a repeatable pipeline rather than a one-time audit. Each stage has a clear owner and a clear exit criterion - the pipeline below reflects exactly where the program stood: foundational stages already running, my own deliverable actively in motion, and the downstream stages scoped and ready to hand off.
01
Data Collection
Foundation in place
Instrument and pull conversion rates at every defined funnel checkpoint - search, restaurant view, menu, checkout, place order. This baseline already existed and fed directly into my diagnosis work.
02
Friction Analysis
Foundation in place
Identify drop-offs and unusually long dwell times at each funnel step, so effort gets pointed at the stages actually losing users rather than the stages that feel most visible.
03
Experiment Definition
My core deliverable
Turn every identified friction point into a written hypothesis, a specific variant to test, and an expected outcome - then tag each one by engineering/data-engineering dependency and priority (P0–P4). This is the stage where PM judgment does the heavy lifting, and it's where this case study concentrates.
04
Implementation
Roadmap handoff
Hand finalized variants to design for UI/UX treatment and to engineering for build - sequenced by the prioritization framework, not by whoever asks first.
05
Monitoring
Roadmap handoff
Watch the funnel dashboard post-launch against the specific metric each experiment was designed to move, and feed results back into the next round of prioritization.
04 — Funnel Diagnosis
Five checkpoints between opening the app and a confirmed order - each with its own friction story.
Before writing a single hypothesis, I mapped the ordering journey into five checkpoints and asked what was specifically breaking at each one - not generically, but in terms an engineer or designer could act on.
Discovery & Search
Static ranking weights, no personalization, discount and offer signals buried in filter order
Restaurant & Menu Browse
Missing or poorly cropped food images, thin menu descriptions, no dietary or delivery-time cues
Add to Cart
No sense of urgency or social proof at the point of decision; limited quick-share or discovery loops
Checkout
Single dense screen with no visual sense of progress, raising perceived complexity
Place Order
Low-contrast CTA; rush-hour rider shortages driving system cancellations instead of a wait option
Visual is illustrative of the funnel shape, not a plot of measured conversion data.
05 — Experiment Portfolio
Two parallel tracks, so quick wins never wait on the engineering queue.
The single biggest structural decision I made was to stop treating "the backlog" as one undifferentiated list. I split it into two tracks that could move at completely different speeds.
Bigger bets that touch UI, backend logic, or the ranking pipeline - each needs a proper hypothesis, a built variant, and a monitored rollout.
Place Order CTA
Color Contrast Variant
Hypothesis: a higher-contrast "Place Order" button draws more attention and reduces last-step hesitation.
Checkout Flow
Multi-Step with Progress Bar
Hypothesis: breaking checkout into steps with a visible progress indicator reduces perceived complexity and abandonment.
Search Results Display
Grid View vs. Detailed List View
Hypothesis: image-led grid results speed up recognition, while a detailed list reduces mis-selection through richer information.
Twenty-three ideas that needed no engineering build - content quality, ordering logic on existing app cards and filters, and small UX affordances. A representative sample:
Complete the menu image gapsEvery menu item gets a photo; visual gaps were quietly costing conversions on otherwise-strong restaurants.
Fix cropped restaurant & food bannersSeveral banners were mis-framed and cropping out the food itself.
Add enticing menu descriptionsShort, appetite-driving copy for every item, not just a name and price.
"Popular for Lunch" app cardA time-boxed (11am–4pm) card tagging restaurants with lunch-appropriate items, mirroring the existing "sokaler nasta" breakfast card.
Reorder the discount filterMove "Offering Discount" earlier in the filter row - it currently sits after Sort, Filter, and Above-4-star and is easy to miss.
Surface restaurants with offers firstWithin "All Restaurants," restaurants with active offers should rank ahead of ones without.
Reorder budget-meal app cardBring the budget-meals card ahead of "Delicious Deals," where it's currently buried in 4th position.
Weather- & time-aware suggestionsMood-based food prompts (juicy/cold items on hot days, warm items on cool days) and dietary icons for quick scanning.
Rush-hour waitlist instead of cancellationOffer an estimated-wait option during rider shortages instead of an automatic system cancellation.
In-app quick tips & delivery-time tagsTooltips for low-adoption features, plus visible delivery-time estimates per item to set expectations upfront.
Why split it this way: Track B could start shipping within days - no roadmap slot required - while Track A's bigger bets went through proper design and engineering scoping in parallel. The two tracks never competed for the same resource, so momentum on quick wins didn't stall while the bigger tests were being built.
06 — Deep-Dives
Three experiments, worked through end-to-end.
Below is the level of detail I took every experiment to before it was ready to hand off - hypothesis, variant, dependency, and the metric that would prove or kill it.
Place Order Button — Color Contrast Test
Track A · P1
Hypothesis
A more visually distinct, contrasting CTA color draws the eye at the final decision point and reduces drop-off right before order confirmation.
Variant
Swap the existing "Place Order" button to a higher-contrast color against its surrounding UI, holding copy and placement constant.
Dependency
Front-end only - a one-sprint change, no backend or data-engineering involvement required.
Success Metric
Cart-to-order conversion rate, measured against the existing button as control.
Checkout — Multi-Step Flow with Progress Bar
Track A · P1
Hypothesis
Splitting checkout into clear steps with a visible progress bar gives users a sense of control and lowers the perceived complexity that causes mid-checkout abandonment.
Variant
Replace the single dense checkout screen with a paged flow (address → payment → review) and a persistent progress indicator across the top.
Dependency
Requires front-end rework and QA across payment states - the largest build in Track A.
Success Metric
Checkout-start-to-completion rate, plus time-to-complete-checkout as a guardrail against added friction.
"Popular for Lunch" App Card
Track B · No Dependency
Hypothesis
Users will discover lunch-appropriate restaurants more easily from a dedicated, time-boxed app card, increasing footfall into those restaurants during lunch hours.
Variant
A new app card, visible 11am–4pm, tagging restaurants with genuinely lunch-appropriate menus (explicitly excluding sweet shops and fast-food-only outlets) - placed beside the existing breakfast-focused card.
Measurement Design
Because no prior baseline existed for tagged restaurants, I specified capturing each restaurant's baseline performance before placement in the card, then comparing post-placement performance against that same baseline - rather than comparing against untagged restaurants, which would confound the read.
Success Metric
Footfall and order conversion lift for tagged restaurants, baseline vs. post-placement.
07 — Search Ranking
The search algorithm nobody was experimenting on - until a prioritized backlog existed for it.
Search ranking scored restaurants on name match, featured status, and primary/secondary cuisine tags, with static weights that hadn't been revisited since launch. I treated the scoring formula itself as an experimentation surface and built a prioritized backlog of weight and logic changes, each with its own rationale.
| Priority | Experiment | What Changes | Expected Outcome |
| P0 |
Badge restaurants I've ordered from before |
Surface a "previously ordered" badge in search results |
Faster re-order decisions, higher repeat-order conversion |
| P0 |
Rank restaurants' active deals higher |
Give live discounts/deals more ranking weight |
Better matching of price-sensitive intent with visible offers |
| P1 |
Increase "is_featured" weight |
Raise featured-tag weight from ~25% toward 30–40% of composite score |
More attention and conversion on featured restaurants |
| P1 |
Reward higher star ratings |
Add tiered bonus points above defined rating thresholds (e.g. 4.0–4.3, >4.3) |
Higher user satisfaction from surfacing better-reviewed restaurants |
| P1 |
Reward higher rating volume |
Add bonus points once a restaurant's rating count clears a threshold |
Trust signal from social proof reflected in ranking, not just average score |
| P1 |
Personalize with order & search history |
Weight results by the individual user's past behavior |
More relevant top-of-search results per user |
| P1 |
Show price signal in food-item results |
Add a visible price indicator directly in search results |
Faster filtering by budget without an extra tap |
| P2 |
Data-quality scoring |
Bonus points for restaurants with complete data and quality images |
Incentivizes restaurants to keep listings updated |
| P2 |
Novelty score for new restaurants |
Temporary ranking boost during a restaurant's first weeks live |
Higher exploration and conversion for new supply |
| P3 |
Time-of-day personalization |
Boost lunch/dinner/late-night tagged restaurants during their relevant hours |
More relevant results tied to search timing |
| P4 |
Seasonal cuisine highlight |
Boost seasonal tags (BBQ in summer, hot beverages in winter) |
Incremental seasonal conversion lift |
08 — Prioritization
A two-axis framework: dependency on the x-axis, priority tier on the y-axis.
Every experiment - whether a funnel fix or a ranking-weight change - got plotted on the same two axes: does it need engineering/DE resourcing, and how high is its priority tier (P0–P4, reflecting my confidence in impact and how directly it addresses a diagnosed friction point). That gave the team a shared, defensible answer to "what do we build next?"
Ship Now
No Dependency · High Priority
Menu image fixes, discount filter reorder, offer-first sorting, "previously ordered" badge logic. Ships within days; no roadmap slot needed.
Roadmap Bets
Requires Eng/DE · High Priority
Checkout progress bar, CTA color test, featured-weight and rating-weight tuning in the ranking algorithm. Scoped, sequenced, and handed to engineering with clear success metrics.
Nice to Have
No Dependency · Lower Priority
Localized loading-screen quotes, weekly digest emails, social check-in rewards. Low cost, but not tied to a diagnosed friction point - queued behind higher-confidence items.
Backlog
Requires Eng/DE · Lower Priority
Seasonal cuisine highlighting, time-of-day ranking personalization. Kept visible in the backlog rather than discarded, so they resurface once the higher-priority queue clears.
Why this mattered: without separating "needs engineering" from "priority," the backlog defaults to whichever idea is loudest, not whichever idea is most likely to move conversion. The quadrant made trade-offs visible and defensible in planning conversations, instead of implicit and political.
09 — Metrics
Every experiment ships with the metric that will kill it, defined before launch.
North Star: Search-to-order conversion rate - the share of search sessions that end in a placed order. Every experiment in this program, whether a funnel fix or a ranking change, ultimately has to move this number or a clearly-linked stage metric feeding into it.
| Stage | Metric | Why it's on the dashboard |
| Search |
Click-through rate on top-3 search results |
Direct read on whether ranking-weight changes are surfacing more relevant restaurants at the top of results. |
| Browse |
Menu-to-cart conversion rate |
Tests whether image, description, and dietary-icon fixes actually translate browsing into an add-to-cart action. |
| Checkout |
Checkout-start-to-completion rate |
The core read on whether the progress-bar variant reduces mid-checkout abandonment. |
| Place Order |
Cart-to-order conversion rate |
Isolates the very last step, where the CTA color test is designed to have its effect. |
| Merchandising |
Footfall lift for restaurants placed in a new app card, baseline vs. post-placement |
The only fair way to evaluate app-card placements where no pre-existing comparison group exists. |
| Guardrail |
System-cancellation rate during rush hours |
Ensures the rush-hour waitlist experiment is actually reducing cancellations, not just delaying them. |
10 — Risks
What could go wrong - and how the program was designed against it.
High
No pre-existing baseline for new app-card placements
Without a "before" state, any lift from a new card like "Popular for Lunch" is unprovable. Mitigated by requiring baseline capture for every tagged restaurant before placement, not after.
High
Quick wins shipped without design review degrade trust in the UI
The speed advantage of Track B could tempt the team to skip design polish entirely. Mitigated by requiring a lightweight design sign-off even for no-dependency changes, so speed didn't come at the cost of visual consistency.
High
Ranking changes reward "featured" or "novel" restaurants at the expense of genuine quality
Over-weighting featured status or novelty could surface paid or new placements over restaurants users actually prefer. Mitigated by pairing every weight increase with a rating-quality and rating-volume signal, so no single lever can dominate the score.
Medium
Running Track A and Track B experiments simultaneously confounds the read
If a checkout redesign and a CTA color test both ship in the same window, it's hard to attribute conversion movement to either. Mitigated by sequencing overlapping-funnel-stage experiments and keeping non-overlapping-stage experiments running in true parallel.
Medium
Rush-hour waitlist experiment shifts poor experience from "cancelled" to "kept waiting"
Replacing cancellations with a wait option only helps if the estimated wait is accurate. Mitigated by tying the offered wait time to real-time rider availability rather than a fixed estimate, and treating cancellation rate as a guardrail metric.
Low
Backlog becomes stale as market conditions shift
A P3/P4 item like seasonal cuisine highlighting can lose relevance if not revisited. Mitigated by keeping the prioritization framework as a living document reviewed each planning cycle, not a one-time ranking exercise.
11 — Lessons
What building an experimentation program - rather than a single feature - taught me.
01
A dashboard shows you a wound; a framework tells you how to treat it
Funnel analytics were already good before I started. What was missing was the discipline of turning "conversion drops here" into a written hypothesis, a scoped variant, and a defined success metric. The real PM deliverable in this program wasn't a chart - it was the translation layer between data and a shippable experiment.
02
Splitting the backlog by dependency protects momentum
The single highest-leverage structural decision in this program was separating no-dependency quick wins from engineering-dependent bets into two tracks. Momentum on the easy wins never had to wait for the hard ones to get scoped - and the hard ones got proper design and engineering time instead of being rushed to keep pace with quick wins.
03
A live ranking algorithm is an experimentation surface, not a fixed asset
Search ranking had run unchanged since launch simply because nobody had framed it as testable. Treating the scoring weights themselves as a backlog of experiments - not a black box - unlocked ten prioritized ranking tests that require no new UI at all, just changes to how existing signals are weighted.
04
Measurement design has to be solved before the experiment, not after
The "Popular for Lunch" app card had no natural control group - there was no historical baseline for restaurants that had never been featured. I had to design the before/after comparison methodology as part of writing the hypothesis, not leave it for whoever analyzed results later. An experiment with no measurement plan isn't an experiment.
05
Prioritization frameworks earn trust by making trade-offs visible, not by being clever
The P0–P4 x dependency quadrant isn't a novel framework - its value was that it gave stakeholders a shared, visible reason for why one experiment shipped before another. Once trade-offs are visible, prioritization debates stop being about whose idea is loudest and start being about the framework's inputs - which is a much healthier argument to have.