Pathao Food is Bangladesh’s largest food delivery platform. Search is where intent converts or dies.
Pathao operates as a super-app across ride-hailing, parcel delivery, and food delivery in Bangladesh. The food vertical competes directly with the local arms of regional players and several domestic challengers. At this scale, even marginal improvements in search quality compound into significant order volume.
I owned the food product experience as PM and identified search as an under-invested surface relative to its conversion importance. A user who types a query and gets irrelevant results doesn’t complain - they leave. The search-to-order funnel was leaking in ways that were invisible without deliberate instrumentation.
Pathao Super-App
Operating across ride-hailing, parcel, and food in Bangladesh. Food delivery is one of the highest-frequency verticals, with returning users placing multiple orders per week. Search is a primary entry point for discovery - not just reordering.
Product Manager, Food
I was the PM responsible for the food delivery product experience. This initiative was self-initiated - I identified search quality as a high-leverage opportunity, built the case, wrote the CRD, and drove the experiment programme from problem framing to spec.
Why search specifically: Conversion funnels in food delivery have a well-understood structure - open app, discover or search, select restaurant, add items, checkout. Most optimisation attention goes to checkout and restaurant pages. Search sits upstream of all of it: a failed search means no restaurant page visit, no menu browse, no order. I treated search quality as a funnel multiplier, not a secondary feature.
Users searching in Banglish, with typos, or without context weren’t getting results that matched their intent.
Food search in Bangladesh has characteristics that make it harder than search in most other markets. Bangladeshi users routinely write in Banglish - Bengali words rendered in English script. “Biriyani” might be typed as “biryani”, “biriani”, “biriany” or “biryanee”. A search engine that requires exact matches fails a significant share of real user intent.
Beyond language, search results at the time were broadly ranked without personalisation. A new user and a returning user with fifty orders see the same list. A restaurant a user orders from weekly appears in the same position as one they’ve never visited. There were no signals for favourites, no indicators of deal availability, no acknowledgment of order history in the results layer.
The four problem areas I scoped
Before writing a single spec, I built the evidence case that each experiment was worth building.
The biggest mistake a PM can make with a list of “improvements” is to start building without validating that the improvements will move the right metrics. I structured this programme around a deliberate sequencing: prove the ROI case first, then write the spec.
For each experiment, I defined the statistics needed to confirm the intervention would produce a return before engineering invested any time. This framework borrowed from classic hypothesis testing: identify what you need to measure, define what outcomes would justify the build, and only proceed when the data supports it.
The discipline I imposed on myself: For every experiment, I asked “what data would tell me this is not worth building?” before I asked “how do we build it?” A search fix that affects 1% of queries and improves conversion by 0.1% is not worth an engineering sprint. A fix that affects 30% of queries and improves conversion by 5% absolutely is. I didn’t want to build a backlog of features - I wanted to build the right features.
How I worked cross-functionally
Quantifying the problem
I worked with the analytics team to pull the baseline metrics: what % of search queries contain typos or Banglish spellings, and what the conversion rate gap looks like between users who search normally vs. those who make typo errors. These numbers determined which experiments to prioritise.
Feasibility and spec review
Each experiment spec went through a feasibility review with engineering before being locked. I wrote the functional requirements; the eng team validated implementation complexity and flagged any dependencies on backend data availability (e.g. whether order history and favourites were already accessible to the search ranking layer).
UI patterns for new signals
Several experiments required new UI elements - icons for previously-ordered items, favourite badges, deal indicators, currency tier symbols. I worked with design on the iconography, ensuring each new signal was visually distinct and legible without cluttering the results card.
I defined a data-first decision matrix before any experiment entered the spec stage.
The most rigorous element of this initiative was the statistical framework I built to gate which experiments we built. Rather than relying on intuition or stakeholder enthusiasm, I defined a two-variable test for each experiment: the prevalence of the problem, and the conversion cost of the problem.
The Banglish/Typo Case - the clearest example: Metric A: what % of search queries are typos or Banglish variations? Metric B: do users who make typo searches convert at a meaningfully lower rate? The outcome matrix below is exactly how I framed it to stakeholders. This wasn’t a gut call - it was a decision tree that any stakeholder could understand and challenge.
Decision matrix for the typo experiment
I applied the same logical framework across all six experiments - defining the measurable conditions that would justify each build before writing any functional requirements. This is the habit that separates roadmaps grounded in evidence from backlogs driven by enthusiasm.
Six experiments. Each specced with a clear hypothesis, measurable success condition, and defined UI behaviour.
Below are all six search experiments from the Change Requirement Document I authored. Each was assigned a spec ID (SE-1 through SE-7, excluding SE-2 which was merged), a priority, and a full behavioural description. The priority was set collaboratively with stakeholders using the ROI framework above.
If a user has ordered from a restaurant previously, the search result for that restaurant displays an icon on the right side of the card indicating prior purchase history. The icon differentiates between a single prior order and multiple prior orders - giving the user an at-a-glance signal of their relationship with that restaurant.
Why this matters: Returning users are the highest-conversion segment in food delivery. They already trust the restaurant. A small icon that reminds them “you’ve ordered here before” reduces decision friction at precisely the moment a user is scanning results. It’s a nudge, not a feature - but nudges at high-intent moments compound into measurable conversion lift.
Map Banglish words (Bengali rendered in English script) and common phonetic variations to canonical search targets. A user typing “biriyani,” “biriany,” or “bryani” should resolve to the same target as “biryani.” Similarly, “kacchi,” “khichuri,” and other culturally-specific dish names need robust variant coverage.
The evidence gate I set: I required two data points before greenlighting this build. First: what % of search queries are Banglish or contain typos? Second: do these queries convert at a lower rate than “clean” queries? High prevalence + low conversion = unambiguous build case. The decision matrix in Section 04 was built specifically for this experiment.
Why this is a Bangladesh-specific problem: Most search algorithms are trained on English-language corpora. Banglish is not phonetically consistent - the same Bengali word is romanised differently by different users. A standard fuzzy-match approach underperforms because the variants aren’t just typos; they’re culturally-specific linguistic patterns. I specified that the mapping layer needed explicit curation for the most common dish and cuisine categories in the Pathao Food catalogue.
Add a quality-weighted scoring component to search ranking. Restaurants above a threshold rating value and above a threshold rating count receive additional ranking score points. This is a combined signal - a restaurant needs both a high rating and a meaningful volume of ratings to receive the boost.
Why the dual threshold matters: A restaurant with 5 stars from 3 orders is statistically meaningless. A restaurant with 4.6 stars from 800 orders is a genuine quality signal. Using count alone disadvantages newer-but-excellent restaurants; using rating alone amplifies noise from low-volume outliers. The dual threshold creates a quality signal that is both reliable and fair to the catalogue.
I defined this as “Should have” rather than “Must have” because the ROI depends on how correlated current ranking quality is with ratings. If the default ranking already surfaces high-rated restaurants, the marginal lift is lower. The data review was required to confirm the gap before committing engineering resources.
Identify restaurants in search results where the user has a prior order and where the restaurant’s rating exceeds a defined threshold. Mark those restaurants with a distinct icon in the search result card. This combines the personalisation signal from SE-1 with a quality filter - surfacing restaurants the user already trusts and that have objectively high satisfaction scores.
The compounding effect: SE-1 tells a user “you’ve ordered here before.” SE-4 tells them “you’ve ordered here before, and this restaurant is highly rated.” The combined signal reduces the research burden for a returning user making a familiar decision. It’s the difference between recognising a restaurant and being confidently directed back to one you already know is good.
If a search result contains a restaurant the user has saved to their favourites list, display a “favourite” icon on that restaurant’s search result card. This makes the favourites feature visible at the discovery layer rather than buried in a separate profile tab.
Why this closes a loop: Users save restaurants to favourites because they intend to return. But if they don’t navigate directly to the favourites tab, that signal is invisible at the point of discovery. Surfacing the favourite badge in search results means the act of saving a restaurant actually changes the future search experience. This makes favourites a more valuable behaviour to encourage, which has upstream effects on engagement and app stickiness.
Restaurants with active deals receive a ranking boost in search results. Additionally, an “Offers” filter chip appears at the top of search results. When activated, the filter shows only restaurants with active deals; when deactivated, all restaurants reappear in the default ranked order.
The conversion rationale: Deals are one of the highest-intent conversion signals in food delivery. A user who can see “20% off” on a restaurant they were already considering will convert at a materially higher rate than one who doesn’t know a deal exists. Burying deals inside restaurant pages means most users never discover them. Surfacing deal availability at the search layer closes the discovery gap.
The filter chip also addresses a specific user behaviour: deal-seekers who open the app with the explicit goal of finding offers. These users currently have no efficient way to filter for deals except by browsing. The filter turns a manual hunt into a one-tap action.
Add a price tier indicator to each restaurant search result card, using one, two, or three currency (৳) symbols to represent the restaurant’s relative price range. The tier is calculated from the restaurant’s median basket size, classified into three buckets: lower, mid-level, and higher median basket.
Why median basket size: A restaurant’s average order value is a better proxy for a user’s expected spend than item prices, because it reflects real ordering behaviour - including add-ons, portions, and typical combination orders. Median (rather than mean) is used to reduce the effect of outlier orders. The three-tier classification keeps the signal simple and actionable without requiring users to process exact numbers.
The discovery problem it solves: Without a price signal in search results, a user browsing results for “pizza” has no way to distinguish a casual fast-food option from a premium restaurant before clicking through. This leads to click-throughs that end in drop-off when the user sees the price range. The ৳/৳৳/৳৳৳ indicator reduces this friction by setting price expectations at the search layer, before a user invests attention in a restaurant page.
Priority summary across all experiments
| ID | Experiment | Priority | Signal category |
|---|---|---|---|
| SE-1 | Indicators for previously ordered items | Must Have | Personalisation |
| SE-Banglish | Banglish & typo correction mapping | Must Have | NLP / Language |
| SE-3 | Rank restaurants with higher ratings higher | Should Have | Quality Ranking |
| SE-4 | Badge: prior order + high-rated restaurant | Must Have | Personalisation |
| SE-5 | Badge for favourite restaurants | Must Have | Personalisation |
| SE-6 | Rank deal restaurants higher + offers filter | Must Have | Deal Discovery |
| SE-7 | Price tier indicator (৳/৳৳/৳৳৳) | Must Have | Price Discovery |
How I would measure success - and what I’d monitor to know if something was going wrong.
North Star metric: Search-to-order conversion rate - the % of search sessions that result in a completed order. All six experiments are designed to move this metric. Sub-metrics exist per experiment to diagnose which lever is working and which isn’t.
| Metric category | Metric | Why it’s tracked |
|---|---|---|
| Core conversion | Search-to-order conversion rate (overall) | The single metric that tells us whether the search improvements are producing business value. Measured pre/post experiment launch and monitored weekly. Any regression triggers immediate investigation. |
| Language quality | Zero-result query rate & typo/Banglish query share | Baseline for the Banglish experiment. Post-launch, zero-result rate should decline as more variant spellings resolve to valid restaurant matches. We also track what % of queries are detected as Banglish or typo-type to validate the initial ROI assumption. |
| Personalisation | Click-through rate on results with personalisation badges vs. without | Validates that the prior-order, favourite, and high-rated badges are actually influencing user behaviour. If badged results don’t show a higher CTR than undecorated results, the badge design may need revisiting before we invest in further personalisation depth. |
| Deal engagement | Offers filter usage rate & conversion rate on filtered sessions | If the filter is never used, the deal-seeker hypothesis was wrong. If it’s used but doesn’t improve conversion, the deal inventory may be too thin. Both scenarios tell us something different about the business health of the deals programme, not just the search UX. |
| Quality | Post-order satisfaction score (ratings given after orders from search) | SE-3 and SE-4 are designed to surface higher-quality restaurants. If they’re working, the orders they generate should produce higher satisfaction scores than the baseline. A drop in post-order ratings after the ranking change is a signal we’ve over-optimised for the wrong quality signal. |
| Discovery | Uninformed click-through bounce rate (users who view & exit restaurant page without ordering) | SE-7 (price tier) is specifically designed to reduce click-throughs that end in drop-off due to price mismatch. If the price indicator is working, the bounce rate from restaurant pages for users who arrived via search should decline. A stable bounce rate post-launch suggests price wasn’t the primary drop-off driver. |