Psychographic Research VALS · Clustering · ML 0→1 Intelligence Layer Pathao · 2025

Psychographic
Clustering
Beyond Behavioral Data

How I led Pathao's research initiative to build a personality-driven intelligence layer - designing VALS-based psychographic segmentation from scratch using employee pilots, gamified surveys, and clustering algorithms to unlock what behavioral data alone never could.

6 clusters
VALS-derived psychographic segments scoped for Bangladesh
15-Q
Validated survey instrument with cluster-mapped answer options
3 phases
Employee pilot → Internal validation → App-scale rollout
CTR
Target: Personalized promo CTR lift via psychographic targeting
01 — Context

Pathao knew what users did. But not who they were.

Pathao is Bangladesh's dominant super-app - rides, food delivery, and logistics serving millions of daily users across Dhaka and beyond. By 2025, the behavioral data was rich: spend patterns, ride frequency, food preferences, cross-service usage. The personalization engine was good at serving users based on what they had done before.

But a recurring gap kept surfacing in strategy discussions. Two users with identical spending patterns might order for completely different reasons. One buys from North End every day because they're time-poor professionals who need reliable food fast. Another orders from the same restaurant because it's the aspirational brand they associate with their lifestyle. Same behavior, radically different psychology - and therefore different ideal messages, promotions, and product experiences.

This is the psychographic gap: behavioral data tells you what people do, but not why they do it. And the "why" is what determines how to speak to someone, what product experience feels right for them, and which promotions will actually convert vs. get ignored.

The core insight driving this initiative: Behavioral clustering groups users by past actions. Psychographic clustering groups them by underlying personality, values, and lifestyle orientation - the drivers that explain why those behaviors exist and predict how they'll evolve. Layering both gives Pathao a significantly stronger intelligence foundation for personalization, cross-sell, and even credit decisions for BNPL users with thin financial histories.

This was not a standard feature build. There was no existing dataset to analyze, no validated instrument to adopt off the shelf, and no precedent within Pathao for this kind of research-to-product work. I owned it end to end: from framing the business case, to designing the research methodology, to building the application framework that connects psychographic segments to real product decisions.


02 — The Gap

Three specific places where behavioral data hit its ceiling.

I didn't start with a solution looking for a problem. I started by mapping where current personalization was visibly failing - where knowing what users had done was insufficient to know what to do next for them.

01
Promo targeting felt like mass marketing dressed up as personalization
Behavioral segmentation grouped users by spend bracket and frequency. But sending the same "20% off your next order" to a budget-maximizing Survivor and a trend-chasing Experiencer is wasteful even if they have identical order frequency. The Survivor converts because the discount is real savings. The Experiencer ignores it because the restaurant isn't trending. Psychographic segmentation unlocks the ability to match not just offer type but copy tone, brand choice, and timing to user personality - significantly improving promo ROI without increasing budget.
02
Thin-file BNPL users couldn't be scored using financial history alone
Pathao's BNPL product faced a fundamental problem: a meaningful portion of its user base were new-to-credit or had minimal financial history. Standard credit scoring fails on these users not because they're risky, but because the signals don't exist. Psychographic data - specifically, dimensions like financial discipline, planning orientation, and attitude toward debt - offer a validated alternative signal set. An Achiever who tracks every taka like a "spreadsheet wizard" is a meaningfully different credit risk than a Striver who "spends first, stresses later," even with identical transaction histories.
03
There was no low-cost, scalable way to generate first-party psychographic data
Third-party psychographic data vendors don't exist for Bangladesh at scale. And traditional market research - focus groups, paid surveys - is expensive, slow, and doesn't integrate into a live product. The opportunity I identified was building an in-app collection mechanism that generates psychographic data as a byproduct of an experience users actually enjoy. A gamified personality quiz with shareable results is sticky, viral, and produces the exact data the personalization engine needs - all without users perceiving it as data collection.

03 — My Approach

Research-to-product in three phases, starting internally.

Most product research fails at the implementation stage because the gap between "interesting findings" and "product decisions" is never properly bridged. I designed this initiative so that every research output fed directly into a product decision - no insights sitting in decks without owners.

A key early decision: Start with Pathao employees, not users. Internal pilots let me test the survey instrument, validate clustering logic, and identify question weaknesses without the consequences of a broken experience reaching paying customers. It also created a natural "friend matching" use case that built internal buy-in and generated genuine engagement data before a single line of production code was written.

Phase 1
Internal Pilot
Employee edition - test the instrument and prove the concept
Deploy the survey to Pathao employees with a "who's most like you at work?" matching mechanic. Target: 400+ complete responses to generate a statistically robust dataset for clustering validation. This phase tests question comprehension, completion rate, answer distribution, and whether the predefined clusters emerge naturally from the data or require redesign. Gamified output: "Your team is mostly Experiencers" type insights create organic virality and make data collection feel like entertainment.
Phase 2
Methodology
Validate and refine - use pilot data to stress-test the clustering logic
Run the pilot responses through two approaches in parallel: predefined group assignment (rules-based, fast) and data-based clustering using K-Modes or hierarchical clustering algorithms. Compare the two outputs - where they agree, the predefined groups are validated; where they diverge, that's a signal that the predefined rules need refinement. Preserve the 6D response vector for every respondent so that no data is lost if we later choose to recluster with different algorithm parameters.
Phase 3
Production
In-app rollout - gamified quiz with real personalization output
Deploy the validated instrument as a gamified in-app quiz for Pathao users. Shareable personality results ("You're an Experiencer - you live for the next trending restaurant!") drive peer-to-peer discovery. Psychographic tags are written to user profiles and consumed by the Loop personalization engine, the CRM targeting workflow, and the BNPL scoring model. The data pipeline transforms quiz answers into actionable segment labels within minutes of completion.

04 — VALS Framework

Why VALS - and how I adapted it for Bangladesh.

VALS (Values, Attitudes, and Lifestyles) is a psychographic segmentation framework with decades of academic validation and commercial application. It segments people across two primary dimensions: primary motivation (ideals, achievement, or self-expression) and resource levels (income, education, confidence, health). The intersection produces eight segments.

I chose VALS over other frameworks (Big Five personality traits, Maslow's hierarchy, OCEAN) for three reasons: it directly maps to consumption behavior rather than abstract personality; it has demonstrated commercial validity in marketing contexts; and its segment definitions translate meaningfully to Bangladeshi consumer culture with moderate adaptation.

A deliberate scoping decision: The standard VALS model has 8 segments. I scoped the Pathao implementation to 6, excluding "Makers" (DIY-oriented, minimal relevance in food delivery) and merging "Innovators" into the Experiencer profile for this context (the two segments overlap significantly in Pathao Food usage patterns). Fewer, cleaner segments produce better ML clustering performance and more actionable personalization rules.

The 6 Pathao Psychographic Segments

Achiever
Busy professional, brand-loyal
Time-starved, reliability-driven, brand-conscious. Orders from the same trusted chains repeatedly. Values speed and consistency over novelty.
Food strategy: Fast reliable chains, loyalty perks, "your regular order is one tap away" notifications.
🌟
Experiencer
Young, trend-driven, impulsive
Trend-obsessed, fun-loving, highly social. First to try viral menu items. Influenced by friends and FOMO. Engages with shareable food content.
Food strategy: Trending/viral menus, time-bound flash deals, meme-based or emoji-heavy notification copy.
🧠
Thinker
Analytical, health-conscious
Educated, rational, health-aware. Researches before ordering. Values nutritional information, reviews, and logical value propositions.
Food strategy: Healthy sections, nutrition filters, meal plan framing with logical discount breakdown.
💼
Striver
Aspirational, offer-sensitive
Aspirational but budget-constrained. Wants premium brands at accessible prices. Highly responsive to discounts at aspirational restaurants.
Food strategy: Discounts at premium brands, BNPL options highlighted, "Sultan's Dine at 30% off" framing.
🏠
Believer
Traditional, loyalty-first
Traditional values, deeply loyal to known restaurants and local favorites. Influenced by familiarity, religion, and community recommendations over trends.
Food strategy: Local Bengali restaurants, neighborhood top picks, iftar deals and prayer-time meal alerts.
💰
Survivor
Price-first, risk-averse
Economic survival is the primary lens. Price is the only decision driver. Minimal experimentation, maximum value extraction from every taka spent.
Food strategy: Cheapest local vendors, micro-promo discounts (Tk 10–30 off), free delivery offers front and center.

05 — Survey Design

15 questions. Every option mapped to a cluster. Nothing wasted.

Survey design for psychographic clustering is not the same as a standard UX research survey. Every question must do double duty: feel natural and engaging to the respondent, while precisely mapping each answer option to a specific psychographic dimension. I owned the question design, the answer mapping logic, and the scoring framework.

Design principles I set: Maximum 15 questions (completion rate drops sharply beyond this). All mandatory - partial responses corrupt the clustering vectors. Mixed question tones - some fun/cultural (superhero alter ego, rainy day food), some direct behavioral (payment method, restaurant exploration style). Answer options must cover all 6 target clusters without any cluster being un-selectable in any given question.

Sample Questions - With Cluster Mapping

Q1. Primary motivation for ordering food online?
A
Convenience for a busy schedule
B
Healthy meal options
C
Exciting social experiences
D
Best available price
Maps to: A → AchieverB → ThinkerC → ExperiencerD → Survivor
Q2. Willingness to try unfamiliar restaurants?
A
Love exploring new options
B
Only with good reviews first
C
Stick to trusted, familiar places
D
Only if very affordable
Maps to: A → ExperiencerB → AchieverC → BelieverD → Survivor
Q5. Key loyalty factor for a food delivery app?
A
Service reliability above all
B
Regular discounts and promos
C
Menu variety and new options
D
Traditional and local food options
Maps to: A → AchieverB → StriverC → ExperiencerD → Believer
Q10. Preferred payment method?
A
Mobile wallet (bKash/Nagad)
B
Debit or credit card
C
Cash on delivery
D
Whatever payment has an incentive today
Maps to: A → AchieverB → StriverC → BelieverD → Survivor
Q15. Attitude toward sharing food experiences?
A
Always love documenting and sharing
B
Share special moments only
C
Keep meals private, not a social thing
D
Food is just fuel
Maps to: A → ExperiencerB → AchieverC → BelieverD → Survivor

The full 15-question instrument covers five psychographic dimensions: service usage and lifestyle preferences (4 questions), social and communication style (3 questions), decision-making and values (4 questions), lifestyle and personal preferences (2 questions), and urban living and convenience orientation (2 questions). Each dimension contributes meaningfully to the cluster assignment without overloading any single behavioral dimension.


06 — Clustering Methodology

A two-stage approach: predefined rules first, then ML refinement.

The core clustering challenge is that you can't wait for 400+ responses to assign a cluster to the first respondent - but you also can't lock in rigid rules without data validation. I designed a two-stage methodology that handles both constraints.

Stage 1: Predefined Cluster Assignment (Bootstrapping)

Every answer option maps to a cluster. When a respondent completes the 15 questions, each answer increments a counter for its mapped cluster. The cluster with the highest score is the primary assignment. The 6D vector is preserved for all respondents.

Simulation - Five Respondents

This simulation illustrates how the predefined rules work before any ML clustering is applied. Each number represents the count of answers mapping to that cluster across 15 questions.

Respondent Striver Achiever Thinker Experiencer Believer Survivor Assigned Ratio
A 4 4 6 1 0 0 Thinker 6:4:4
B 0 1 0 9 0 5 Experiencer 9:5:1
C 1 6 1 2 3 1 Achiever 6:3:2
D 0 0 0 0 4 11 Survivor 11:4:0
E 6 2 3 2 0 0 Striver 6:3:2

Stage 2: ML-Based Cluster Refinement

Once 400+ responses are collected from the employee pilot, the predefined rules are stress-tested against data-driven clustering. The 6D response vectors are run through K-Modes clustering (the correct algorithm for categorical data - unlike K-Means, which assumes numerical distance) and hierarchical clustering to see whether natural clusters emerge that match the predefined segments.

Where the two approaches agree

Predefined rules validated

When the algorithmic clusters match the predefined segments, this confirms that the question-to-cluster mapping is correct and that the predefined groups are meaningful real-world categories rather than arbitrary constructs.

Where the two approaches diverge

Predefined rules need revision

Divergence signals either that a question's answer options don't cleanly map to the intended cluster, or that two of our predefined segments are actually the same cluster in the data. These divergences inform the survey instrument refinement before production rollout.

The KNN fallback: For respondents who fall at the boundary between two clusters (e.g., Respondent E above: Striver 6 vs Thinker 3 is a meaningful gap, but Respondent C's Achiever 6 vs Believer 3 might warrant review), a KNN (K-Nearest Neighbors) pass on the full response vector provides a distance-based assignment that's more nuanced than the simple maximum-score rule. The 6D vector preserved for every respondent makes this possible without re-running the survey.


07 — Personalization Application

Psychographic clusters translated into concrete product decisions.

Research that doesn't change product behavior is not product research - it's academic. I built the application framework to ensure every segment definition has a corresponding action across the three primary Pathao Food surfaces where personalization matters.

AI Product Work

Homepage Feed Personalization

The food tile sort order on the Pathao homepage is the highest-leverage personalization surface - it's the first thing every user sees. Psychographic segment determines which restaurant types appear first, not just behavioral recency or rating score.

SegmentHomepage Feed StrategyRationale
Achiever Fast, reliable chains first. KFC, North End, established brands with strong delivery SLA. "Your regular order" shortcut prominent. Achievers optimize for time and reliability. Novelty is noise. Surfacing their established patterns first removes friction from the repeat-order flow.
Experiencer Trending and viral menus at the top. "New this week" and "Everyone's ordering" labels. Flash deal countdown timers visible immediately. FOMO is the primary driver. Social proof and time pressure convert Experiencers more than discount depth. The feed should feel like a live trend feed, not a menu.
Thinker Healthy restaurant section elevated. Nutrition tags visible on food tiles. Curated "Editor's selection" with rationale (not just rating). Thinkers research before deciding. They need information architecture that supports rational evaluation - not emotional triggers. Nutrition data is a feature, not a filter.
Believer Local neighborhood restaurants first. Bengali cuisine prioritized. Prayer-time meal alerts (iftar, sehri). "Your neighborhood's favorites" section. Believers value familiarity, community, and tradition. The feed should feel like a trusted local guide, not a discovery engine pushing unfamiliar brands.
Striver Premium brands at discount. "Normally X, today Y" price framing. BNPL eligibility surfaced on aspirational restaurant tiles. Strivers want access to aspirational brands without the full price. Making the gap feel closeable (discount + BNPL) is the conversion trigger, not the food itself.
Survivor Lowest-cost nearby vendors first. Free delivery options prominently labeled. No-minimum-order restaurants surfaced. Micro-promo (Tk 10–30) visible immediately. Price is the only decision variable. The feed needs to surface the maximum value options without any browsing friction - every scroll is a potential dropout.

Push Notification Personalization

Beyond what to show, psychographic segments determine how to say it. The same offer lands differently depending on the copy tone, urgency framing, and social context used.

Experiencer
🔥 The new tea is breaking the internet - Try before it sells out!
Achiever
⚡ Busy day? Your regular lunch is just one tap away.
Believer
🍛 Your trusted biryani place is back - with a special discount for you.
Striver
✨ Sultan's Dine - now 30% off. Upgrade your lunch without upgrading your budget.
Survivor
💚 Save 20 taka today + free delivery. Order now before the offer expires.
Thinker
🥗 Meal plans that make sense - healthy options with a logical price breakdown.

BNPL Credit Scoring Integration

Psychographic clustering also feeds into Pathao's BNPL underwriting model for thin-file users. Specific question answers correlate with financial behavior dimensions that traditional credit scoring uses financial history to capture - allowing psychographic data to serve as a proxy signal when transaction history is insufficient.

Financial Discipline Signal

"I track every BDT like a spreadsheet wizard"

This answer maps to Achiever or Thinker - both associated with higher financial conscientiousness and lower default risk. In thin-file users, this single response meaningfully shifts the prior for creditworthiness.

Risk Appetite Signal

"I spend first, stress later"

This maps to Experiencer or Striver - clusters with higher risk appetite and impulse-driven spending. For BNPL underwriting, this is a negative signal that increases the default risk prior, enabling tighter credit limits for this cohort.


08 — Metrics

What success looks like - and how to measure a research initiative.

Measuring a psychographic clustering initiative is more complex than measuring a feature launch. The primary output - a richer user intelligence layer - only creates value when downstream systems consume and act on it. I defined metrics across three horizons: data quality, engagement, and business impact.

HorizonMetricWhy it's on the dashboard
Data Quality Psychographic coverage - % of active users with a cluster assignment Personalization only works where data exists. Coverage is the precondition for every downstream application. Low coverage means the system is personalized for a minority of users.
Data Quality Survey completion rate - % of quiz starters who complete all 15 questions Partial responses corrupt the clustering vectors. A completion rate below ~70% signals that the survey is too long, too confusing, or not engaging enough to see through - all solvable product problems.
Engagement Quiz viral coefficient - shares per quiz completion The gamified result-sharing mechanic is both the growth engine and a signal that users found the output meaningful enough to broadcast. A high viral coefficient means low CAC for psychographic data collection.
Engagement CTR lift on psychographically-targeted promos vs. behavioral-only baseline The most direct test of whether psychographic clusters are producing better-matched messaging. If targeted promos don't outperform the behavioral baseline on CTR, the segmentation is either wrong or not being applied correctly.
Business BNPL approval rate uplift for thin-file users with psychographic scores Psychographic data's most commercially significant application. If approval rates increase without a corresponding increase in default rates, the clustering is genuinely adding predictive signal to the credit model.
Business Retention - return order rate within 7 days for psychographically-personalized users vs. control The feed personalization application has a compounding retention effect. Users who consistently see food options aligned with their personality are more likely to order again, more quickly. This metric captures the long-term business case.

09 — Risks

What could go wrong - and how I designed against it.

High
Survey answer bias - respondents gaming toward socially desirable segments
Users who understand the quiz is about personality may answer how they want to be seen (Achiever, Thinker) rather than how they actually behave (Survivor, Striver). Mitigated by: disguising the cluster mapping intent - questions about rainy day food and superhero alter egos feel like entertainment, not self-assessment. Behavioral data cross-referencing during validation catches systematic mismatches between stated and revealed preferences.
High
Low survey completion rate making coverage insufficient for production use
A quiz that 10% of users complete produces psychographic data for a non-representative minority. Mitigated by: gamified mechanics (shareable results, "who's most like you" matching) that make completion intrinsically valuable; limiting questions to 15 with an engaging tone; piloting with employees to identify and fix friction points before public launch.
High
Clustering model producing segments that don't align with actual behavioral differences
If the psychographic segments don't correlate with meaningfully different product behaviors (order frequency, promo sensitivity, restaurant choice), the intelligence layer adds no value regardless of how well the survey works. Mitigated by: systematic cross-referencing of cluster assignments with existing behavioral data during the employee pilot phase - if Achievers aren't actually more likely to repeat-order, the segment definition needs revision before it enters any production system.
Medium
Segment drift - users assigned once but whose personality actually changes over time
A university student classified as Experiencer today may be an Achiever two years into their career. Static clusters assigned at survey completion will become stale. Mitigated by: designing the system to allow re-survey after 6–12 months, monitoring for behavioral divergence from cluster prediction as an early signal of drift, and weighting behavioral signals more heavily as they accumulate to override stale psychographic tags.
Medium
Personalization backfiring by feeling intrusive or stereotyping to users
If a user receives "budget-first" messaging and identifies as aspirational, the experience can feel insulting rather than relevant. Mitigated by: applying psychographic targeting primarily to structural elements (feed order, restaurant type) rather than explicit copy ("because you're a budget user"), and A/B testing personalization depth to find the level of differentiation that improves engagement without creating negative reactions.
Low
VALS framework not translating accurately to Bangladesh consumer culture
VALS was developed and validated in Western consumer markets. Some dimensions (e.g., Maker/DIY orientation) may not map cleanly to Bangladeshi consumer behavior patterns. Mitigated by: the scoping decision to exclude Makers from the implementation and adapt segment definitions to local cultural context (e.g., Believers explicitly incorporates religious practice orientation as a dimension, which VALS standard doesn't emphasize).

10 — Lessons

What this initiative taught me about research-to-product work.

01
Framing the "why" before touching the "what" is the most important PM contribution
Psychographic clustering is a well-established research technique. The harder and more important question was: why does Pathao need this now, and where specifically will it change product decisions? I spent significant time building the business case - mapping it to specific gaps in promo targeting, BNPL underwriting, and personalization feed logic - before any survey design happened. Without a clear application framework, research generates interesting findings that nobody acts on. I built the application plan and the research instrument in parallel to ensure the two stayed aligned throughout.
02
The internal pilot is a genuine product ship, not a prototype throwaway
I could have treated the employee survey as a throwaway test and designed it loosely. Instead, I scoped it as a real product deliverable - designing the full 15-question instrument, the matching mechanics, and the cluster logic from the start. This meant the employee pilot generated production-quality learning, not just directional signals. The data from 400 employees is actually useful for validating the ML clustering methodology, not just for checking whether the survey is "too long."
03
Adapting a global framework for local context is product work, not just localization
Taking VALS and applying it to Bangladeshi users required substantive decisions: which segments to include, how to redefine segment boundaries for local consumer behavior, and which behavioral correlates to test for. The decision to scope to 6 segments, exclude Makers, and merge Innovators into Experiencer wasn't a minor adjustment - it was a fundamental product architecture decision that affected the survey design, the clustering algorithm parameters, and the downstream personalization logic. Frameworks are starting points, not specifications.
04
Data collection UX and data quality are the same problem
The decision to make the survey gamified wasn't a UX preference - it was a data quality decision. High completion rates (all 15 questions, mandatory) produce valid clustering vectors. Drop-off at question 8 produces corrupted data that can't be clustered. Every design decision that improves engagement directly improves data quality, and vice versa. The shareable personality result isn't a marketing add-on; it's the mechanism that makes 15 mandatory questions feel worth completing rather than abandoning.
05
The two-stage clustering methodology is a governance decision as much as a technical one
The choice to start with predefined rules and transition to ML-based clustering after data accumulation looks like a technical architecture decision. It's actually a governance decision about who owns segment definitions. Predefined rules are PM-owned and transparent; ML clusters can drift in ways that are hard to audit or explain to stakeholders. Starting with transparent rules, validating them against ML outputs, and only transitioning to ML where the data supports it means segment definitions remain explainable and defensible to business stakeholders throughout the process.
Appendix — The Evidence

BCD & Research Artifacts

The Business Concept Document, questionnaire framework, and cluster mapping logic - as produced during the research initiative.

Appendix A — Business Concept Document (BCD) Summary
SectionContent
BCD.1 Problem: Behavioral data (spend, frequency) reveals what users do but not why - limiting personalization beyond transactional history. Psychographic gap prevents delivery of true lifestyle-aligned experiences.
BCD.2 Primary objective: Develop psychographic segments and map users into clusters at scale. Metrics: CTR on personalized promos, BNPL approval rate for thin-file users, psychographic data coverage %, quiz engagement rate, behavioral prediction accuracy lift when psychographic traits are layered.
BCD.3 Evidence for prioritization: Behavioral limits (spend/usage don't explain "why"), BNPL credit context (thin-file users need alternative signals), engagement potential (gamified quizzes are sticky and drive peer comparison virality).
BCD.4 Solution: Lightweight psychographic segmentation powered by surveys. Questions map to Big Five/VALS models. Hierarchical clustering or factor analysis defines segments. Clusters cross-referenced with behavioral data and tagged to user profiles for use in recommendation systems and credit engines.
BCD.5 Business impact: Higher marketing ROI from accurate personalization, greater conversion on cross-vertical campaigns from better understanding of user values, BNPL product expansion without compromising default risk, enhanced brand stickiness from gamified interactions.
Appendix B — Full 15-Question Cluster Mapping
#QuestionPsychographic IntentCluster Mapping
Q1 Primary motivation for ordering food online Core value driver identification A→Achiever · B→Thinker · C→Experiencer · D→Survivor
Q2 Willingness to try unfamiliar restaurants Risk tolerance and novelty appetite A→Experiencer · B→Achiever · C→Believer · D→Survivor
Q3 Most important meal consideration Benefit prioritization A→Thinker · B→Striver · C→Experiencer · D→Survivor
Q4 Group ordering frequency Social orientation A→Experiencer · B→Achiever · C→Believer · D→Survivor
Q5 Key app loyalty factor Brand relationship depth A→Achiever · B→Striver · C→Experiencer · D→Believer
Q6 Cuisine selection approach Cultural openness A→Striver · B→Believer · C→Believer · D→Survivor
Q7 Health consciousness level Ideals strength A→Thinker · B→Achiever · C→Experiencer · D→Survivor
Q8 Decision-making process Information processing style A→Thinker · B→Achiever · C→Believer · D→Survivor
Q9 Food trend adoption Innovation readiness A→Achiever · B→Striver · C→Believer · D→Survivor
Q10 Preferred payment method Digital adoption and resources A→Achiever · B→Striver · C→Believer · D→Survivor
Q11 Festival/holiday ordering behavior Tradition vs. convenience A→Achiever · B→Believer · C→Striver · D→Survivor
Q12 Environmental consideration in orders Values consistency A→Thinker · B→Achiever · C→Striver · D→Survivor
Q13 Premium restaurant frequency Resource level indicator A→Experiencer · B→Achiever · C→Striver · D→Survivor
Q14 Meal planning style Structure vs. spontaneity A→Thinker · B→Experiencer · C→Believer · D→Survivor
Q15 Attitude toward sharing food experiences Social expression A→Experiencer · B→Achiever · C→Believer · D→Survivor
Appendix C — Data-to-Segment Pipeline

How a completed quiz response becomes a psychographic segment tag on a user profile and drives personalization decisions across Pathao's product surfaces.

1
Survey completion - 15 mandatory responses collected
User completes the in-app gamified quiz. All 15 questions are mandatory. Partial responses are rejected. The raw answer vector (Q1=C, Q2=A, etc.) is stored against the user ID.
2
Answer-to-cluster translation - 6D vector computed
Each answer is mapped to its cluster weight. The 6D response vector [Achiever, Experiencer, Thinker, Striver, Believer, Survivor] is computed by counting how many answers mapped to each cluster across the 15 questions.
3
Primary cluster assignment - max-score rule with KNN boundary resolution
The cluster with the highest count in the 6D vector is assigned as the primary segment. For boundary cases (top two clusters within 2 points of each other), KNN on the full 6D vector determines final assignment using nearest-neighbor distance across the existing respondent population.
4
Segment tag written to user profile
The primary cluster label (e.g., "Achiever") and the full 6D vector are written to the user's profile. The 6D vector is preserved for future reclassification as the ML model matures. A timestamp is stored for drift monitoring and re-survey scheduling.
5
Downstream consumption - Loop, CRM, BNPL
The Loop personalization engine reads the segment tag to reorder the food feed. The CRM targeting workflow selects promo type, tone, and brand based on segment rules. The BNPL underwriting model reads specific 6D vector dimensions as credit behavior proxy signals for thin-file users.
User experiences a Pathao that knows who they are
Homepage feed sorted by personality-aligned restaurants · Promo copy in the right tone for their motivation style · BNPL access decisions informed by psychographic signals · Shareable personality result driving organic peer acquisition. One quiz, permanent intelligence layer.