Overview
This started with a very practical question:
“What do customers usually buy together?”
The store had a wide variety of imported snacks, but everything was treated independently.
Each product was stocked, priced, and displayed in isolation.
But in reality, customers don’t shop like that.
They build baskets.
The goal of this project was simple: Understand those baskets — and use that insight to drive better decisions.
Why This Was Interesting
At first glance, this seems like a straightforward analytics problem.
But a few challenges made it more nuanced:
1. High Product Variety
- Hundreds of SKUs
- Many low-frequency items
2. Sparse Combinations
- Not every product pair appears frequently
- Long-tail combinations dominate
3. No Explicit Customer Data
- Only transaction-level data
- No user-level tracking
So the system had to rely purely on:
Patterns within transactions — not users.
Design Philosophy
- Focus on practical associations, not statistical novelty
- Prioritize high-impact combinations, not all combinations
- Ensure outputs are directly actionable
Architecture
Data Layer
- Transaction-level sales data
- Each bill mapped to list of purchased SKUs
Processing Layer (BigQuery)
- Basket creation (grouping SKUs per transaction)
Analysis Layer (Python)
- Association rule mining
- Support, confidence, lift calculations
Output Layer
- Top product combinations
- Cross-sell recommendations
Core Approach
1. Basket Construction
Each transaction was converted into a basket:
basket = ["chocolate_A", "candy_B", "drink_C"]
2. Association Rules
The system evaluated relationships like:
“If a customer buys A, how likely are they to buy B?”
Key metrics:
- Support → how often the combination appears
- Confidence → probability of B given A
- Lift → strength of association beyond random chance
3. Filtering Useful Rules
Not all associations are useful.
Filtered by:
- Minimum support threshold
- High confidence
- Lift > 1 (meaning meaningful relationship)
4. Actionable Pairing
Instead of generating hundreds of rules, the system outputs:
- Top cross-sell pairs
- Bundle suggestions
- Placement recommendations
Key Decisions
Why Not Personalization?
- No user-level data
- Store operates mostly offline
So focus stayed on:
- Aggregate behavior → store-level optimization
Why Focus on Lift?
High frequency alone is misleading.
Lift ensures:
- The relationship is meaningful
- Not just due to popularity of individual items
Why Keep Output Small?
Too many rules = no adoption.
Instead:
- Focused on top ~10–20 actionable insights
Tradeoffs
- No personalization → less targeted recommendations
- Aggregate analysis → misses individual preferences
- Threshold filtering → may miss niche but valuable combos
Results
Operational
- Identified strong product pairings
- Improved product placement strategy
- Enabled simple bundling ideas
Business
- Increased average basket value
- Encouraged impulse purchases
- Better utilization of shelf space
What This Enables Next
- Bundle pricing strategies
- Recommendation systems (online)
- Store layout optimization
- Campaign design (combo offers)
Takeaway
Customers don’t buy products — they buy combinations.
Understanding those combinations turns raw sales data into:
- Better placement
- Smarter bundling
- Higher revenue per transaction
Even simple association rules can unlock insights that are immediately usable.