Preference Learning and Hybrid Combinatorial Optimization for Intelligent Group-Buying Platforms: A Prototype-Calibrated Study of SmartBuy Connect
Abstract
1. Introduction
2. Literature Review
3. Architecture and Software Implementation of SmartBuy Connect
3.1. Architectural Requirements
3.2. Lifecycle of a Group Purchase
3.3. Service Layer and Separation of Responsibilities
3.4. Data Layer
3.5. Intelligent Layer
3.6. Closed Decision Loop
3.7. Match Between the Architecture and the Research Model
4. Materials and Methods
4.1. Study Design
4.2. Problem Statement and Decision Variables
4.3. Building the Behavioral Signal
4.4. Learning User–Product Preferences
4.4.1. Latent Model
4.4.2. Social Component
4.4.3. Cold-Start Handling
4.5. From a Product Score to a Lot Utility
4.6. Complete PA-HCGO Optimization Problem
4.6.1. Objective Function
4.6.2. Lot Activation and Capacity
4.6.3. Spending Constraint
4.6.4. Multi-Group Participation Constraint
4.6.5. Several Lots of the Same Product
4.6.6. Product Stock Constraint
4.6.7. Assignment Feasibility
4.6.8. Complete Mathematical Formulation
4.7. Pre-Reduction in the Solution Space
Illustrative Example of Candidate Set Construction
4.8. Hybrid Solution Procedure
| Algorithm 1. PA-HCGO solution procedure |
| Input: User set ; lot set ; product set ; learned utilities ; feasibility indicators ; initial lot participants ; minimum group sizes ; lot capacities ; product stocks ; joining costs ; user spending limits ; maximum simultaneous participations ; utility threshold ; activation weight ; exact/heuristic switching threshold ; solver time limit ; and maximum number of local search iterations . Output: Assignment matrix , activation vector , and objective value . 1. Construct the reduced candidate set: 2. If , return , , and . 3. Compute the normalization constant: 4. If , submit the binary integer program in Section 4.6.8 to the exact solver with time limit . 5. If the exact solver proves optimality within , return the optimal , , and . 6. If the exact solver does not prove optimality within , store the best feasible solution and solver gap. If a feasible incumbent exists, use it as the initial solution for the heuristic mode. 7. Initialize , , remaining user budgets , remaining user participation slots , remaining lot capacities , and remaining product stocks . 8. For each lot , compute its current deficit: 9. For each lot , compute an activation priority score: where is the average utility of the best available candidate users for lot . 10. Sort lots in descending order of . 11. For each lot in the sorted list: 11.1. Select candidate users with , which are sorted by decreasing . 11.2. Add the highest-utility feasible users to lot while all constraints remain satisfied. 11.3. Stop adding users to when the lot reaches its activation threshold, reaches capacity, or no feasible candidate remains. 11.4. If set . 11.5. If lot does not reach the activation threshold after the tentative additions, set for all tentative assignments to this lot and restore the corresponding budget, participation slot, capacity, and stock counters. 12. Utility-based completion stage. For each remaining candidate pair not yet selected, compute the marginal objective gain: where if adding user makes lot activatable, and otherwise. 13. Sort remaining candidate pairs by decreasing . 14. For each remaining candidate pair : 14.1. If assigning user to lot preserves all feasibility constraints and , set . 14.2. Update remaining budget, user participation slots, lot capacity, product stock, and activation status . 15. Local improvement stage. Repeat for at most iterations: 15.1. Generate feasible add, drop, and swap moves over selected and unselected candidate pairs. 15.2. For each move, compute the change in the full objective . 15.3. Apply the best feasible move if it increases . 15.4. Stop if no improving move exists. 16. Recompute each activation variable: 17. Return , , and . |
4.9. Computational Complexity of PA-HCGO
4.10. Reproducibility and Ethics
4.11. Section Summary
5. Experimental Base and Evaluation Protocol
5.1. Research Questions
5.2. Data Origin and Limits of Applicability
5.3. Structure of the Prototype Dataset
5.4. Distribution of Behavioral Events
5.5. Transactional Outcomes
5.6. Data Preprocessing
- Checking the uniqueness of primary identifiers;
- Checking referential integrity between tables;
- Removing exact duplicate events;
- Converting timestamps to a single time zone;
- Checking the logical order of lot and order states;
- Checking that prices, group sizes, and capacities are valid;
- Encoding categorical features;
- Building aggregated features only from events available at the prediction moment.
5.7. Lot Timestamp Consistency and Use of Lot Context Features
5.8. Temporal Split
5.9. Calibrated Scaling Scenarios
5.10. Compared Methods
5.11. Parameter Settings
5.12. Recommendation Quality Metrics
5.13. Group Formation Metrics
5.14. Computational Efficiency Metrics
5.15. Ablation and Sensitivity Analysis Protocol
- Without the referral-based social component: the model uses behavioral MF-BPR and content/cold-start scores;
- Without the content/cold-start component: the model uses only the behavioral MF-BPR and referral-based social scores;
- Without the behavioral MF-BPR component: the model uses the referral-based social and content/cold-start scores;
- Popularity-only preference scoring.
5.16. Statistical Testing
5.17. Section Summary
6. Results
6.1. Data Checking and Preparation Results
6.2. Quality of the Recommendation Layer
6.3. Group Formation in the Prototype Scenario
6.4. Scaled Scenario Results
6.5. Comparison with the Exact Solution
6.6. Computational Scalability

6.7. Ablation and Sensitivity Analysis Results
6.8. Answers to the Research Questions
6.9. Section Summary
7. Discussion
7.1. Interpreting the Results Against the Research Gap
7.2. Theoretical Contribution
7.3. Practical Relevance for a Digital Platform
7.4. Architectural and Deployment Implications
7.5. Interpreting the Individual Model Components
7.6. Threats to Validity
7.7. From a Static to a Learned User–Lot Utility
7.8. Managerial and Platform–Finance Implications
7.9. Directions for Further Research
7.10. Section Summary
8. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AU | Average Utility of Assignments |
| BIP | Binary Integer Programming |
| GCR | Group Completion Rate |
| GNN | Graph Neural Network |
| ILP | Integer Linear Programming |
| MF | Matrix Factorization |
| NDCG | Normalized Discounted Cumulative Gain |
| PA-HCGO | Preference-Aware Hybrid Combinatorial Group Optimization |
| SLA | Service-Level Agreement |
References
- Anand, K.S.; Aron, R. Group-Buying on the Web: A Comparison of Price-Discovery Mechanisms. Manag. Sci. 2003, 49, 1546–1562. [Google Scholar] [CrossRef] [Scilit]
- Jing, X.; Xie, J. Group Buying: A New Mechanism for Selling through Social Interactions. Manag. Sci. 2011, 57, 1354–1372. [Google Scholar] [CrossRef] [Scilit]
- Yamamoto, J.; Sycara, K.A. Stable and Efficient Buyer Coalition Formation Scheme for E-Marketplaces. In Proceedings of the Fifth International Conference on Autonomous Agents, Montreal, QC, Canada, 28 May–1 June 2001; ACM Press: New York, NY, USA, 2001; pp. 576–583. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Sycara, K.; Scheller-Wolf, A. Combinatorial Coalition Formation for Multi-Item Group-Buying with Heterogeneous Customers. Decis. Support Syst. 2010, 49, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Hsieh, F.-S.; Lin, J.-B. Assessing the benefits of group-buying-based combinatorial reverse auctions. Electron. Commer. Res. Appl. 2012, 11, 407–419. [Google Scholar] [CrossRef] [Scilit]
- Kauffman, R.J.; Lai, H.; Ho, C.-T. Incentive Mechanisms, Fairness and Participation in Online Group-Buying Auctions. Electron. Commer. Res. Appl. 2010, 9, 249–262. [Google Scholar] [CrossRef] [Scilit]
- Liang, T.-P.; Turban, E. Introduction to the Special Issue Social Commerce: A Research Framework for Social Commerce. Int. J. Electron. Commer. 2011, 16, 5–14. [Google Scholar] [CrossRef] [Scilit]
- Stephen, A.T.; Toubia, O. Deriving Value from Social Commerce Networks. J. Mark. Res. 2010, 47, 215–228. [Google Scholar] [CrossRef] [Scilit]
- Bawack, R.E.; Wamba, S.F.; Carillo, K.D.A.; Akter, S. Artificial Intelligence in E-Commerce: A Bibliometric Study and Literature Review. Electron. Mark. 2022, 32, 297–338. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Koren, Y.; Bell, R.; Volinsky, C. Matrix Factorization Techniques for Recommender Systems. Computer 2009, 42, 30–37. [Google Scholar] [CrossRef] [Scilit]
- Rendle, S.; Freudenthaler, C.; Gantner, Z.; Schmidt-Thieme, L. BPR: Bayesian Personalized Ranking from Implicit Feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, 18–21 June 2009; pp. 452–461. [Google Scholar]
- He, X.; Liao, L.; Zhang, H.; Nie, L.; Hu, X.; Chua, T.-S. Neural Collaborative Filtering. In Proceedings of the 26th International Conference on World Wide Web, Perth, Australia, 3–7 April 2017; pp. 173–182. [CrossRef] [Scilit]
- Zhang, S.; Yao, L.; Sun, A.; Tay, Y. Deep Learning Based Recommender System: A Survey and New Perspectives. ACM Comput. Surv. 2019, 52, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations, Toulon, France, 24–26 April 2017. [Google Scholar]
- Hamilton, W.L.; Ying, R.; Leskovec, J. Inductive Representation Learning on Large Graphs. In Advances in Neural Information Processing Systems; NeurIPS: Sydney, Australia, 2017; Volume 30, pp. 1024–1034. [Google Scholar]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph Attention Networks. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Wang, X.; He, X.; Wang, M.; Feng, F.; Chua, T.-S. Neural Graph Collaborative Filtering. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Paris, France, 21–25 July 2019; pp. 165–174. [Google Scholar] [CrossRef] [Scilit]
- He, X.; Deng, K.; Wang, X.; Li, Y.; Zhang, Y.; Wang, M. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, China, 25–30 July 2020; pp. 639–648. [Google Scholar] [CrossRef] [Scilit]
- Fan, W.; Ma, Y.; Li, Q.; He, Y.; Zhao, E.; Tang, J.; Yin, D. Graph Neural Networks for Social Recommendation. In Proceedings of the World Wide Web Conference, San Francisco, CA, USA, 13–17 May 2019; pp. 417–426. [Google Scholar] [CrossRef] [Scilit]
- Wu, S.; Sun, F.; Zhang, W.; Xie, X.; Cui, B. Graph Neural Networks in Recommender Systems: A Survey. ACM Comput. Surv. 2023, 55, 97. [Google Scholar] [CrossRef] [Scilit]
- Gao, C.; Zheng, Y.; Li, N.; Li, Y.; Qin, Y.; Piao, J.; Quan, Y.; Chang, J.; Jin, D.; He, X.; et al. A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions. ACM Trans. Recomm. Syst. 2023, 1, 1–51. [Google Scholar] [CrossRef] [Scilit]
- Jameson, A.; Smyth, B. Recommendation to Groups. In The Adaptive Web: Methods and Strategies of Web Personalization; Brusilovsky, P., Kobsa, A., Nejdl, W., Eds.; Springer: Berlin/Heidelberg, Germany, 2007; pp. 596–627. [Google Scholar] [CrossRef] [Scilit]
- Amer-Yahia, S.; Roy, S.B.; Chawla, A.; Das, G.; Yu, C. Group Recommendation: Semantics and Efficiency. Proc. VLDB Endow. 2009, 2, 754–765. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; He, X.; Miao, L.; Xiao, G.; Chen, H.; Xu, J. Attentive Group Recommendation. In Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, Ann Arbor, MI, USA, 8–12 July 2018; pp. 645–654. [Google Scholar] [CrossRef] [Scilit]
- Schein, A.I.; Popescul, A.; Ungar, L.H.; Pennock, D.M. Methods and Metrics for Cold-Start Recommendations. In Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Tampere, Finland, 11–15 August 2002; pp. 253–260. [Google Scholar] [CrossRef] [Scilit]
- Volkovs, M.; Yu, G.W.; Poutanen, T. DropoutNet: Addressing Cold Start in Recommender Systems. In Advances in Neural Information Processing Systems; NeurIPS: Sydney, Australia, 2017; Volume 30, pp. 4957–4966. [Google Scholar]
- Wolsey, L.A. Integer Programming; Wiley: New York, NY, USA, 1998. [Google Scholar]
- Blum, C.; Roli, A. Metaheuristics in Combinatorial Optimization: Overview and Conceptual Comparison. ACM Comput. Surv. 2003, 35, 268–308. [Google Scholar] [CrossRef] [Scilit]
- Bengio, Y.; Lodi, A.; Prouvost, A. Machine Learning for Combinatorial Optimization: A Methodological Tour d’Horizon. Eur. J. Oper. Res. 2021, 290, 405–421. [Google Scholar] [CrossRef] [Scilit]
- Kotary, J.; Fioretto, F.; Van Hentenryck, P.; Wilder, B. End-to-End Constrained Optimization Learning: A Survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada, 19–27 August 2021; pp. 4475–4482. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mazyavkina, N.; Sviridov, S.; Ivanov, S.; Burnaev, E. Reinforcement Learning for Combinatorial Optimization: A Survey. Comput. Oper. Res. 2021, 134, 105400. [Google Scholar] [CrossRef] [Scilit]
- Talbi, E.-G. Machine Learning into Metaheuristics: A Survey and Taxonomy. ACM Comput. Surv. 2022, 54, 129. [Google Scholar] [CrossRef] [Scilit]
- Nowak, M.; Pawłowska-Nowak, M. Dynamic Pricing Method in the E-Commerce Industry Based on Machine Learning. Appl. Sci. 2024, 14, 11668. [Google Scholar] [CrossRef] [Scilit]
- Chenavaz, R.Y.; Dimitrov, S. Artificial Intelligence and Dynamic Pricing: A Systematic Literature Review. J. Appl. Econ. 2025, 28, 2466140. [Google Scholar] [CrossRef] [Scilit]
- Platt, J.C. Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods. In Advances in Large Margin Classifiers; Smola, A.J., Bartlett, P.L., Schölkopf, B., Schuurmans, D., Eds.; MIT Press: Cambridge, MA, USA, 1999; pp. 61–74. [Google Scholar]
- Järvelin, K.; Kekäläinen, J. Cumulated Gain-Based Evaluation of IR Techniques. ACM Trans. Inf. Syst. 2002, 20, 422–446. [Google Scholar] [CrossRef] [Scilit]
- Virtanen, P.; Gommers, R.; Oliphant, T.E.; Haberland, M.; Reddy, T.; Cournapeau, D.; Burovski, E.; Peterson, P.; Weckesser, W.; Bright, J.; et al. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nat. Methods 2020, 17, 261–272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Perron, L.; Furnon, V. OR-Tools. Google. 2019. Available online: https://developers.google.com/optimization (accessed on 18 June 2026).
- Efron, B.; Tibshirani, R.J. An Introduction to the Bootstrap; Chapman & Hall/CRC: New York, NY, USA, 1993. [Google Scholar]
- Harris, C.R.; Millman, K.J.; van der Walt, S.J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N.J.; et al. Array Programming with NumPy. Nature 2020, 585, 357–362. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-Learn: Machine Learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Lam, S.K.; Pitrou, A.; Seibert, S. Numba: A LLVM-Based Python JIT Compiler. In Proceedings of the Second Workshop on the LLVM Compiler Infrastructure in HPC, Austin, TX, USA, 15 November 2015; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]






| Direction/Source | Main Idea | Strength | Limitation | How PA-HCGO Differs |
|---|---|---|---|---|
| Anand and Aron [1] | Group buying as a price-discovery mechanism | Explains the economics of collective demand | Does not solve personalized distribution of users across lots | Uses group-buying logic as a base, but adds learning and optimization |
| Jing and Xie [2] | Group buying through social interaction | Shows the role of invitations and social interaction | Does not formalize a constrained assignment problem | Treats social and referral links as part of utility learning |
| Yamamoto and Sycara [3] | Buyer coalition formation in an e-marketplace | Addresses coalition formation among buyers | Limited link to modern AI recommendation | Extends the coalition idea through preference-aware optimization |
| Li et al. [4] | Multi-item group buying with heterogeneous buyers | Accounts for differences in preferences and products | Does not use learned behavioral utility | Replaces fixed utility with learned scores |
| Hsieh and Lin [5] | Group-buying-based combinatorial reverse auction | Provides a coordinated procurement model and solution algorithms | Procurement-oriented setting with fixed requirements; no learned consumer preferences | Learns individual user–lot utility and assigns consumers to active lots under platform constraints |
| Bawack et al. [9] | AI in e-commerce literature review | Shows the systemic role of AI in e-commerce | Offers no concrete group-buying model | Provides a concrete model, data, and architecture |
| Koren et al. [10] | Matrix factorization for recommendation | Model’s latent preferences | Ignores the hard constraints of a group deal | Uses preference learning as input to the optimization engine |
| Rendle et al. [11] | BPR for implicit feedback | Suits clicks, views, and purchases | Optimizes ranking, not group activation | Turns behavioral signals into utility scores for constrained assignment |
| He et al. [12] | Neural Collaborative Filtering | Captures non-linear user–item interactions | Does not guarantee budget, capacity, or group-size constraints | Integrates learned scores into BIP and heuristic optimization |
| Wang et al. [17] | NGCF for the user–item graph | Captures high-order graph relations | Focused on recommendation accuracy | Uses the graph signal to improve lot formation |
| He et al. [18] | LightGCN | An efficient graph recommendation model | Does not address group activation | Uses graph proximity as a component of utility estimation |
| Fan et al. [19] | GraphRec for social recommendation | Accounts for user–user and user–item graphs | Does not model group-buying constraints | Includes the social graph in a constrained group-buying framework |
| Jameson and Smyth [22] | Group recommendation | Selects an item for an existing group | The group is fixed in advance | PA-HCGO forms groups, not only recommends an item to a group |
| Cao et al. [24] | Attentive group recommendation | Accounts for unequal member contribution | Does not solve multi-lot assignment | Distributes users across many lots |
| Bengio et al. [29] | ML for combinatorial optimization | Justifies learning-augmented optimization | Not aimed specifically at group buying | Adapts this logic to intelligent group buying |
| Kotary et al. [30] | End-to-end constrained optimization learning | Links ML and constrained optimization | General in nature, no e-commerce platform case | Provides an applied framework for SmartBuy Connect |
| Mazyavkina et al. [31] | RL for combinatorial optimization | Shows the potential of learned heuristics | RL is hard to deploy reproducibly at an early stage | Uses an interpretable ILP-plus-heuristic hybrid |
| Nowak and Pawłowska-Nowak [33] | Dynamic pricing in e-commerce | Shows the role of AI in adaptive pricing | Focused on price, not on group formation | Treats pricing as a future extension |
| This study | PA-HCGO + SmartBuy Connect | Integrates preference learning, content/cold-start scoring, an optional exploratory referral signal, and constrained optimization | Needs further validation using industrial data | Proposes a prototype-calibrated three-track evaluation design |
| Symbol | Meaning |
|---|---|
| Set of users | |
| Set of products | |
| Set of active group-buying lots | |
| Index of a user, | |
| Index of a product, | |
| Index of a lot, | |
| Product linked to lot | |
| Initial number of participants in lot before the current assignment decision | |
| Minimum number of participants required to activate lot | |
| Maximum capacity of lot | |
| Available stock of product | |
| Joining cost for user if assigned to lot | |
| Experimental spending limit of user | |
| Maximum number of simultaneous group participations allowed for user | |
| Learned utility score for user and lot | |
| Feasibility indicator for assigning user to lot | |
| Binary assignment variable equal to 1 if user is recommended lot | |
| Binary activation variable equal to 1 if lot is algorithmically activatable | |
| Final number of participants in lot after assignment | |
| Minimum utility threshold for candidate inclusion | |
| Reduced candidate set of feasible user–lot pairs | |
| Number of reduced candidate user–lot pairs | |
| Normalization constant for assignment utility | |
| Lot-specific Big-M activation-linking constant | |
| Weight controlling the trade-off between assignment utility and lot activation | |
| Empirical switching threshold between exact and heuristic modes | |
| Maximum number of local search iterations | |
| Number of users | |
| Number of active lots | |
| Number of products | |
| Number of binary variables in the reduced BIP | |
| Number of constraints in the reduced BIP | |
| Set of observed events of user related to product | |
| Individual user event | |
| Event type of event | |
| Timestamp of event | |
| Sampling time | |
| Weight assigned to the event type | |
| Temporal decay coefficient | |
| Aggregated behavioral signal of user for product | |
| Latent vector of user | |
| Latent vector of product | |
| User bias | |
| Product bias | |
| Predicted latent compatibility score | |
| Observed positive product for user | |
| Sampled unobserved product used as a negative example | |
| Set of sampled BPR training triples | |
| Set of trainable model parameters | |
| Regularization coefficient | |
| Training loss function | |
| Internal mixing weight between behavioral MF-BPR and referral-based social scores | |
| Number of historical interactions of user | |
| Minimum history threshold for trusting collaborative information | |
| Accumulated history coefficient for user | |
| Relative discount of lot | |
| Fill-level feature of lot | |
| Urgency feature of lot | |
| Parameters of the lot context utility model | |
| Activation deficit of lot before additional assignment | |
| Global mixing weight between the collaborative/social block and the content score | |
| Number of distinct user–product combinations induced by the reduced candidate set |
| Event Type | Weight | Interpretation |
|---|---|---|
| product_view | 1.0 | Weak positive signal |
| search | 1.2 | Weak-to-moderate interest signal |
| product_share | 1.5 | Socially expressed interest |
| lot_create | 2.5 | Strong intent to initiate group purchase |
| lot_join | 3.0 | Strong participation intent |
| checkout | 4.0 | Very strong purchase intent |
| completed/confirmed/processing order | 5.0 | Strongest positive transactional signal |
| lot_leave | −2.0 | Negative or weakening signal |
| Formula | What It Computes | Why It is Needed | Example Interpretation |
|---|---|---|---|
| Formula (9) | Weighted time-decayed behavioral signal | Converts heterogeneous implicit events into a comparable user–product signal | Recent lot joins contribute more than old product views |
| Formula (14) | Latent compatibility score | Places users and products in a shared preference space | A higher dot product means stronger compatibility |
| Formula (16) | BPR pairwise ranking loss | Trains the model to rank observed products above sampled unobserved products | Product should score higher than product for the same user |
| Lot | Product | Initial Participants | Minimum Group Size | Capacity | Joining Cost |
|---|---|---|---|---|---|
| 1 | 2 | 3 | 20 | ||
| 0 | 2 | 2 | 15 |
| User–Lot Pair | Learned Utility | Feasibility | Included in ? |
|---|---|---|---|
| 0.82 | 1 | Yes | |
| 0.45 | 1 | Yes | |
| 0.76 | 1 | Yes | |
| 0.55 | 1 | Yes | |
| 0.30 | 1 | No, because | |
| 0.63 | 1 | Yes |
| Solution | Selected Assignments | Activated Lots | Sum of Utility | Interpretation |
|---|---|---|---|---|
| A | 1 | 1.58 | High utility for , but remains inactive | |
| B | 2 | 2.00 | Activates both lots and better satisfies group-buying logic |
| Stage | Complexity | Explanation |
|---|---|---|
| Candidate set construction | Worst-case check of all user–lot pairs before filtering | |
| Reduced BIP variables | One variable for each and one variable for each lot | |
| Reduced BIP constraints | Constraints depend on candidate pairs, users, lots, products, and user–product combinations in | |
| Candidate sorting | Sorting candidate pairs by utility or marginal objective gain | |
| Lot-priority sorting | Sorting lots by activation priority | |
| Activation-oriented initialization | after sorting | Each candidate pair is checked at most a bounded number of times using running feasibility counters |
| Utility-based completion | after sorting | Remaining candidate pairs are added only if they preserve feasibility and improve the objective |
| Bounded local search | Prototype implementation uses a bounded list of feasible add, drop, and swap moves | |
| Exhaustive swap upper bound | Conservative upper bound if all pairwise swaps are examined | |
| Memory requirement | Storage for candidate pairs, utilities, assignments, lot states, user counters, and product stock counters |
| Table | Main Content | Records | Role in the Study |
|---|---|---|---|
| Sellers | Anonymized seller profiles | 15 | Linking products and orders to suppliers |
| Products | Product cards, categories, prices, and sellers | 200 | Forming the product set and product features |
| Users | Anonymized user profiles | 150 | Forming the user set and context features |
| Lots | Parameters and states of group lots | 150 | Activation thresholds, capacity, and current group size |
| User Events | Views, search, lot participation, and checkout | 4000 | Training on implicit feedback |
| Orders | Transactional outcomes and fulfillment states | 500 | Checking conversion and the strength of the positive signal |
| User Preferences | Aggregated user–category scores | 1415 | Content and cold-start components |
| Referral Links | Directed links between users | 200 | Building the social graph |
| Indicator | Value | Interpretation |
|---|---|---|
| Users | 150 | Nodes of the referral graph |
| Directed referral links | 200 | Observed directed referral relations |
| Directed graph density | 0.009 | Sparse directed graph |
| Undirected graph density | 0.018 | Sparse weak-tie structure |
| Average out-degree | 1.33 | Low referral activity per user |
| Average in-degree | 1.33 | Low average received referrals |
| Isolated users | 12 | Users with no referral connection |
| Weakly connected components | 16 | Fragmented graph structure |
| Largest weak component | 132 users | Most users are connected through one weak component |
| Maximum total degree | 10 | Few highly connected users |
| Record Group | Number of Records | Used for Product Preference Layer | Used for Dynamic Lot Context Layer | Interpretation |
|---|---|---|---|---|
| User events with a lot identifier | 1060 | Yes, if user and product identifiers are valid | Not in the reported experiments; used only for timestamp consistency diagnostics. | Lot identifier exists, but temporal validity must be verified. |
| Temporally consistent lot-linked events | 79 | Yes | Diagnostic assessment only. Not used to estimate the reported coefficients. | Only subset supporting time-dependent lot context analysis. |
| Lot-linked events outside recorded lot window | 981 | Yes, as user–product interest signals | No | Retained for preference learning, excluded from fill-level and urgency estimation. |
| Discount information from product/lot tables | 150 lots | Not applicable | Yes, as a static attribute. | Does not require reconstruction of the lot state over time. |
| Fill-level and urgency variables | 150 lots/79 consistent events | Not applicable | Not in the reported experiments. | Retained in the formal model for future evaluation after a corrected event-level lot timeline becomes available. |
| Scenario | Users | Products | Active Lots | Demand Character | Purpose |
|---|---|---|---|---|---|
| P0 | 150 | 200 | 150 | Prototype-calibrated counterfactual pre-activation state | Counterfactual group activation analysis |
| S1 | 500 | 300 | 250 | Sparse demand | Robustness when suitable participants are scarce |
| S2 | 1000 | 500 | 500 | Moderate demand | Comparison of the exact and hybrid modes |
| S3 | 5000 | 1000 | 1000 | High load | Testing computational scalability |
| Symbol | Method | Short Description |
|---|---|---|
| POP | Popularity ranking | Ranking products by the frequency of positive interactions, without personalization |
| MF-BPR | Matrix factorization with BPR | Personalized ranking from implicit feedback |
| MF-SOC | MF-BPR with a social component | Preference model with the referral graph, without group optimization |
| GREEDY | Greedy feasible assignment | Sequential choice of feasible pairs by decreasing utility |
| HCGO-STATIC | Hybrid combinatorial optimization | Optimization with a fixed, manually set utility |
| ILP-PA | Exact preference-aware ILP | Exact solution of the learned utility model for feasibly sized instances |
| PA-HCGO | Proposed method | Learned user–lot utility, filtering, an exact or heuristic mode, and local search |
| ITEM-KNN | Cosine item-based nearest-neighbor recommendation | Ranks candidate products using cosine item–item similarities computed from training-period implicit interactions; |
| Baseline Group | Included in This Study | Purpose | Advanced Alternatives Not Included | Reason for Exclusion/Future Work |
|---|---|---|---|---|
| Non-personalized ranking | POP | Tests whether simple popularity explains user activity | Category popularity, time-aware popularity | POP is sufficient as a lower-complexity popularity baseline |
| Latent personalization | MF-BPR | Tests implicit feedback personalization | NCF, factorization machines, gradient-boosted ranking | Future work; current focus is integration, not recommender benchmarking |
| Social preference signal | MF-SOC | Tests whether referral information improves ranking | LightGCN, NGCF, GraphSAGE, GraphRec | Prototype referral graph is small; advanced GNNs require denser graph data |
| Feasible assignment | GREEDY | Tests independent assignment under constraints | Coverage heuristics, matching heuristics | GREEDY isolates the cost of ignoring global group activation |
| Static constrained optimization | HCGO-STATIC | Tests constrained assignment with fixed utility | Min-cost flow, Lagrangian relaxation | Useful future solvers, but not central to the learned utility contribution |
| Exact reference | ILP-PA | Gives optimal solutions for small instances | Commercial-solver warm starts | Exact comparison is limited to small subproblems |
| Neighborhood recommendation | ITEM-KNN | Tests item-based similarity from training-period implicit interactions | User-KNN, SLIM, graph-based neighborhood models | ITEM-KNN provides a transparent low-cost neighborhood baseline; more advanced neighborhood models require a broader recommender benchmark |
| Parameter | Meaning | Base Value | Selection/Role |
|---|---|---|---|
| Temporal decay coefficient in implicit feedback | 0.05 per day | Selected based on the validation interval | |
| Latent embedding dimension | 16 | Used in MF-BPR preference learning | |
| Regularization coefficient | 0.002 | Used in pairwise MF-BPR training | |
| Learning rate | Initial learning rate | 0.02 | Used in MF-BPR training |
| Epochs | Number of pairwise training epochs | 60 | Fixed after validation |
| Negative samples | Negatives per positive pair | 5 | Used in BPR training |
| Graph layers | LightGCN propagation depth | 2 | Social/referral component |
| Cold-start history threshold | 5 | Controls transition from content to collaborative score | |
| Utility threshold for candidate filtering | 0.40 | Base feasibility/relevance threshold | |
| Activation weight in objective | 0.4 | Controls utility–activation trade-off | |
| Exact/heuristic switching threshold | 2000 candidate pairs | Determines solution mode | |
| ILP time limit | Exact solver time limit | 60 s | Used in the exact comparison protocol |
| Bootstrap resamples | Number of bootstrap repetitions | 10,000 | Used for uncertainty intervals |
| MF-BPR runs | Number of independent MF-BPR training runs with different random seeds; final score is the average of the prediction matrices | 5 runs, seeds 101–105 | Reduces seed-dependent variance of the preference layer |
| Global mixing weight between content and collaborative/social scores | 0.5 | Selected based on the validation interval | |
| Internal mixing weight between behavioral MF-BPR and referral-based social scores | 0.8 | Gives effective weights 0.4 MF-BPR, 0.1 social, and 0.5 content | |
| Number of item neighbors | 50 | ITEM-KNN baseline | |
| Content profile mode | Temporal handling of user preferences table | As of training boundary | 952 of 1415 records retained |
| Content profile cutoff | Latest allowed updated_at timestamp | 13 February 2026 | Prevents post-training profile information |
| Lot utility intercept | Fitted on training interval | ||
| Preference coefficient | 5.4668 | Fitted on training interval | |
| Discount coefficient | Fitted on training interval | ||
| Fill-level coefficient | 0, inactive | Not estimated | |
| Urgency coefficient | 0, inactive | Not estimated | |
| Maximum number of local search iterations | 100 | Bonus heuristic refinement | |
| Experimental spending limit in P0 | 40,000 monetary units | Applied uniformly to all users | |
| Maximum simultaneous participations in P0 | 3 lots | Applied uniformly to all users | |
| P0 initial deficit | Initial pre-activation rule | 0.50 | |
| Joining cost | Discounted price of the product linked to lot | Same construction for all methods |
| Setting | Value Used in the Experiment | Purpose |
|---|---|---|
| Number of exact comparison subproblems | 20 | Estimating heuristic gap on reproducible small instances |
| Users per subproblem | 50 | Keeps exact ILP solvable within the time limit |
| Lots per subproblem | 20 | Keeps binary assignment problem interpretable |
| Candidate filtering | Same filtered candidate set for ILP-PA and PA-HCGO | Ensures fair comparison of exact and heuristic solution modes |
| Objective function | Normalized PA-HCGO objective in Equation (35) | Same criterion for exact and heuristic solutions |
| Solver | SciPy optimize.milp with HiGHS | Exact backend used for the reported v1.2.0 results |
| Time limit per subproblem | 60 s | Prevents unbounded exact search |
| Relative optimality tolerance | mip_rel_gap = 0.0 | Requests proven optimality |
| Random seeds | 1–20 | Reproducibility |
| Reported exact cases | All 20 cases with proven optimum | No unfinished run treated as an optimum |
| Indicator | Value | Use in the Experiment |
|---|---|---|
| Total events and orders | 4500 | Building the temporal sequence of interactions |
| Training interval | 3150 records (70%) | Training the models and building user profiles |
| Validation interval | 675 records (15%) | Choosing hyperparameters and component weights |
| Test interval | 675 records (15%) | Final evaluation without further tuning |
| Test user–product interactions | 580 items for 148 users | Diagnostic evaluation of general engagement |
| Test transactional items | 40 items for 36 users | Main evaluation of purchase intent |
| Events with a lot identifier | 1060 | Checking temporal consistency |
| Temporally consistent lot events | 79 | Lot context analysis |
| Target Event | Method | Precision@10 | Recall@10 | NDCG@10 | Hit Rate@10 |
|---|---|---|---|---|---|
| General engagement | POP | 0.0385 | 0.1024 | 0.0802 | 0.3243 |
| General engagement | MF-BPR | 0.0318 | 0.0942 | 0.0619 | 0.2770 |
| General engagement | MF-SOC | 0.0277 | 0.0817 | 0.0544 | 0.2432 |
| General engagement | CONTENT | 0.0324 | 0.0776 | 0.0524 | 0.2905 |
| General engagement | PA-PREF | 0.0324 | 0.0804 | 0.0515 | 0.2973 |
| General engagement | ITEM-KNN | 0.0304 | 0.0837 | 0.0545 | 0.2703 |
| Transactional event | POP | 0.0083 | 0.0694 | 0.0259 | 0.0833 |
| Transactional event | MF-BPR | 0.0083 | 0.0694 | 0.0266 | 0.0833 |
| Transactional event | MF-SOC | 0.0056 | 0.0417 | 0.0153 | 0.0556 |
| Transactional event | CONTENT | 0.0278 | 0.2639 | 0.1034 | 0.2778 |
| Transactional event | PA-PREF | 0.0389 | 0.3611 | 0.1319 | 0.3889 |
| Transactional event | ITEM-KNN | 0.0111 | 0.0833 | 0.0323 | 0.1111 |
| Method | AU | GCR | Coverage | Assignments | Active Lots | Objective | Time, s |
|---|---|---|---|---|---|---|---|
| GREEDY | 0.6841 | 0.5667 | 1.0000 | 405 | 85 | 0.5961 | 0.0028 |
| HCGO-STATIC | 0.5632 | 0.8133 | 1.0000 | 403 | 122 | 0.6280 | 0.0356 |
| PA-HCGO | 0.6710 | 0.7800 | 1.0000 | 396 | 117 | 0.6663 | 0.0158 |
| Scenario | Method | AU | GCR | Coverage | Assignments | Objective | Time, s |
|---|---|---|---|---|---|---|---|
| S1 | GREEDY | 0.7418 ± 0.0069 | 0.7512 ± 0.0252 | 0.9987 | 1175.15 | 0.6491 | 0.0167 |
| S1 | HCGO-STATIC | 0.5598 ± 0.0048 | 0.8110 ± 0.0306 | 0.9986 | 1227.85 | 0.5993 | 0.0937 |
| S1 | PA-HCGO | 0.7339 ± 0.0059 | 0.8112 ± 0.0297 | 0.9986 | 1185.70 | 0.6726 | 0.0986 |
| S2 | GREEDY | 0.7653 ± 0.0049 | 0.7502 ± 0.0164 | 0.9991 | 2376.15 | 0.6638 | 0.0810 |
| S2 | HCGO-STATIC | 0.5606 ± 0.0035 | 0.8167 ± 0.0163 | 0.9999 | 2476.70 | 0.6044 | 0.4466 |
| S2 | PA-HCGO | 0.7573 ± 0.0044 | 0.8171 ± 0.0163 | 0.9995 | 2393.75 | 0.6894 | 0.4453 |
| S3 | GREEDY | 0.8442 ± 0.0035 | 0.8108 ± 0.0133 | 0.6027 | 5613.00 | 0.7271 | 1.0750 |
| S3 | HCGO-STATIC | 0.5595 ± 0.0023 | 0.8227 ± 0.0131 | 0.4770 | 5610.60 | 0.5959 | 4.2706 |
| S3 | PA-HCGO | 0.8441 ± 0.0037 | 0.8229 ± 0.0128 | 0.6018 | 5610.95 | 0.7317 | 4.7174 |
| Indicator | ILP-PA | PA-HCGO | Interpretation |
|---|---|---|---|
| GCR | 0.8025 ± 0.0981 | 0.7850 ± 0.1050 | Heuristic GCR was 1.75 percentage points lower on average |
| AU | 0.5927 ± 0.0213 | 0.6075 ± 0.0231 | The heuristic selected slightly higher-utility assignments |
| Objective gap | 0 | 0.0316 ± 0.0243 | Mean relative gap was 3.16% |
| Mean solve time, s | 0.0133 | 0.0006 | Ratio of means ; mean per-instance ratio |
| 95th-percentile solve time, s | 0.0345 | 0.0008 | Tail latency of the two modes |
| Configuration | Recall@10 | NDCG@10 | Hit Rate@10 | NDCG Change vs. Full Model |
|---|---|---|---|---|
| Full PA-PREF | 0.3611 | 0.1319 | 0.3889 | 0.0% |
| Without the social component | 0.3611 | 0.1403 | 0.3889 | +6.3% |
| Without content/cold-start fallback | 0.0417 | 0.0153 | 0.0556 | −88.4% |
| Without the behavioral MF component | 0.3056 | 0.1202 | 0.3333 | −8.9% |
| Popularity only | 0.0694 | 0.0259 | 0.0833 | −80.4% |
| Participant | Practical Problem | How PA-HCGO Helps | Limit of the Interpretation |
|---|---|---|---|
| Buyer | Many irrelevant lots and uncertainty about group completion | Prioritizes lots by user interest and the current member deficit | A recommendation does not guarantee that the user actually wants to buy |
| Seller | Weak predictability of collective demand | Tops up groups to the minimum threshold and raises the share of activated lots | Needs testing on real commercial flows and stock |
| Platform | Conflict between personalization and business constraints | A single model that joins utility scores, budgets, capacity, and participations | The choice of and affects the balance between metrics |
| AI/Analytics team | Hard to evaluate the algorithm rigorously at an early prototype stage | Uses prototype-calibrated evaluation instead of a fully synthetic experiment | Needs further validation once industrial data accumulate |
| Threat Type | Potential Problem | How It Is Handled in This Paper | What Is Needed in Future Work |
|---|---|---|---|
| Internal validity | Temporal mismatch of part of the events with a lot identifier | Such events are excluded from the dynamic context analysis but kept as user–product signals | Synchronize the logging of lots, events, and orders in an industrial version |
| Construct validity | Views and searches are not equal to purchase intent | Metrics are split into general engagement and transactional events | Add explicit intent signals: cart, abandonment, return visit, payment |
| External validity | The dataset reflects a prototype, not mass operation | Observed, counterfactual, and synthetic evidence are reported as three explicitly separated evaluation tracks | Run a pilot with real users and sellers |
| Algorithmic validity | The referral-based social component provides no confirmed positive contribution on the current sparse graph | An ablation analysis is done and the limits of interpretation are stated | Repeat the test after extending the referral graph and the number of orders |
| Statistical validity | The transactional test set is small | Bootstrap intervals and cautious interpretation are used | Accumulate more orders and run A/B testing |
| Deployment validity | Single-digit second prototype solve times do not guarantee an industrial SLA | Results are named a characteristic of the prototype implementation | Run load testing, latency monitoring, and an API stress test |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kassymova, A.; Uskenbayeva, R.; Cho, Y.I.; Elle, V.; Anartayeva, A.; Smakhanova, A. Preference Learning and Hybrid Combinatorial Optimization for Intelligent Group-Buying Platforms: A Prototype-Calibrated Study of SmartBuy Connect. Information 2026, 17, 768. https://doi.org/10.3390/info17080768
Kassymova A, Uskenbayeva R, Cho YI, Elle V, Anartayeva A, Smakhanova A. Preference Learning and Hybrid Combinatorial Optimization for Intelligent Group-Buying Platforms: A Prototype-Calibrated Study of SmartBuy Connect. Information. 2026; 17(8):768. https://doi.org/10.3390/info17080768
Chicago/Turabian StyleKassymova, Aizhan, Raissa Uskenbayeva, Young Im Cho, Venera Elle, Aizhan Anartayeva, and Aizhan Smakhanova. 2026. "Preference Learning and Hybrid Combinatorial Optimization for Intelligent Group-Buying Platforms: A Prototype-Calibrated Study of SmartBuy Connect" Information 17, no. 8: 768. https://doi.org/10.3390/info17080768
APA StyleKassymova, A., Uskenbayeva, R., Cho, Y. I., Elle, V., Anartayeva, A., & Smakhanova, A. (2026). Preference Learning and Hybrid Combinatorial Optimization for Intelligent Group-Buying Platforms: A Prototype-Calibrated Study of SmartBuy Connect. Information, 17(8), 768. https://doi.org/10.3390/info17080768

