Nutrient-Aware Food Recommendation for Diabetes Using Item–Item Collaborative Filtering
DOI:
https://doi.org/10.58414/SCIENTIFICTEMPER.2026.17.9.2557Keywords:
collaborative filtering, item–item similarity, cosine similarity, food recommender system, diabetes diet management, nutrient-aware recommendation, synthetic rating data.Abstract
Diet management is an important part of diabetes care, but food recommendation is more demanding than ordinary item recommendation because a useful suggestion must satisfy both user preference and nutritional constraints. This study develops an item–item collaborative filtering (CF) model for diabetes-oriented food recommendation by combining rule-based nutrient filtering with cosine-similarity-based collaborative prediction. A nutrient composition dataset containing 12,044 food records and 29 raw attributes was cleaned and mapped to canonical nutrient fields. After removal of malformed records and duplicate food names, 1,781 unique foods remained; a sugar threshold of 12 g per reported unit reduced the candidate catalogue to 1,779 items. Because public data linking diabetic-user food preferences with detailed nutrient profiles were not available, 100 synthetic users and a controlled 1–5 rating matrix were generated from nutrient-weighted preference profiles with Gaussian noise. The resulting matrix contained 21,203 observed ratings at 88.08% sparsity. An 80:20 per-user train/test split was used, and item–item cosine similarity with an eight-neighbourhood was applied for prediction and Top-10 recommendation. Among 4,200 held-out ratings, predictions were available for 1,640 user–item pairs, producing RMSE = 0.9463 and MAE = 0.7101. Across 100 users, Precision@10 = 0.0091, Recall@10 = 0.0075, and HitRate@10 = 0.0909. The representative Top-10 list was dominated by low-recorded-sugar, protein- and fibre-rich foods. The findings support the feasibility of the proposed pipeline as a reproducible baseline, while the modest ranking performance, synthetic ratings, sparse matrix, mean-imputation strategy, and incomplete source fields limit clinical interpretation and motivate validation with real preference data.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 The Scientific Temper

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
