STAT3009 · Recommender Systems
CUHK · Department of Statistics and Data Science
User 1960 gave movie 670 four stars. The model may learn from this row.
The model sees the user and movie. The rating stays hidden until evaluation.
train has 51,161 observed triples with user, movie and rating.
Fit shared ID maps. Create X_train, y_train and X_test.
Learn an abstract scoring rule f_theta from the training arrays.
Compute y_pred. Reveal y_test only when calculating RMSE.
For user 1960, movies 2098, 898 and 2195 are ordered by y_pred.
Row alignment stays fixed: X_test[j] receives y_pred[j], which is later compared with y_test[j].
| u | i | r |
|---|---|---|
| 1960 | 670 | 4 |
| 1346 | 152 | 4 |
| u | i | r |
|---|---|---|
| 1574 | 2956 | ? |
| 1670 | 791 | ? |
fit on concat(train, test)
1960u_k
670i_l
Raw IDs are labels. Encoded indices are contiguous and start at zero.
The same raw ID must receive the same index in both splits.
X_train[j] = [u_idx, i_idx] y_train[j] = rating
Each row remains aligned.
X_test[j] = [u_idx, i_idx] y_test stays hidden
Keep the original test-row order.
f_theta = learn(X_train, y_train)y_pred[j] = f_theta(X_test[j])one score per test pair
We will define how f_theta works later.
RMSE(y_test, y_pred)
One score over 51,161 test rows.
20983.81
8983.73
21953.28
The same ID maps transform both splits. The model learns from training ratings, then returns one prediction for every test pair.
Information the model may use.
A complete rule, including fallbacks.
Pairs excluded from model fitting.
One comparable number for every method.
SimpleWe can explain every prediction.
CompleteNew users and movies still receive a score.
ReproducibleThe same training data gives the same benchmark.
User and item IDs are labels. The model needs stable array indices, but the spelling or numerical size of an ID carries no preference information.
USER LABELS ["U_A", "U_Z", "U_A"] fit user vocabulary USER INDICES [0, 1, 0]
MOVIE LABELS ["M_B", "M_A", "M_C"] fit movie vocabulary MOVIE INDICES [1, 0, 2]
“U_Z”
An identifier. Its characters carry no distance or order.
1
A valid zero-based position in a NumPy array.
Use two encoders: users and movies have separate vocabularies, even when both happen to be integers.
from sklearn.preprocessing import LabelEncoder
01LabelEncoder()Creates an encoder with no learned vocabulary.
02fit(labels)Finds the unique labels and stores them in classes_.
03transform(labels)Applies the stored mapping to known labels.
04inverse_transform(indices)Converts integer indices back to the original labels.
fit changes the encoder’s state. transform reuses that state and cannot invent an index for an unseen label.
| user | movie | rating |
|---|---|---|
| U_A | M_A | 4 |
| U_C | M_B | 5 |
| U_A | M_B | 3 |
FIT ON TRAIN ONLY User vocabulary U_A: 0 U_C: 1 Movie vocabulary M_A: 0 M_B: 1
| user | movie |
|---|---|
| U_B unseen | M_A |
| U_C | M_C unseen |
TRANSFORM CANNOT CONTINUE U_B is absent from the user vocabulary M_C is absent from the movie vocabulary
The course split contains test-only IDs. A train-only encoder cannot assign every scheduled test pair an index.
U_A, U_C, U_A
U_B, U_C
M_A, M_B, M_B
M_A, M_C
FIT ONCE Shared user map U_A: 0 U_B: 1 U_C: 2 Shared movie map M_A: 0 M_B: 1 M_C: 2
[0,0] [2,1] [0,1]
[1,0] [2,2]
WHAT ENTERS THE ENCODER? user_id and movie_id WHAT STAYS OUT? test.rating
Concatenation defines a common vocabulary. It does not merge the ratings or move test rows into the training set.
from sklearn.preprocessing import LabelEncoder
all_users = pd.concat([train["user_id"], test["user_id"]])
all_items = pd.concat([train["movie_id"], test["movie_id"]])
user_enc = LabelEncoder().fit(all_users)
item_enc = LabelEncoder().fit(all_items)
train["u"] = user_enc.transform(train["user_id"])
test["u"] = user_enc.transform(test["user_id"])
train["i"] = item_enc.transform(train["movie_id"])
test["i"] = item_enc.transform(test["movie_id"])
n_users = len(user_enc.classes_)
n_items = len(item_enc.classes_)
01Collect ID columnsNo ratings enter all_users or all_items.
02Fit onceEach encoder stores one shared vocabulary.
03Transform separatelyTrain and test remain different datasets.
fit defines the mapping. transform applies the same mapping wherever that entity appears.
51,161 observed triples
MODEL INPUTX_train(51161, 2)columns [u, i]
TRAINING TARGETy_train(51161,)observed ratings
MODEL INPUTX_test(51161, 2)columns [u, i]
MODEL OUTPUTy_pred(51161,)one score per row
X_train = train[["u", "i"]].to_numpy(dtype=np.int64)
y_train = train["rating"].to_numpy(dtype=float)
X_test = test[["u", "i"]].to_numpy(dtype=np.int64)
y_test = test["rating"].to_numpy(dtype=float) # evaluation only
y_pred.
During prediction, the method receives X_test. We reveal y_test only when calculating RMSE.
Understand the observations, the missing pairs, and the information available at prediction time.
for reaching the competition target
over the Cinematch benchmark
for specified customer–movie pairs
The competition reduced recommendation to a precise offline problem: estimate a customer’s missing movie rating.
Source: Netflix Prize and the Netflix dataset README.
| Netflix Prize | STAT3009 subset | |
|---|---|---|
| Training ratings | 100,480,507 | 51,161 |
| Users | 480,189 | 2,000 |
| Movies | 17,770 | 3,568 |
| Rating scale | 1–5 stars | 1–5 stars |
| Test labels | Hidden | Visible for teaching |
both = pd.concat([train, test])
len(train)
both[["user_id", "movie_id"]].nunique()
train["rating"].agg(["min", "max"])
51,161 rows
2,000 users · 3,568 movies
ratings from 1 to 5
How to read the comparisonPython gives the subset counts. The original counts come from the competition README.
Sources: Netflix dataset README and STAT3009 course data.
FITTraining · 99,072,112Released ratings used to estimate the model.
VALIDATEProbe · 1,408,395Known ratings held out for local model checks.
PUBLIC FEEDBACKQuiz · 1,408,342Hidden ratings produced the leaderboard score.
FINAL DECISIONTest · 1,408,789Hidden ratings determined the final ranking.
100,480,507 = 99,072,112 + 1,408,395The released rating total includes the fitting and probe portions.
Course versiontrain.csv fits the model. We hide the test.csv rating column while predicting, then reveal it to calculate RMSE.
| movie_id | user_id | rating | |
|---|---|---|---|
| 0 | 670 | 1960 | 4 |
| 1 | 152 | 1346 | 4 |
| 2 | 1741 | 785 | 4 |
| 3 | 3032 | 686 | 5 |
| 4 | 536 | 1894 | 4 |
(1960, 670, 4) means customer 1960 gave movie 670 four stars.
IDs identify categories. Their numerical distance carries no meaning.
train["rating"].value_counts(
normalize=True
).sort_index()
train["rating"].agg(["mean", "std"])
TRAINING MEAN3.621
TRAINING SD1.088
3–5 STAR RATINGS85.5%
The target is not centered at three stars. A credible first prediction should reflect this.
both = pd.concat([train, test])
n_user = both["user_id"].nunique()
n_movie = both["movie_id"].nunique()
density = len(train) / (n_user * n_movie)
100 * density
0.7169%
0.717% observed 99.283% unobserved
A blank cell means “no observed rating.” It does not mean that the customer assigned zero stars.
| Median | 13 |
|---|---|
| 90th percentile | 65 |
| Maximum | 535 |
| Median | 3 |
|---|---|
| 90th percentile | 41 |
| Maximum | 411 |
user_n = train.groupby("user_id").size()
movie_n = train.groupby("movie_id").size()
user_n.quantile([.5, .9]), user_n.max()
movie_n.quantile([.5, .9]), movie_n.max()
The group means summarize different amounts of evidence: some users and movies have many more ratings than others.
Used to estimate every mean and model parameter.
Ratings enter the RMSE calculation only.
USERS74 unseen in train1,843 users occur in both files
MOVIES584 unseen in train2,415 movies occur in both files
keys = ["user_id", "movie_id"]
n_overlap = train.merge(test, on=keys).shape[0]
new_users = len(set(test.user_id) - set(train.user_id))
new_movies = len(set(test.movie_id) - set(train.movie_id))
0 overlapping pairs · 74 new users · 584 new movies
| Test pair | User / movie seen? | Rows | Share |
|---|---|---|---|
| Warm pair | Yes / Yes | 50,159 | 98.04% |
| New movie | Yes / No | 850 | 1.66% |
| New user | No / Yes | 149 | 0.29% |
| New user and movie | No / No | 3 | 0.01% |
seen_u = test.user_id.isin(train.user_id)
seen_i = test.movie_id.isin(train.movie_id)
pd.crosstab(seen_u, seen_i)
A held-out pair can still contain familiar IDs.Fallbacks cover the 1,002 rows with a new user or movie.
global mean, user effects, movie effects
one score for every test row
compare truth with predictions
pd.concat([train, test]) lets test ratings influence the prediction.
Use both files only to describe the dataset universe, never to fit a predictive quantity.
\[ \operatorname{RMSE} = \sqrt{\frac{1}{n_{\mathrm{test}}} \sum_{j=1}^{n_{\mathrm{test}}}(y_j-\hat y_j)^2} \]
InterpretationTypical prediction error, expressed on the 1–5 rating scale, with larger misses penalized more heavily.
| u | i | r |
|---|---|---|
| 1960 | 670 | 4 |
| 1346 | 152 | 4 |
| 785 | 1741 | 4 |
| 686 | 3032 | 5 |
Learn which systematic patterns explain r.
Use quantities calculated only from train.
Row j in the vector predicts row j in test.
len(y_pred) == len(test)
test[“rating”]. We use that column only after the vector is complete.
Use progressively more specific averages while retaining a prediction for every test row.
ONE NUMBER FROM TRAIN \[ \mu=\frac{1}{|\Omega^{\mathrm{tr}}|}\sum_{(u,i)\in\Omega^{\mathrm{tr}}}r_{ui} \qquad \hat r_{ui}=\mu \]
TRAINING MEAN3.621
TEST RMSE1.085
UNSEEN IDSAlways defined
The rating distribution gives us a stronger constant prediction than an arbitrary midpoint such as three stars.
| u | i | r |
|---|---|---|
| 1960 | 670 | 4 |
| 1346 | 152 | 4 |
| 785 | 1741 | 4 |
| 686 | 3032 | 5 |
| 1894 | 536 | 4 |
| ⋮ | ⋮ | ⋮ |
The global mean uses the r column. It does not group by u or i.
SUM OF ALL RATINGS185,262y_train.sum()
NUMBER OF RATINGS51,161y_train.size
185,26251,161
=3.621156…
CALCULATEglobal_mean = y_train.mean()
FILL THE TEST VECTORy_pred_global = np.full(X_test.shape[0], global_mean)
User 58974 ratings2.01
User 119074 ratings4.99
Movie 234164 ratings2.75
Movie 20581 ratings4.49
One global number leaves predictable structure in the prediction errors: some users rate low, and some movies receive stronger ratings.
ONE NUMBER PER OBSERVED USER \[ \bar r_u=\frac{1}{|\mathcal I_u^{\mathrm{tr}}|} \sum_{i\in\mathcal I_u^{\mathrm{tr}}}r_{ui} \]
KNOWN USERr̂ui = r̄usame prediction for every movie
NEW USERr̂ui = μfall back to the global mean
TEST RMSE1.017
CHANGE FROM GLOBAL−0.069
| u | i | r |
|---|---|---|
| 1960 | 670 | 4 |
| 589 | 3181 | 1 |
| 1346 | 152 | 4 |
| 1960 | 2597 | 4 |
| 589 | 422 | 2 |
| 1346 | 2102 | 5 |
Rows with the same u belong to one group. The movie IDs may differ.
| User | Σr / nu | User mean |
|---|---|---|
| 1960 | 1,324 / 347 | 3.816 |
| 589 | 149 / 74 | 2.014 |
| 1346 | 398 / 107 | 3.720 |
Same calculation, separate groupsEach observed user receives one value. User 1960 gets 1324 / 347 = 3.815562.
user_mean[1960] = 3.816
user_mean[589] = 2.014
user_mean[1346] = 3.720
User column[0, 1, 0, 2, 1]
Boolean mask[F, T, F, F, T]
Selected ratings[1, 2]
user_mean[1]1.5
Build the lookup vector from trainX_train[:, 0] == u selects the ratings from user u.
Predict by array indexingEach encoded test user retrieves one value from user_mean.
ONE NUMBER PER OBSERVED MOVIE \[ \bar r_i=\frac{1}{|\mathcal U_i^{\mathrm{tr}}|} \sum_{u\in\mathcal U_i^{\mathrm{tr}}}r_{ui} \]
KNOWN MOVIEr̂ui = r̄isame prediction for every user
NEW MOVIEr̂ui = μfall back to the global mean
TEST RMSE1.052
CHANGE FROM GLOBAL−0.034
| u | i | r |
|---|---|---|
| 1960 | 670 | 4 |
| 1346 | 152 | 4 |
| 785 | 1741 | 4 |
| 1882 | 670 | 4 |
| 657 | 152 | 2 |
| 961 | 1741 | 4 |
Rows with the same i belong to one group. The user IDs may differ.
| Movie | Σr / ni | Item mean |
|---|---|---|
| 670 | 145 / 47 | 3.085 |
| 152 | 292 / 74 | 3.946 |
| 1741 | 416 / 108 | 3.852 |
Only the grouping column changesUser mean groups by column 0. Item mean groups by column 1.
item_mean = np.full(n_items, global_mean)
for i in np.unique(X_train[:, 1]):
item_mean[i] = np.mean(y_train[X_train[:, 1] == i])
y_pred_item = item_mean[X_test[:, 1]]
| Method | What varies? | Fallback | Test RMSE |
|---|---|---|---|
| Global mean | Nothing | Always defined | 1.085 |
| User mean | User | Global mean for new user | 1.017 |
| Item mean | Movie | Global mean for new movie | 1.052 |
User effects explain more error in this splitThe user mean improves more than the movie mean.
Neither model combines both sourcesUser mean ignores which movie is scored. Movie mean ignores which user is scored.
| Test case | Global | User mean | Item mean |
|---|---|---|---|
| Known user, known movie | μ | r̄u | r̄i |
| Known user, new movie | μ | r̄u | μ |
| New user, known movie | μ | μ | r̄i |
| New user, new movie | μ | μ | μ |
The global mean is more than a weak model. It is the fallback that makes the other baselines complete.
Take the arithmetic average of the user mean, item mean, and global mean.
r̄uUser meanUser u’s average training rating.
r̄iItem meanMovie i’s average training rating.
μGlobal meanThe average over every training rating.
\[\hat r_{ui}=\frac{\bar r_u+\bar r_i+\mu}{3}\]
\[ \hat r_{ui} = w_u\bar r_u+w_i\bar r_i+w_0\mu, \qquad w_u+w_i+w_0=1 \]
USER WEIGHTwuControls how much the user’s rating tendency contributes.
ITEM WEIGHTwiControls how much the movie’s average rating contributes.
GLOBAL WEIGHTw0Pulls the prediction toward the overall rating level.
Equal weights give a simple average. Other weights give different averages without changing the prediction workflow.
Missing user or item meanPut the global mean into the missing slot, then average the same three positions.
OUTPUT 1.005
Equal weightseach mean contributes one third
Missing group meanreplace it with the global mean
Complete outputevery test row receives a score
| Method | Test RMSE | Reduction from global | Information used |
|---|---|---|---|
| Global mean | 1.085 | reference | overall rating level |
| Item mean | 1.052 | 0.034 | movie history |
| User mean | 1.017 | 0.069 | user history |
| Three-mean average | 1.005 | 0.081 | user, movie, and global means |
The equal-weight three-mean baseline reduces test RMSE by about 7.4% relative to the global baseline on this course split.
| Test-pair type | Rows | Three-mean RMSE | Prediction used |
|---|---|---|---|
| Known user, known movie | 50,159 | 1.000 | (r̄u + r̄i + μ) / 3 |
| Known user, new movie | 850 | 1.264 | (r̄u + 2μ) / 3 |
| New user, known movie | 149 | 1.114 | (r̄i + 2μ) / 3 |
| New user, new movie | 3 | 0.621 | μ only |
The final row has only three observations. Its small RMSE is sampling noise, not evidence that the global mean handles full cold start well.
Return from rating prediction to the ordered recommendation list.
347 training ratings · training mean 3.82
| Candidate movie | Global | User mean | Item mean | Three-mean | Observed test rating |
|---|---|---|---|---|---|
| Movie 2098 | 3.62 | 3.82 | 4.00 | 3.81 | 4 |
| Movie 898 | 3.62 | 3.82 | 3.75 | 3.73 | 4 |
| Movie 2195 | 3.62 | 3.82 | 2.40 | 3.28 | 2 |
Global and user means produce tiesThey cannot order movies for one user.
Three-mean scores vary by movieSorting the three-mean column gives 2098, 898, 2195.
For one user, the user and global terms are constant. The three-mean score therefore preserves the item-mean order while pulling scores toward the center.
Source: predictions calculated from the STAT3009 course files. Test ratings appear only for retrospective evaluation.
Separate average tendencies with equal weights.
Preference specific to this user–movie pairing.
A strict rater can still love one movie.The user’s low average pulls every prediction down, but it cannot represent a special match.
A popular movie can still be wrong for one user.The movie’s high average pulls every prediction up, but it cannot represent individual incompatibility.
Collaborative filtering models the user–movie interaction that separate averages cannot represent.
01 Rating imbalance sets the global level 02 User and movie variation motivates group effects 03 Cold-start rows require explicit fallbacks 04 Unexplained user–movie interaction motivates collaborative filtering
CUHK · STAT3009 · Recommender Systems