Baseline Methods for Rating Prediction

STAT3009 · Recommender Systems

Ben Dai

CUHK · Department of Statistics and Data Science

The Netflix rating-prediction pipeline

Row alignment stays fixed: X_test[j] receives y_pred[j], which is later compared with y_test[j].

Netflix data preparation and prediction

The same ID maps transform both splits. The model learns from training ratings, then returns one prediction for every test pair.

A baseline must score every test pair

SimpleWe can explain every prediction.

CompleteNew users and movies still receive a score.

ReproducibleThe same training data gives the same benchmark.

Data preprocessing: categorical IDs

User and item IDs are labels. The model needs stable array indices, but the spelling or numerical size of an ID carries no preference information.

RAW LABEL“U_Z”

An identifier. Its characters carry no distance or order.

ARRAY INDEX1

A valid zero-based position in a NumPy array.

Use two encoders: users and movies have separate vocabularies, even when both happen to be integers.

sklearn.preprocessing.LabelEncoder

PACKAGEsklearn
.
MODULEpreprocessing
.
CLASSLabelEncoder

from sklearn.preprocessing import LabelEncoder

enc = LabelEncoder()
labels = ["U_A", "U_Z", "U_A"]

enc.fit(labels)
enc.classes_                       # ["U_A", "U_Z"]

enc.transform(["U_Z", "U_A"])    # [1, 0]
enc.inverse_transform([0, 1])      # ["U_A", "U_Z"]

01LabelEncoder()Creates an encoder with no learned vocabulary.

02fit(labels)Finds the unique labels and stores them in classes_.

03transform(labels)Applies the stored mapping to known labels.

04inverse_transform(indices)Converts integer indices back to the original labels.

fit changes the encoder’s state. transform reuses that state and cannot invent an index for an unseen label.

Train-only ID vocabulary

TRANSFORM CANNOT CONTINUE U_B is absent from the user vocabulary M_C is absent from the movie vocabulary

The course split contains test-only IDs. A train-only encoder cannot assign every scheduled test pair an index.

Shared train–test ID vocabulary

WHAT ENTERS THE ENCODER? user_id and movie_id WHAT STAYS OUT? test.rating

Concatenation defines a common vocabulary. It does not merge the ratings or move test rows into the training set.

LabelEncoder implementation

from sklearn.preprocessing import LabelEncoder

all_users = pd.concat([train["user_id"], test["user_id"]])
all_items = pd.concat([train["movie_id"], test["movie_id"]])

user_enc = LabelEncoder().fit(all_users)
item_enc = LabelEncoder().fit(all_items)

train["u"] = user_enc.transform(train["user_id"])
test["u"]  = user_enc.transform(test["user_id"])
train["i"] = item_enc.transform(train["movie_id"])
test["i"]  = item_enc.transform(test["movie_id"])

n_users = len(user_enc.classes_)
n_items = len(item_enc.classes_)

01Collect ID columnsNo ratings enter all_users or all_items.

02Fit onceEach encoder stores one shared vocabulary.

03Transform separatelyTrain and test remain different datasets.

fit defines the mapping. transform applies the same mapping wherever that entity appears.

Data preprocessing: NumPy arrays

X_train = train[["u", "i"]].to_numpy(dtype=np.int64)
y_train = train["rating"].to_numpy(dtype=float)

X_test = test[["u", "i"]].to_numpy(dtype=np.int64)
y_test = test["rating"].to_numpy(dtype=float)  # evaluation only
IMPLEMENTATION RULE Pandas prepares the tables. NumPy fits every baseline and produces y_pred.

During prediction, the method receives X_test. We reveal y_test only when calculating RMSE.

Netflix data and evaluation

01
CUHK emblem

Understand the observations, the missing pairs, and the information available at prediction time.

The Netflix Prize made rating prediction measurable

PRIZEUS$1 million

for reaching the competition target

TARGET10% improvement

over the Cinematch benchmark

TASKPredict hidden ratings

for specified customer–movie pairs

The competition reduced recommendation to a precise offline problem: estimate a customer’s missing movie rating.

Source: Netflix Prize and the Netflix dataset README.

The full data and the course subset

Netflix Prize STAT3009 subset
Training ratings 100,480,507 51,161
Users 480,189 2,000
Movies 17,770 3,568
Rating scale 1–5 stars 1–5 stars
Test labels Hidden Visible for teaching
PYTHON · COURSE FILES
both = pd.concat([train, test])

len(train)
both[["user_id", "movie_id"]].nunique()
train["rating"].agg(["min", "max"])

51,161 rows
2,000 users · 3,568 movies
ratings from 1 to 5

How to read the comparisonPython gives the subset counts. The original counts come from the competition README.

Sources: Netflix dataset README and STAT3009 course data.

Competition split and course split

Netflix Prize training, probe, quiz, and test partitions
Source: Pantelis Monogioudis, Netflix Prize lecture notes.

FITTraining · 99,072,112Released ratings used to estimate the model.

VALIDATEProbe · 1,408,395Known ratings held out for local model checks.

PUBLIC FEEDBACKQuiz · 1,408,342Hidden ratings produced the leaderboard score.

FINAL DECISIONTest · 1,408,789Hidden ratings determined the final ranking.

100,480,507 = 99,072,112 + 1,408,395The released rating total includes the fitting and probe portions.

Course versiontrain.csv fits the model. We hide the test.csv rating column while predicting, then reveal it to calculate RMSE.

The course tables

import numpy as np
import pandas as pd

base = (
    "https://raw.githubusercontent.com/"
    "statmlben/CUHK-STAT3009/main/"
    "dataset/netflix/"
)

fields = ["movie_id", "user_id", "rating"]
train = pd.read_csv(base + "train.csv", usecols=fields)
test = pd.read_csv(base + "test.csv", usecols=fields)
train.head()
movie_id user_id rating
0 670 1960 4
1 152 1346 4
2 1741 785 4
3 3032 686 5
4 536 1894 4

(1960, 670, 4) means customer 1960 gave movie 670 four stars.

IDs identify categories. Their numerical distance carries no meaning.

Rating distribution in the training set

PYTHON
train["rating"].value_counts(
    normalize=True
).sort_index()

train["rating"].agg(["mean", "std"])

TRAINING MEAN3.621

TRAINING SD1.088

3–5 STAR RATINGS85.5%

The target is not centered at three stars. A credible first prediction should reflect this.

The observed matrix is 0.717% full

POSSIBLE USER–MOVIE PAIRS2,000 × 3,5687,136,000 cells
OBSERVED TRAINING RATINGS51,161one rating per observed pair
PYTHON
both = pd.concat([train, test])
n_user = both["user_id"].nunique()
n_movie = both["movie_id"].nunique()

density = len(train) / (n_user * n_movie)
100 * density

0.7169%

0.717% observed 99.283% unobserved

A blank cell means “no observed rating.” It does not mean that the customer assigned zero stars.

Interaction counts vary sharply

RATINGS PER USER IN TRAIN
Median 13
90th percentile 65
Maximum 535
RATINGS PER MOVIE IN TRAIN
Median 3
90th percentile 41
Maximum 411
PYTHON · GROUP SIZE, THEN SUMMARIZE
user_n = train.groupby("user_id").size()
movie_n = train.groupby("movie_id").size()
user_n.quantile([.5, .9]), user_n.max()
movie_n.quantile([.5, .9]), movie_n.max()

The group means summarize different amounts of evidence: some users and movies have many more ratings than others.

Train and test contain different pairs

TRAIN51,161 pairs

Used to estimate every mean and model parameter.

0overlapping user–movie pairs
TEST51,161 pairs

Ratings enter the RMSE calculation only.

USERS74 unseen in train1,843 users occur in both files

MOVIES584 unseen in train2,415 movies occur in both files

PYTHON
keys = ["user_id", "movie_id"]
n_overlap = train.merge(test, on=keys).shape[0]
new_users = len(set(test.user_id) - set(train.user_id))
new_movies = len(set(test.movie_id) - set(train.movie_id))

0 overlapping pairs · 74 new users · 584 new movies

Four types of test pair

Test pair User / movie seen? Rows Share
Warm pair Yes / Yes 50,159 98.04%
New movie Yes / No 850 1.66%
New user No / Yes 149 0.29%
New user and movie No / No 3 0.01%
PYTHON
seen_u = test.user_id.isin(train.user_id)
seen_i = test.movie_id.isin(train.movie_id)

pd.crosstab(seen_u, seen_i)

A held-out pair can still contain familiar IDs.Fallbacks cover the 1,002 rows with a new user or movie.

Leakage-free evaluation

DATA LEAKAGE Computing a user or movie mean from pd.concat([train, test]) lets test ratings influence the prediction.

Use both files only to describe the dataset universe, never to fit a predictive quantity.

RMSE as the common scoreboard

\[ \operatorname{RMSE} = \sqrt{\frac{1}{n_{\mathrm{test}}} \sum_{j=1}^{n_{\mathrm{test}}}(y_j-\hat y_j)^2} \]

InterpretationTypical prediction error, expressed on the 1–5 rating scale, with larger misses penalized more heavily.

def rmse(y_true, y_pred):
    error = np.asarray(y_true) - np.asarray(y_pred)
    return np.sqrt(np.mean(error ** 2))

# test ratings enter only here
score = rmse(y_test, y_pred)
EVERY METHOD RETURNS51,161 predictions

in the same row order as test

From training triples to a prediction vector

len(y_pred) == len(test)
The model never reads test[“rating”]. We use that column only after the vector is complete.

Mean baselines

02
CUHK emblem

Use progressively more specific averages while retaining a prediction for every test row.

Global mean baseline

ONE NUMBER FROM TRAIN \[ \mu=\frac{1}{|\Omega^{\mathrm{tr}}|}\sum_{(u,i)\in\Omega^{\mathrm{tr}}}r_{ui} \qquad \hat r_{ui}=\mu \]

TRAINING MEAN3.621

TEST RMSE1.085

UNSEEN IDSAlways defined

The rating distribution gives us a stronger constant prediction than an arbitrary midpoint such as three stars.

Global mean from the training triples

Systematic variation across users and movies

USER RATING LEVELS

User 58974 ratings2.01

User 119074 ratings4.99

MOVIE RATING LEVELS

Movie 234164 ratings2.75

Movie 20581 ratings4.49

One global number leaves predictable structure in the prediction errors: some users rate low, and some movies receive stronger ratings.

User mean baseline

ONE NUMBER PER OBSERVED USER \[ \bar r_u=\frac{1}{|\mathcal I_u^{\mathrm{tr}}|} \sum_{i\in\mathcal I_u^{\mathrm{tr}}}r_{ui} \]

KNOWN USERr̂ui = r̄usame prediction for every movie

NEW USERr̂ui = μfall back to the global mean

TEST RMSE1.017

CHANGE FROM GLOBAL−0.069

User means from training triples

User mean in NumPy

user_mean = np.full(n_users, global_mean)

for u in np.unique(X_train[:, 0]):
    user_mean[u] = np.mean(y_train[X_train[:, 0] == u])

y_pred_user = user_mean[X_test[:, 0]]
EXAMPLE · USER u = 1

User column[0, 1, 0, 2, 1]

Boolean mask[F, T, F, F, T]

Selected ratings[1, 2]

user_mean[1]1.5

Build the lookup vector from trainX_train[:, 0] == u selects the ratings from user u.

Predict by array indexingEach encoded test user retrieves one value from user_mean.

Item mean baseline

ONE NUMBER PER OBSERVED MOVIE \[ \bar r_i=\frac{1}{|\mathcal U_i^{\mathrm{tr}}|} \sum_{u\in\mathcal U_i^{\mathrm{tr}}}r_{ui} \]

KNOWN MOVIEr̂ui = r̄isame prediction for every user

NEW MOVIEr̂ui = μfall back to the global mean

TEST RMSE1.052

CHANGE FROM GLOBAL−0.034

Item mean uses the same pattern

Three mean baselines on the test set

Method What varies? Fallback Test RMSE
Global mean Nothing Always defined 1.085
User mean User Global mean for new user 1.017
Item mean Movie Global mean for new movie 1.052

User effects explain more error in this splitThe user mean improves more than the movie mean.

Neither model combines both sourcesUser mean ignores which movie is scored. Movie mean ignores which user is scored.

Fallback behavior for unseen IDs

Test case Global User mean Item mean
Known user, known movie μ r̄u r̄i
Known user, new movie μ r̄u μ
New user, known movie μ μ r̄i
New user, new movie μ μ μ

The global mean is more than a weak model. It is the fallback that makes the other baselines complete.

Three-mean baseline

03
CUHK emblem

Take the arithmetic average of the user mean, item mean, and global mean.

Arithmetic average of three means

THREE NUMBERS · EQUAL WEIGHT
user mean
+
item mean
+
global mean

r̄uUser meanUser u’s average training rating.

r̄iItem meanMovie i’s average training rating.

μGlobal meanThe average over every training rating.

\[\hat r_{ui}=\frac{\bar r_u+\bar r_i+\mu}{3}\]

Flexible weights

\[ \hat r_{ui} = w_u\bar r_u+w_i\bar r_i+w_0\mu, \qquad w_u+w_i+w_0=1 \]

USER WEIGHTwuControls how much the user’s rating tendency contributes.

ITEM WEIGHTwiControls how much the movie’s average rating contributes.

GLOBAL WEIGHTw0Pulls the prediction toward the overall rating level.

CHOICE USED TODAY wu = wi = w0 = 1/3

Equal weights give a simple average. Other weights give different averages without changing the prediction workflow.

A three-number prediction

Missing user or item meanPut the global mean into the missing slot, then average the same three positions.

Three-mean baseline in Python

u_test = X_test[:, 0]
i_test = X_test[:, 1]

# The mean arrays already contain the global fallback
user_term = user_mean[u_test]
item_term = item_mean[i_test]

y_pred_three = (user_term + item_term + global_mean) / 3

round(rmse(y_test, y_pred_three), 3)

OUTPUT 1.005

Equal weightseach mean contributes one third

Missing group meanreplace it with the global mean

Complete outputevery test row receives a score

Test-set results

Method Test RMSE Reduction from global Information used
Global mean 1.085 reference overall rating level
Item mean 1.052 0.034 movie history
User mean 1.017 0.069 user history
Three-mean average 1.005 0.081 user, movie, and global means

The equal-weight three-mean baseline reduces test RMSE by about 7.4% relative to the global baseline on this course split.

Performance by test-pair type

Test-pair type Rows Three-mean RMSE Prediction used
Known user, known movie 50,159 1.000 (r̄u + r̄i + μ) / 3
Known user, new movie 850 1.264 (r̄u + 2μ) / 3
New user, known movie 149 1.114 (r̄i + 2μ) / 3
New user, new movie 3 0.621 μ only

The final row has only three observations. Its small RMSE is sampling noise, not evidence that the global mean handles full cold start well.

Prediction scores and ranking

04
CUHK emblem

Return from rating prediction to the ordered recommendation list.

From prediction scores to a ranking

EXAMPLE TEST USERUser 1960

347 training ratings · training mean 3.82

Candidate movie Global User mean Item mean Three-mean Observed test rating
Movie 2098 3.62 3.82 4.00 3.81 4
Movie 898 3.62 3.82 3.75 3.73 4
Movie 2195 3.62 3.82 2.40 3.28 2

Global and user means produce tiesThey cannot order movies for one user.

Three-mean scores vary by movieSorting the three-mean column gives 2098, 898, 2195.

For one user, the user and global terms are constant. The three-mean score therefore preserves the item-mean order while pulling scores toward the center.

Source: predictions calculated from the STAT3009 course files. Test ratings appear only for retrospective evaluation.

What the three-mean model cannot explain

THREE-MEAN MODEL(μ + r̄u + r̄i) / 3

Separate average tendencies with equal weights.

leaves
UNEXPLAINED ERRORrui − r̂ui

Preference specific to this user–movie pairing.

A strict rater can still love one movie.The user’s low average pulls every prediction down, but it cannot represent a special match.

A popular movie can still be wrong for one user.The movie’s high average pulls every prediction up, but it cannot represent individual incompatibility.

Collaborative filtering models the user–movie interaction that separate averages cannot represent.

Baselines turn data properties into modeling choices

01 Rating imbalance sets the global level 02 User and movie variation motivates group effects 03 Cold-start rows require explicit fallbacks 04 Unexplained user–movie interaction motivates collaborative filtering