Machine Learning I: Models and Estimators

From prediction rules to reusable code

Ben Dai

CUHK · Department of Statistics and Data Science

Today’s question

WE CAN ALREADY PREDICT global mean · user mean · item mean

How do we move from hand-designed averages to a general way of learning prediction rules from data?

BEFOREChoose a formula

Then calculate its averages from the training triples.

→
TODAYChoose a model family

Then learn its parameters from examples.

→
THE REAL CHALLENGEGeneralize

Work well on ratings the model has never seen.

The learning loop

🔒X_test is available. y_test stays out of fitting and model selection.

ML I · A language for prediction

01
CUHK emblem

Examples, models, losses, parameters—and the separation between fitting and evaluation.

Rating prediction is supervised learning

RECOMMENDER SYSTEM
user\(u\)
item\(i\)
→
rating\(r_{ui}\)
HOUSING REGRESSION
income\(x_1\)
location\(x_2\)
→
value\(y\)

Supervised learning: learn a mapping from inputs to a known numerical target, then apply it to new inputs.

One row becomes one learning example

RATING DATA
user item rating
1960 670 4
1346 152 4
785 1741 4

\(X\) contains the inputs; \(y\) contains the answer.

ONE EXAMPLE
\(x_1=(1960,670)\)

the user–item pair

paired with
\(y_1=4\)

the observed rating

The row number is not a feature. It only identifies an observation.

Four questions decode every estimator

1 · MODEL\(f_{\theta,h}(x)\)

What family of prediction rules can this estimator represent?

2 · PARAMETERS\(\theta\)

What quantities will fit learn from the training data?

3 · HYPERPARAMETERS\(h\)

What choices are fixed before fit and control the model or fitting rule?

4 · LOSS / CRITERION\(L(y,\widehat y)\)

What does the fitting procedure regard as a better solution?

\[h\ \text{chosen before fit}\quad\Longrightarrow\quad \widehat\theta=\arg\min_{\theta}\frac{1}{n}\sum_{i=1}^{n}L\!\left(y_i,f_{\theta,h}(x_i)\right) \quad\Longrightarrow\quad \widehat f=f_{\widehat\theta,h}\]

Parameters are learned; hyperparameters are chosen

BEFORE fit · CHOSEN hyperparameter \(h\) fit_intercept=True

Set by us—or later selected with validation—to control the model or fitting process.

fit(X, y)uses training data
DURING fit · LEARNED parameter \(\widehat\theta\) coef_ · intercept_

Estimated from the observed examples and stored in the fitted object.

One word, two conventionsscikit-learn calls constructor arguments “parameters” in get_params(). In our ML template, we call model settings hyperparameters to distinguish them from quantities learned during fit.

Quick test: did the data determine the value? Yes → learned parameter. No, fixed before fitting → hyperparameter.

A model is a family—not one fitted line

LINEAR FAMILY \(f_{\theta}(x)=\theta_0+\theta_1x\)

The formula gives a family. Each value of \(\theta=(\theta_0,\theta_1)\) gives a different member.

Training decides which member of the family to use.

Training minimizes average loss

1 · SCORE ONE PARAMETER SETTING

Fix \(\theta^{(B)}\). For every training row, predict \(\widehat y_i=f_{\theta^{(B)}}(x_i)\) and calculate \(\ell_i=(y_i-\widehat y_i)^2\).

row \(i\) observed \(y_i\) prediction \(\widehat y_i\) loss \(\ell_i\)
1 4.0 3.5 0.25
2 2.0 2.5 0.25
3 5.0 4.0 1.00

AVERAGE THE ROW LOSSES \[J\!\left(\theta^{(B)}\right)=\frac{0.25+0.25+1.00}{3}=0.50\]

For squared loss, minimizing MSE also minimizes RMSE. A small training loss still does not guarantee good predictions on new rows.

Prediction freezes the fitted rule

AFTER FITTING \(\widehat f(x)=f_{\widehat\theta}(x)\)

\(\widehat\theta\) is now fixed.

apply
NEW INPUT\(x_{new}\)

features only

→
PREDICTION\(\widehat y_{new}\)

\(=\widehat f(x_{new})\)

.fit(X_train, y_train)changes the model by learning parameters

.predict(X_new)uses the fitted model without learning from new labels

Training and test answer different questions

TRAINING SET How should the model fit?
  • X_train and y_train are available
  • estimates parameters
  • may be reused during fitting
TEST INPUT AND LABEL What may the model see?
  • X_test is available for prediction
  • y_test does not estimate parameters
  • y_test is opened only for final evaluation

Generalization is performance on new examples drawn from the task we care about.

Keep the test labels closed

AVAILABLE FROM THE START Inputs for fitting and prediction

X_traintraining inputs

y_traintraining answers

X_testtest user–item pairs

HIDDEN ANSWERS y_test

opened only after all modeling choices are fixed

MODEL SELECTION QUESTION How can we compare model A and model B while y_test remains hidden? ML II introduces holdout validation and cross-validation.

A complete regression example

02
CUHK emblem

California housing: define \(X\) and \(y\), fit linear regression, predict, and calculate RMSE.

California housing is a regression task

ROWS20,640

California block groups

FEATURES8

demographic and geographic measurements

TARGETMedHouseVal

median house value in units of $100,000

MedIncHouseAgeAveRoomsAveBedrmsPopulationAveOccupLatitudeLongitude

Source: scikit-learn California housing dataset documentation.

\(X\) is a matrix; \(y\) is a vector

FEATURE MATRIX \(X_{train}\in\mathbb{R}^{13828\times 8}\)

one row = one block group
one column = one feature

paired row by row
TARGET VECTOR \(y_{train}\in\mathbb{R}^{13828}\)
4.5263.5853.5213.413

one numerical answer
for every row of \(X\)

X_train.shape(13828, 8)y_train.shape(13828,)

Linear regression combines the features

\[\widehat y=\theta_0+\theta_1x_1+\theta_2x_2+\cdots+\theta_8x_8\]

\(\theta_0\)intercept

the model’s reference level

\(\theta_j\)coefficient

how prediction changes with feature \(j\), holding others fixed

\(\widehat y\)prediction

a weighted sum calculated for one row

The model family is linear. The coefficients are learned from the training data.

Fitting makes the training residuals small

ONE RESIDUAL \(e_i=y_i-\widehat y_i\)

Vertical distance between the observed target and the fitted prediction.

LEAST SQUARES FIT \[\widehat\theta=\arg\min_{\theta}\sum_{i=1}^{n}e_i^2\]

Squaring makes large errors count more and prevents positive and negative errors from cancelling.

Read LinearRegression with the four questions

MODEL\(\widehat y=\theta_0+\sum_{j=1}^{8}\theta_jx_j\)

A linear family: one intercept plus one coefficient for each feature.

LEARNED PARAMETERS\(\widehat\theta_0,\ldots,\widehat\theta_8\)

Stored after fit as intercept_ and coef_.

HYPERPARAMETERfit_intercept=True

Chosen before fitting; controls whether the model includes an intercept.

LOSS / CRITERION\(\sum_i(y_i-\widehat y_i)^2\)

Ordinary least squares chooses coefficients with the smallest residual sum of squares.

THE TEMPLATE ANSWERchoose the hyperparameter → fit the learned parameters → obtain one fitted linear rule

Fit and predict in scikit-learn

from sklearn.linear_model import LinearRegression
from sklearn.metrics import root_mean_squared_error

model = LinearRegression()

model.fit(X_train, y_train)
y_pred = model.predict(X_test)

rmse = root_mean_squared_error(y_test, y_pred)

1ConstructChoose the model family.

2FitLearn coefficients from training examples.

3PredictProduce one value for every test row.

4EvaluateSummarize the prediction errors.

fit sees targets. predict sees features only.

Many errors become one RMSE

TRUE\(y\)2.41.83.22.6

−

PREDICTED\(\widehat y\)2.12.02.72.8

→

ERROR\(e\)0.3−0.20.5−0.2

01square

0.09, 0.04, 0.25, 0.04

02average

\((0.42)/4=0.105\)

03square root

\(\mathrm{RMSE}=0.324\)

The scikit-learn estimator interface

03
CUHK emblem

A shared contract turns different learning rules into reusable prediction objects.

One contract for every prediction model

LinearRegression()KNeighborsRegressor()YourRegressor()can share the same interface

The contract specifies how we use a model. Each estimator supplies its own learning and prediction rules.

The estimator as an octopus

One object stores the state. Its methods provide controlled ways to use or update that state.Different estimator classes may add more “tentacles,” while the common sklearn methods keep familiar names.

Octopus illustration: Jason J. Patterson / Open Clip Art Library, CC0.

One format, many regression algorithms

The “template” is an API contract.It standardizes names, inputs, outputs, and object state. It does not prescribe the learning algorithm.

Class, estimator, and fitted estimator

model.fit(X, y) is modelTruefit conventionally returns self.

LinearRegression as a running example

import numpy as np
from sklearn.linear_model import LinearRegression

X_train = np.array([
    [1., 0.],
    [0., 1.],
    [1., 1.],
    [2., 1.],
])
y_train = np.array([3., 2., 4., 6.])

model = LinearRegression(fit_intercept=True)

X_trainfour samples, two featuresEach row describes one training case.

y_trainfour known targetsPosition j matches row j of X_train.

modelan unfitted estimatorIt knows its settings but has not learned a line.

LinearRegression(…) creates a Python object. Training happens only when we call fit.

Hyperparameters configure fitting

model = LinearRegression(
    fit_intercept=True,
    positive=False,
)

model.get_params()["fit_intercept"]
# True

model.set_params(positive=True)
# LinearRegression(positive=True)

init(…)records hyperparametersThese settings describe how a future fit should behave.

get_params()returns constructor settingsScikit-learn calls them estimator “parameters” in its API.

set_params(…)changes hyperparametersFit the estimator again after changing a setting.

Training data do not belong in init.X and y enter through fit(X, y).

fit learns from paired rows of X and y

X_train[j] and y_train[j] form one supervised example.fit returns self, which makes LinearRegression().fit(X, y) valid syntax.

Fitted attributes expose the learned equation

model = LinearRegression().fit(X_train, y_train)

model.coef_
# array([2., 1.])

model.intercept_
# 1.0

model.n_features_in_
# 2
LEARNED PREDICTION RULE

\(\widehat y = 1.0 + 2.0x_1 + 1.0x_2\)

coef_one coefficient per feature

intercept_the fitted intercept

n_features_in_number of features seen by fit

The trailing underscore marks an attribute created from data. Accessing coef_ before fit fails because it does not exist yet.

predict applies the fitted equation row by row

X_new = np.array([
    [0., 2.],
    [3., 0.],
])

y_pred = model.predict(X_new)
# array([3., 7.])
X_new.shape(2, 2)y_pred.shape(2,)
ROW 0[0., 2.]\(1 + 2(0) + 1(2)\)3.0
ROW 1[3., 0.]\(1 + 2(3) + 1(0)\)7.0
Feature order must match training.Column 0 still means \(x_1\) and column 1 still means \(x_2\).

predict returns one value for each input row and preserves the row order.

Configuration and learned state

HYPERPARAMETERfit_intercept=True

A constructor setting chosen before training. It controls how fit behaves.

LEARNED MODEL PARAMETER\(\widehat\theta\)

The mathematical quantity estimated from X_train and y_train.

FITTED ATTRIBUTEcoef_, intercept_

The Python attributes that store the learned parameter after fit.

fit_interceptcontrols fitting\(\widehat\theta\)stored inside the objectcoef_ + intercept_

Hyperparameters control learning. Learned parameters are stored as fitted attributes.

Anatomy of a regressor estimator

from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.utils.validation import check_is_fitted

class RegressorTemplate(RegressorMixin, BaseEstimator):
    def __init__(self, setting=1):
        self.setting = setting

    def fit(self, X, y):
        self.learned_value_ = ...
        return self

    def predict(self, X):
        check_is_fitted(self, "learned_value_")
        return ...

initstores settingsNo training data and no fitted values.

fitlearns stateLearned public attributes end in _.

predictuses fitted stateIt never receives the unknown targets.

fit returns self, so Estimator().fit(X, y) is the fitted object.

Recommender baselines: same template, new models

04
CUHK emblem

Keep the four questions and estimator lifecycle fixed; change only the prediction rule and learned state.

Rating triples match the estimator input contract

TRAINING TRIPLES

useritemrating

19606704

13461524

78517414

split columns
X_trainuser–item pairs
[[1960,  670],
 [1346,  152],
 [ 785, 1741]]
shape: (n_samples, 2)
y_trainratings
[4, 4, 4]
shape: (n_samples,)
X_new[j] = (u, i)producesy_pred[j] = r_hat[u, i]same row, same position

fit(X, y) receives pairs and known ratings. predict(X_new) receives pairs only.

GlobalMeanRS: specify before coding

1 · MODEL\(f_c(u,i)=c\)

The family contains every constant prediction rule.

2 · HYPERPARAMETERSnone

No setting is chosen before fitting.

3 · LOSS / CRITERION\(J(c)=\sum_{(u,i,r)\in\mathcal R_{tr}}(r-c)^2\)

Compare candidate constants by their total squared error.

FIT · OPTIMIZE \[\widehat\mu =\arg\min_{c\in\mathbb R}\sum_{(u,i,r)\in\mathcal R_{tr}}(r-c)^2 =\frac{1}{|\mathcal R_{tr}|}\sum_{(u,i,r)\in\mathcal R_{tr}}r\]

4 · LEARNED PARAMETER AFTER fit\(\widehat\mu\)

stored as global_mean_ · predict returns it once per new row

Translate the specification into estimator methods

HYPERPARAMETERS→init(…)These basic baselines have none, so an explicit constructor is unnecessary.

LEARNED PARAMETERS→fit(X, y)Calculate them from training ratings and store them in attributes ending in _.

MODEL RULE→predict(X_new)Read the fitted attributes and return one prediction for each input row.

LOSS / CRITERION→explains the fitting ruleIt tells us why the learned mean is the chosen parameter value.

check_is_fitted(…)guards the boundarypredict should fail clearly when the learned parameters do not exist yet.

First specify the mathematics. Then each method has one clear job.

Now implement GlobalMeanRS

import numpy as np
from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.utils.validation import check_is_fitted

class GlobalMeanRS(RegressorMixin, BaseEstimator):
    def fit(self, X, y):
        self.global_mean_ = np.asarray(y).mean()
        return self

    def predict(self, X):
        check_is_fitted(self, "global_mean_")
        return np.full(len(X), self.global_mean_)

FITone learned valueglobal_mean_ comes from y.

PREDICTone value per rowX determines only how many predictions to return.

CHECKfit must happen firstOtherwise scikit-learn raises NotFittedError.

One object moves through the full lifecycle

01model = GlobalMeanRS()configured, not fitted
02model.fit(X_train, y_train)learns global_mean_
03model.global_mean_inspect fitted state
04y_pred = model.predict(X_test)one prediction per test pair
SAME OBJECTConfiguration remains. Fitted attributes appear after fit.

UserMeanRS: specify before coding

1 · MODEL\(f_{\mathbf a}(u,i)=a_u\)

One candidate constant per user; use the global rule for an unseen user.

2 · HYPERPARAMETERSnone

The grouping and fallback rule are fixed in this definition.

3 · LOSS / CRITERION\(J_u(c)=\sum_{i\in\mathcal I_u^{tr}}(r_{ui}-c)^2\)

For each observed user, compare candidate constants by squared error.

FIT · OPTIMIZE FOR EACH OBSERVED USER \(u\) \[\widehat a_u =\arg\min_{c\in\mathbb R}\sum_{i\in\mathcal I_u^{tr}}(r_{ui}-c)^2 =\frac{1}{|\mathcal I_u^{tr}|}\sum_{i\in\mathcal I_u^{tr}}r_{ui}\]

4 · LEARNED PARAMETERS AFTER fit\(\widehat\mu,\{\widehat a_u\}\)

stored as global_mean_ and user_means_ · retrieve by user ID

User mean learns a lookup vector

GROUP TRAINING RATINGS BY USER
\(u_0\)423mean = 3.0
\(u_1\)543mean = 4.0
\(u_2\)2mean = 2.0
store
FITTED ATTRIBUTES
global_mean_ = 3.3

user ID        0    1    2
user_means_ = [3.0, 4.0, 2.0]

Only a user absent from training uses global_mean_.

Every observed user receives a mean. The global mean handles users absent from training.

Now implement UserMeanRS

class UserMeanRS(RegressorMixin, BaseEstimator):
    def fit(self, X, y):
        X, y = np.asarray(X), np.asarray(y)
        users = X[:, 0].astype(int)
        self.global_mean_ = y.mean()
        self.user_means_ = np.full(
            users.max() + 1, self.global_mean_
        )

        for user in set(users):
            ratings = y[users == user]
            self.user_means_[user] = ratings.mean()
        return self

    def predict(self, X):
        check_is_fitted(self, ["global_mean_", "user_means_"])
        users = np.asarray(X)[:, 0].astype(int)
        y_pred = np.full(len(users), self.global_mean_)
        valid = (users >= 0) & (users < len(self.user_means_))
        y_pred[valid] = self.user_means_[users[valid]]
        return y_pred

global_mean_fallback learned by fit

user_means_NumPy lookup vector learned by fit

users == userselect that user’s ratings

validsafe array indices or fallback

Teaching implementation: expects a two-column array with user ID in column 0.

User mean predicts by lookup

model = UserMeanRS().fit(X_train, y_train)

X_new = np.array([
    [0, 17],   # user seen in training
    [71, 17],  # user absent from training
])

model.predict(X_new)
# array([3.36, 3.62])

ROW 0 · USER 0 user_means_[0] training ratings replaced the initial value 3.36

ROW 1 · USER 71 user_means_[71] no training ratings, so the global mean remains 3.62

The item ID keeps the input contract consistent, but this model’s prediction depends only on the user ID.

Values rounded from the STAT3009 Netflix training split.

ItemMeanRS: transfer the same template

1 · MODEL\(f_{\mathbf b}(u,i)=b_i\)

One candidate constant per item; use the global rule for an unseen item.

2 · HYPERPARAMETERSnone

The grouping and fallback rule are fixed in this definition.

3 · LOSS / CRITERION\(J_i(c)=\sum_{u\in\mathcal U_i^{tr}}(r_{ui}-c)^2\)

For each observed item, compare candidate constants by squared error.

FIT · OPTIMIZE FOR EACH OBSERVED ITEM \(i\) \[\widehat b_i =\arg\min_{c\in\mathbb R}\sum_{u\in\mathcal U_i^{tr}}(r_{ui}-c)^2 =\frac{1}{|\mathcal U_i^{tr}|}\sum_{u\in\mathcal U_i^{tr}}r_{ui}\]

4 · LEARNED PARAMETERS AFTER fit\(\widehat\mu,\{\widehat b_i\}\)

stored as global_mean_ and item_means_ · retrieve by item ID

What the base classes provide

YOUR CLASSfit and predict

The recommender-specific learning rule lives here.

RegressorMixinregressor identity and default score

The default regression score is \(R^2\).

BaseEstimatorget_params, set_params, representation

These support cloning and model-selection tools.

class MyRS(RegressorMixin, BaseEstimator):mixin on the left, BaseEstimator on the right

Evaluation metrics stay outside the estimator

from sklearn.metrics import root_mean_squared_error

model = UserMeanRS()
model.fit(X_train, y_train)

y_pred = model.predict(X_test)
rmse = root_mean_squared_error(y_test, y_pred)
ESTIMATORlearns and predicts

fit and predict define the shared model interface.

METRICjudges predictions

The course reports RMSE so we call the metric explicitly.

Do not read model.score(…) as RMSE.RegressorMixin.score returns \(R^2\) by default, where larger is better.

One workflow can use every estimator

models = [
    GlobalMeanRS(),
    UserMeanRS(),
]

for model in models:
    fitted_model = model.fit(X_train, y_train)
    y_pred = fitted_model.predict(X_new)

    print(type(model).__name__, y_pred.shape)
OUTPUT
GlobalMeanRS  (m,)
UserMeanRS    (m,)

Same inputsX_train, y_train, X_new

Same methodsfit and predict

Same output contractone prediction per new row

Model-specific logic stays inside each class. Workflow code depends only on the shared estimator interface.

One interface—and one way to read every estimator

ASK FOUR QUESTIONS model learned parameters hyperparameters loss / criterion

then use
fit(X, y)predict(X_new)
next lecture
ML IIWhich fitted estimator will generalize?

Holdout validation and cross-validation compare candidates without using the final test labels.

For every new estimator: answer the four questions to understand it, then use the familiar fit / predict interface.