From prediction rules to reusable code
CUHK · Department of Statistics and Data Science
How do we move from hand-designed averages to a general way of learning prediction rules from data?
Then calculate its averages from the training triples.
Then learn its parameters from examples.
Work well on ratings the model has never seen.
features \(X_{train}\) + targets \(y_{train}\)
\(\widehat f\) stores patterns learned from training data
\(X_{new} \rightarrow \widehat y\)
compare \(\widehat y\) with untouched answers
X_test is available. y_test stays out of fitting and model selection.
Examples, models, losses, parameters—and the separation between fitting and evaluation.
Supervised learning: learn a mapping from inputs to a known numerical target, then apply it to new inputs.
| user | item | rating |
|---|---|---|
| 1960 | 670 | 4 |
| 1346 | 152 | 4 |
| 785 | 1741 | 4 |
\(X\) contains the inputs; \(y\) contains the answer.
the user–item pair
the observed rating
The row number is not a feature. It only identifies an observation.
What family of prediction rules can this estimator represent?
What quantities will fit learn from the training data?
What choices are fixed before fit and control the model or fitting rule?
What does the fitting procedure regard as a better solution?
\[h\ \text{chosen before fit}\quad\Longrightarrow\quad \widehat\theta=\arg\min_{\theta}\frac{1}{n}\sum_{i=1}^{n}L\!\left(y_i,f_{\theta,h}(x_i)\right) \quad\Longrightarrow\quad \widehat f=f_{\widehat\theta,h}\]
fit · CHOSEN hyperparameter \(h\) fit_intercept=True
Set by us—or later selected with validation—to control the model or fitting process.
fit(X, y)uses training data
fit · LEARNED parameter \(\widehat\theta\) coef_ · intercept_
Estimated from the observed examples and stored in the fitted object.
get_params(). In our ML template, we call model settings hyperparameters to distinguish them from quantities learned during fit.
Quick test: did the data determine the value? Yes → learned parameter. No, fixed before fitting → hyperparameter.
The formula gives a family. Each value of \(\theta=(\theta_0,\theta_1)\) gives a different member.
candidate A candidate B candidate C
Training decides which member of the family to use.
Fix \(\theta^{(B)}\). For every training row, predict \(\widehat y_i=f_{\theta^{(B)}}(x_i)\) and calculate \(\ell_i=(y_i-\widehat y_i)^2\).
| row \(i\) | observed \(y_i\) | prediction \(\widehat y_i\) | loss \(\ell_i\) |
|---|---|---|---|
| 1 | 4.0 | 3.5 | 0.25 |
| 2 | 2.0 | 2.5 | 0.25 |
| 3 | 5.0 | 4.0 | 1.00 |
AVERAGE THE ROW LOSSES \[J\!\left(\theta^{(B)}\right)=\frac{0.25+0.25+1.00}{3}=0.50\]
For squared loss, minimizing MSE also minimizes RMSE. A small training loss still does not guarantee good predictions on new rows.
\(\widehat\theta\) is now fixed.
features only
\(=\widehat f(x_{new})\)
.fit(X_train, y_train)changes the model by learning parameters
.predict(X_new)uses the fitted model without learning from new labels
X_train and y_train are available
X_test is available for prediction
y_test does not estimate parameters
y_test is opened only for final evaluation
Generalization is performance on new examples drawn from the task we care about.
X_traintraining inputs
y_traintraining answers
X_testtest user–item pairs
y_test
opened only after all modeling choices are fixed
MODEL SELECTION QUESTION How can we compare model A and model B while y_test remains hidden? ML II introduces holdout validation and cross-validation.
California housing: define \(X\) and \(y\), fit linear regression, predict, and calculate RMSE.
California block groups
demographic and geographic measurements
median house value in units of $100,000
MedIncHouseAgeAveRoomsAveBedrmsPopulationAveOccupLatitudeLongitude
Source: scikit-learn California housing dataset documentation.
one row = one block group
one column = one feature
one numerical answer
for every row of \(X\)
X_train.shape(13828, 8)y_train.shape(13828,)
\[\widehat y=\theta_0+\theta_1x_1+\theta_2x_2+\cdots+\theta_8x_8\]
the model’s reference level
how prediction changes with feature \(j\), holding others fixed
a weighted sum calculated for one row
The model family is linear. The coefficients are learned from the training data.
feature combinationhouse value
Vertical distance between the observed target and the fitted prediction.
LEAST SQUARES FIT \[\widehat\theta=\arg\min_{\theta}\sum_{i=1}^{n}e_i^2\]
Squaring makes large errors count more and prevents positive and negative errors from cancelling.
LinearRegression with the four questionsA linear family: one intercept plus one coefficient for each feature.
Stored after fit as intercept_ and coef_.
fit_intercept=True
Chosen before fitting; controls whether the model includes an intercept.
Ordinary least squares chooses coefficients with the smallest residual sum of squares.
1ConstructChoose the model family.
2FitLearn coefficients from training examples.
3PredictProduce one value for every test row.
4EvaluateSummarize the prediction errors.
fit sees targets. predict sees features only.
TRUE\(y\)2.41.83.22.6
PREDICTED\(\widehat y\)2.12.02.72.8
ERROR\(e\)0.3−0.20.5−0.2
0.09, 0.04, 0.25, 0.04
\((0.42)/4=0.105\)
\(\mathrm{RMSE}=0.324\)
A shared contract turns different learning rules into reusable prediction objects.
model = Estimator(…)
Choose the model and its settings.
model.fit(X, y)
Store patterns learned from the training data.
model.predict(X_new)
Return one prediction for each new row.
LinearRegression()KNeighborsRegressor()YourRegressor()can share the same interface
The contract specifies how we use a model. Each estimator supplies its own learning and prediction rules.
model
CONFIGURATION attributes before fit fit_intercept positive
LEARNED STATE attributes after fit coef_ intercept_
1 2 3 4 5
fit(X, y)writes learned state
predict(X_new)reads learned state
score(X, y)evaluates
get_params()reads configuration
set_params(…)changes configuration
Octopus illustration: Jason J. Patterson / Open Clip Art Library, CC0.
Estimator(hyperparameters)configuration
fit(X, y) returns selflearn from data
predict(X_new) returns y_predapply learned state
learned_attribute_fitted state
LinearRegression
learn coefficients
coef_
KNeighborsRegressor
store training cases
algorithm-specific stateYourRegressor
supply your own rule
learned_state_
LinearRegression
class LinearRegression(...):
def __init__(...)
def fit(self, X, y)
def predict(self, X)
Defines which parameters, methods, and fitted attributes an object may have.
call the classcreates
model = LinearRegression(
fit_intercept=True)
fit_intercept
coef_
A configured Python object, ready to receive data.
model.fit(X, y)updates
model
fit_intercept
coef_
intercept_
The same object now contains fitted state.
model.fit(X, y) is modelTruefit conventionally returns self.
X_trainfour samples, two featuresEach row describes one training case.
y_trainfour known targetsPosition j matches row j of X_train.
modelan unfitted estimatorIt knows its settings but has not learned a line.
LinearRegression(…) creates a Python object. Training happens only when we call fit.
init(…)records hyperparametersThese settings describe how a future fit should behave.
get_params()returns constructor settingsScikit-learn calls them estimator “parameters” in its API.
set_params(…)changes hyperparametersFit the estimator again after changing a setting.
init.X and y enter through fit(X, y).
X_train
[[1., 0.], [0., 1.], [1., 1.], [2., 1.]]
shape (4, 2)
y_train
[3., 2., 4., 6.]
shape (4,)
model.fit(X_train, y_train)
model
learned coefficients now live inside the estimator
returned is model # True
X_train[j] and y_train[j] form one supervised example.fit returns self, which makes LinearRegression().fit(X, y) valid syntax.
\(\widehat y = 1.0 + 2.0x_1 + 1.0x_2\)
coef_one coefficient per feature
intercept_the fitted intercept
n_features_in_number of features seen by fit
The trailing underscore marks an attribute created from data. Accessing coef_ before fit fails because it does not exist yet.
X_new.shape(2, 2)y_pred.shape(2,)
[0., 2.]\(1 + 2(0) + 1(2)\)3.0
[3., 0.]\(1 + 2(3) + 1(0)\)7.0
predict returns one value for each input row and preserves the row order.
fit_intercept=True
A constructor setting chosen before training. It controls how fit behaves.
The mathematical quantity estimated from X_train and y_train.
coef_, intercept_
The Python attributes that store the learned parameter after fit.
fit_interceptcontrols fitting\(\widehat\theta\)stored inside the objectcoef_ + intercept_
Hyperparameters control learning. Learned parameters are stored as fitted attributes.
from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.utils.validation import check_is_fitted
class RegressorTemplate(RegressorMixin, BaseEstimator):
def __init__(self, setting=1):
self.setting = setting
def fit(self, X, y):
self.learned_value_ = ...
return self
def predict(self, X):
check_is_fitted(self, "learned_value_")
return ...
initstores settingsNo training data and no fitted values.
fitlearns stateLearned public attributes end in _.
predictuses fitted stateIt never receives the unknown targets.
fit returns self, so Estimator().fit(X, y) is the fitted object.
Keep the four questions and estimator lifecycle fixed; change only the prediction rule and learned state.
useritemrating
19606704
13461524
78517414
X_trainuser–item pairs
[[1960, 670], [1346, 152], [ 785, 1741]]shape:
(n_samples, 2)
y_trainratings
[4, 4, 4]shape:
(n_samples,)
X_new[j] = (u, i)producesy_pred[j] = r_hat[u, i]same row, same position
fit(X, y) receives pairs and known ratings. predict(X_new) receives pairs only.
GlobalMeanRS: specify before codingThe family contains every constant prediction rule.
No setting is chosen before fitting.
Compare candidate constants by their total squared error.
FIT · OPTIMIZE \[\widehat\mu =\arg\min_{c\in\mathbb R}\sum_{(u,i,r)\in\mathcal R_{tr}}(r-c)^2 =\frac{1}{|\mathcal R_{tr}|}\sum_{(u,i,r)\in\mathcal R_{tr}}r\]
fit\(\widehat\mu\)
stored as global_mean_ · predict returns it once per new row
HYPERPARAMETERS→init(…)These basic baselines have none, so an explicit constructor is unnecessary.
LEARNED PARAMETERS→fit(X, y)Calculate them from training ratings and store them in attributes ending in _.
MODEL RULE→predict(X_new)Read the fitted attributes and return one prediction for each input row.
LOSS / CRITERION→explains the fitting ruleIt tells us why the learned mean is the chosen parameter value.
check_is_fitted(…)guards the boundarypredict should fail clearly when the learned parameters do not exist yet.
First specify the mathematics. Then each method has one clear job.
GlobalMeanRSimport numpy as np
from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.utils.validation import check_is_fitted
class GlobalMeanRS(RegressorMixin, BaseEstimator):
def fit(self, X, y):
self.global_mean_ = np.asarray(y).mean()
return self
def predict(self, X):
check_is_fitted(self, "global_mean_")
return np.full(len(X), self.global_mean_)
FITone learned valueglobal_mean_ comes from y.
PREDICTone value per rowX determines only how many predictions to return.
CHECKfit must happen firstOtherwise scikit-learn raises NotFittedError.
model = GlobalMeanRS()configured, not fitted
model.fit(X_train, y_train)learns global_mean_
model.global_mean_inspect fitted state
y_pred = model.predict(X_test)one prediction per test pair
fit.
UserMeanRS: specify before codingOne candidate constant per user; use the global rule for an unseen user.
The grouping and fallback rule are fixed in this definition.
For each observed user, compare candidate constants by squared error.
FIT · OPTIMIZE FOR EACH OBSERVED USER \(u\) \[\widehat a_u =\arg\min_{c\in\mathbb R}\sum_{i\in\mathcal I_u^{tr}}(r_{ui}-c)^2 =\frac{1}{|\mathcal I_u^{tr}|}\sum_{i\in\mathcal I_u^{tr}}r_{ui}\]
fit\(\widehat\mu,\{\widehat a_u\}\)
stored as global_mean_ and user_means_ · retrieve by user ID
global_mean_ = 3.3 user ID 0 1 2 user_means_ = [3.0, 4.0, 2.0]
Only a user absent from training uses global_mean_.
Every observed user receives a mean. The global mean handles users absent from training.
UserMeanRSclass UserMeanRS(RegressorMixin, BaseEstimator):
def fit(self, X, y):
X, y = np.asarray(X), np.asarray(y)
users = X[:, 0].astype(int)
self.global_mean_ = y.mean()
self.user_means_ = np.full(
users.max() + 1, self.global_mean_
)
for user in set(users):
ratings = y[users == user]
self.user_means_[user] = ratings.mean()
return self
def predict(self, X):
check_is_fitted(self, ["global_mean_", "user_means_"])
users = np.asarray(X)[:, 0].astype(int)
y_pred = np.full(len(users), self.global_mean_)
valid = (users >= 0) & (users < len(self.user_means_))
y_pred[valid] = self.user_means_[users[valid]]
return y_pred
global_mean_fallback learned by fit
user_means_NumPy lookup vector learned by fit
users == userselect that user’s ratings
validsafe array indices or fallback
Teaching implementation: expects a two-column array with user ID in column 0.
ROW 0 · USER 0 user_means_[0] training ratings replaced the initial value 3.36
ROW 1 · USER 71 user_means_[71] no training ratings, so the global mean remains 3.62
The item ID keeps the input contract consistent, but this model’s prediction depends only on the user ID.
Values rounded from the STAT3009 Netflix training split.
ItemMeanRS: transfer the same templateOne candidate constant per item; use the global rule for an unseen item.
The grouping and fallback rule are fixed in this definition.
For each observed item, compare candidate constants by squared error.
FIT · OPTIMIZE FOR EACH OBSERVED ITEM \(i\) \[\widehat b_i =\arg\min_{c\in\mathbb R}\sum_{u\in\mathcal U_i^{tr}}(r_{ui}-c)^2 =\frac{1}{|\mathcal U_i^{tr}|}\sum_{u\in\mathcal U_i^{tr}}r_{ui}\]
fit\(\widehat\mu,\{\widehat b_i\}\)
stored as global_mean_ and item_means_ · retrieve by item ID
fit and predict
The recommender-specific learning rule lives here.
RegressorMixinregressor identity and default score
The default regression score is \(R^2\).
BaseEstimatorget_params, set_params, representation
These support cloning and model-selection tools.
class MyRS(RegressorMixin, BaseEstimator):mixin on the left, BaseEstimator on the right
fit and predict define the shared model interface.
The course reports RMSE so we call the metric explicitly.
model.score(…) as RMSE.RegressorMixin.score returns \(R^2\) by default, where larger is better.
GlobalMeanRS (m,) UserMeanRS (m,)
Same inputsX_train, y_train, X_new
Same methodsfit and predict
Same output contractone prediction per new row
Model-specific logic stays inside each class. Workflow code depends only on the shared estimator interface.
ASK FOUR QUESTIONS model learned parameters hyperparameters loss / criterion
fit(X, y)predict(X_new)
Holdout validation and cross-validation compare candidates without using the final test labels.
For every new estimator: answer the four questions to understand it, then use the familiar fit / predict interface.
CUHK · STAT3009 · Recommender Systems