Comprehensive AI & Machine Learning Documentation

Complete AI & ML Guide

From NumPy arrays to Transformers and Generative AI. This documentation covers modern AI frameworks, deep learning techniques, and practical ML workflows — with practical examples, deep explanations, and best practices used across research and industry.

1 Introduction to AI & Machine Learning

What AI and ML are, and how they relate

Artificial Intelligence (AI) is the broad field of building systems that perform tasks normally requiring human intelligence. Machine Learning (ML) is a subset of AI in which systems learn patterns from data instead of following explicitly programmed rules. Deep Learning is, in turn, a subset of ML built on multi-layered neural networks capable of learning very complex patterns directly from raw data.

The AI Hierarchy

LayerDescriptionExample
Artificial IntelligenceAny technique that lets machines mimic intelligent behaviorRule-based chess engines, expert systems
Machine LearningSystems that learn from data rather than fixed rulesSpam filters, recommendation engines
Deep LearningML using multi-layer neural networksImage recognition, language models

Why Python Dominates AI/ML

  • Rich ecosystem: NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch all mature and interoperable
  • Readable syntax: Lets researchers focus on ideas instead of boilerplate
  • Community & research adoption: Most papers ship Python reference implementations
  • Interoperability: Easy to bind to fast C/C++/CUDA backends under the hood

A minimal "hello world" of machine learning.

python
from sklearn.linear_model import LinearRegression
python
X = [[1], [2], [3], [4]] y = [2, 4, 6, 8]
python
model = LinearRegression() model.fit(X, y) print(model.predict([[5]]))
Note: This documentation uses Python throughout, since it's the de-facto standard language across research and industry for AI and machine learning.

2 Types of Machine Learning

The three core learning paradigms

TypeDataGoalExamples
Supervised LearningLabeled (input, correct output) pairsPredict labels for new inputsSpam detection, price prediction
Unsupervised LearningUnlabeled dataDiscover hidden structureCustomer segmentation, anomaly detection
Reinforcement LearningEnvironment + reward signalLearn a policy maximizing rewardGame-playing agents, robotics
Semi-supervisedSmall labeled + large unlabeled setCombine both signalsMedical imaging with scarce labels
Self-supervisedUnlabeled data with generated pseudo-labelsLearn general representationsPretraining LLMs, contrastive vision models

Supervised Learning: Two Main Tasks

  • Regression: Predict a continuous number (e.g., house price)
  • Classification: Predict a discrete category (e.g., spam or not spam)
Tip: When starting a new project, first ask: "Do I have labels?" — that single question usually decides between supervised and unsupervised approaches.

3 Math Foundations

The minimum math you need to understand what's happening

Linear Algebra

Data in ML is represented as vectors and matrices. A neural network layer is fundamentally a matrix multiplication followed by a non-linear function.

python
import numpy as np

Vector and matrix.

python
v = np.array([1, 2, 3]) M = np.array([[1, 2], [3, 4]])

Dot product (core operation in neural network layers).

python
result = np.dot(M, v[:2])

Matrix transpose and inverse.

python
M.T np.linalg.inv(M)

Calculus (Gradients)

Training a model means minimizing a loss function. Gradient descent uses the derivative of the loss with respect to each parameter to know which direction reduces error, and steps the parameters that way, repeatedly.

Probability & Statistics

  • Distributions: Understanding how data is spread (normal, uniform, etc.)
  • Mean, variance, standard deviation: Core descriptive statistics used everywhere in preprocessing
  • Bayes' theorem: Foundation of probabilistic models and Naive Bayes classifiers
  • Conditional probability: Central to language models predicting the next token
Key Insight: You don't need to derive backpropagation by hand to use deep learning frameworks effectively — but understanding gradients conceptually helps you debug training issues like exploding or vanishing gradients.

4 Environment Setup

Getting your AI/ML development environment ready

Create an isolated virtual environment.

bash
python3 -m venv ai-env
bash
source ai-env/bin/activate

Linux/Mac.

bash
ai-env\Scripts\activate

Windows.

Or with conda (popular in the ML community)

bash
conda create -n ai-env python=3.11
bash
conda activate ai-env

Install the core stack.

bash
pip install numpy pandas matplotlib scikit-learn

Deep learning frameworks.

bash
pip install tensorflow
bash
pip install torch torchvision torchaudio

Jupyter for interactive notebooks.

bash
pip install jupyterlab
bash
jupyter lab

Checking GPU Availability

python
import torch print(torch.cuda.is_available())
python
import tensorflow as tf print(tf.config.list_physical_devices('GPU'))
Tip: Use Google Colab or Kaggle Notebooks for free GPU access if you don't have a local GPU — both come with the core AI/ML stack pre-installed.

5 NumPy Fundamentals

The array library everything else in AI/ML is built on

NumPy provides the ndarray, a fast, memory-efficient, multi-dimensional array type with vectorized operations that avoid slow Python loops. Every major ML/DL framework either uses NumPy directly or mirrors its API.

python
import numpy as np

Creating arrays.

python
a = np.array([1, 2, 3]) zeros = np.zeros((3, 4)) ones = np.ones((2, 2)) rand = np.random.rand(3, 3)

Shape, dtype, reshape.

python
print(a.shape, a.dtype) b = np.arange(12).reshape(3, 4)

Vectorized operations (no explicit loops).

python
c = a * 2 + 1 d = np.sqrt(a)

Broadcasting: operate on arrays of different shapes.

python
M = np.ones((3, 3)) row = np.array([1, 2, 3]) result = M + row

Indexing and slicing.

python
b[0, :] b[:, 1] b[b > 5]

Common Operations

FunctionPurpose
np.mean / np.stdDescriptive statistics along an axis
np.dot / @Matrix multiplication
np.concatenate / np.stackCombine arrays
np.whereConditional element selection
np.linalg.normVector/matrix norms
Key Insight: Prefer vectorized NumPy operations over Python for loops — they run in optimized C code and can be 10-100x faster on large datasets.

6 Pandas for Data Manipulation

Loading, cleaning, and exploring tabular data

Pandas introduces the DataFrame, a labeled 2D table structure ideal for real-world, messy datasets — CSVs, SQL query results, spreadsheets, and logs.

python
import pandas as pd

Loading data.

python
df = pd.read_csv('data.csv')

Quick exploration.

python
df.head() df.info() df.describe() df.shape

Selecting columns and rows.

python
df['age'] df[['age', 'income']] df.loc[df['age'] > 30] df.iloc[0:5]

Handling missing values.

python
df.isnull().sum() df.dropna() df.fillna(df.mean(numeric_only=True))

Grouping and aggregation.

python
df.groupby('category')['sales'].sum()

Merging datasets.

python
merged = pd.merge(df1, df2, on='id', how='left')
Tip: Most real-world ML time is spent here, not on modeling — data cleaning and exploration typically take 60-80% of a project's time.

7 Data Visualization

Seeing your data before modeling it

Visualization surfaces distributions, outliers, correlations, and class imbalance that summary statistics alone can hide. Matplotlib is the low-level foundation; Seaborn builds statistical plots on top of it with far less code.

python
import matplotlib.pyplot as plt import seaborn as sns

Distribution of a single feature.

python
sns.histplot(df['age'], kde=True)

Relationship between two features.

python
sns.scatterplot(data=df, x='income', y='spending', hue='segment')

Correlation heatmap.

python
sns.heatmap(df.corr(numeric_only=True), annot=True, cmap='coolwarm')

Box plot to spot outliers.

python
sns.boxplot(data=df, x='category', y='price')
python
plt.show()

Common Plot Types

PlotBest For
Histogram / KDEFeature distributions
Scatter plotRelationship between two variables
Box / violin plotOutliers and spread by category
Correlation heatmapFeature interdependence
Confusion matrixClassification error patterns

8 Data Preprocessing & Feature Engineering

Turning raw data into model-ready features

Scaling & Encoding

python
from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.model_selection import train_test_split

Feature scaling (mean 0, std 1) — required for many algorithms.

python
scaler = StandardScaler() X_scaled = scaler.fit_transform(X_train)

One-hot encoding categorical variables.

python
encoder = OneHotEncoder(sparse_output=False) X_cat = encoder.fit_transform(df[['category']])

Train/test split.

python
X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 )

Common Preprocessing Steps

  • Handling missing data: Imputation (mean/median/mode) or removal
  • Feature scaling: Standardization or min-max normalization
  • Encoding categories: One-hot, ordinal, or target encoding
  • Feature engineering: Creating new informative features (ratios, dates split into day/month, text length)
  • Outlier handling: Clipping, transformation, or removal
  • Class imbalance: Oversampling (SMOTE), undersampling, or class weights
Warning: Always fit scalers/encoders on the training set only, then transform the test set with the same fitted object — fitting on the full dataset leaks test information into training and inflates reported performance.

9 Scikit-learn Basics

The standard toolkit for classical machine learning

Scikit-learn provides a consistent API — fit, predict, transform — across dozens of algorithms, plus utilities for preprocessing, model selection, and evaluation.

python
from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler from sklearn.ensemble import RandomForestClassifier

Pipelines chain preprocessing + model into one object.

python
pipeline = Pipeline([ ('scaler', StandardScaler()), ('model', RandomForestClassifier(n_estimators=100)) ])
python
pipeline.fit(X_train, y_train) predictions = pipeline.predict(X_test) accuracy = pipeline.score(X_test, y_test)
Tip: Wrapping preprocessing and modeling in a Pipeline prevents data leakage automatically and makes cross-validation and deployment far simpler.

10 Regression Algorithms

Predicting continuous values

AlgorithmIdeaWhen to Use
Linear RegressionFits a straight line/hyperplane minimizing squared errorSimple, interpretable baselines
Ridge / LassoLinear regression with L2/L1 penalty on coefficientsMany features, need regularization
Polynomial RegressionLinear regression on polynomial feature expansionsNon-linear but smooth relationships
Decision Tree RegressorSplits data into regions, predicts region averageNon-linear, interpretable
Random Forest / Gradient BoostingEnsembles of treesStrong tabular baselines
SVRSupport vector machine adapted for regressionSmall-to-medium, high-dimensional data
python
from sklearn.linear_model import Ridge from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import mean_squared_error, r2_score
python
model = RandomForestRegressor(n_estimators=200, random_state=42) model.fit(X_train, y_train) preds = model.predict(X_test)
python
print("RMSE:", mean_squared_error(y_test, preds, squared=False)) print("R2:", r2_score(y_test, preds))

11 Classification Algorithms

Predicting discrete categories

AlgorithmIdeaWhen to Use
Logistic RegressionLinear decision boundary via sigmoid functionFast, interpretable baseline
k-Nearest NeighborsClassifies by majority vote of nearest pointsSmall datasets, simple boundaries
Decision TreeRecursive feature-based splitsInterpretable, non-linear
Random ForestEnsemble of decision trees, majority voteStrong general-purpose baseline
Gradient Boosting (XGBoost/LightGBM)Sequentially corrects previous trees' errorsTop performer on tabular data
SVMMaximizes margin between classesHigh-dimensional, clear-margin data
Naive BayesProbabilistic, assumes feature independenceText classification, spam filtering
python
from sklearn.linear_model import LogisticRegression from sklearn.metrics import classification_report, confusion_matrix
python
clf = LogisticRegression(max_iter=1000) clf.fit(X_train, y_train) preds = clf.predict(X_test)
python
print(classification_report(y_test, preds)) print(confusion_matrix(y_test, preds))

12 Clustering & Dimensionality Reduction

Finding structure without labels

Clustering

AlgorithmIdea
K-MeansPartitions data into k clusters by minimizing distance to centroids
Hierarchical ClusteringBuilds a tree of nested clusters (dendrogram)
DBSCANDensity-based; finds arbitrarily shaped clusters and outliers
Gaussian Mixture ModelsSoft, probabilistic clustering

Dimensionality Reduction

python
from sklearn.decomposition import PCA from sklearn.cluster import KMeans

PCA: project high-dimensional data onto principal components.

python
pca = PCA(n_components=2) X_reduced = pca.fit_transform(X_scaled) print(pca.explained_variance_ratio_)

K-Means clustering.

python
kmeans = KMeans(n_clusters=3, random_state=42, n_init='auto') labels = kmeans.fit_predict(X_scaled)
Key Insight: PCA and t-SNE/UMAP are used both to speed up modeling on high-dimensional data and to visualize it in 2D/3D for exploration.

13 Ensemble Methods

Combining multiple models for better performance

  • Bagging: Train many models on random data subsets in parallel and average results (e.g., Random Forest)
  • Boosting: Train models sequentially, each correcting the previous one's errors (e.g., XGBoost, LightGBM, CatBoost, AdaBoost)
  • Stacking: Train a meta-model on the outputs of several base models
  • Voting: Combine predictions from different model types by majority vote or averaging
python
from xgboost import XGBClassifier
python
model = XGBClassifier( n_estimators=300, learning_rate=0.05, max_depth=6, subsample=0.8 ) model.fit(X_train, y_train)
Tip: Gradient-boosted trees (XGBoost, LightGBM, CatBoost) consistently win on structured/tabular data — reach for them before deep learning in that setting.

14 Model Evaluation Metrics

Measuring what actually matters for your problem

Classification Metrics

MetricFormula / IdeaBest For
AccuracyCorrect / TotalBalanced classes
PrecisionTP / (TP + FP)Cost of false positives is high
RecallTP / (TP + FN)Cost of false negatives is high
F1-scoreHarmonic mean of precision/recallImbalanced classes
ROC-AUCArea under TPR vs FPR curveRanking quality across thresholds

Regression Metrics

MetricIdea
MAEAverage absolute error, robust to outliers
MSE / RMSEPenalizes larger errors more heavily
R²Proportion of variance explained by the model
Warning: On imbalanced datasets, accuracy can be misleading — a model predicting "no fraud" 99% of the time can be 99% accurate while being useless. Use precision, recall, F1, or ROC-AUC instead.

15 Overfitting, Underfitting & Regularization

Building models that generalize

ProblemSymptomFix
UnderfittingPoor performance on both train and test dataMore complex model, more/better features, less regularization
OverfittingGreat on training data, poor on test dataMore data, regularization, simpler model, early stopping

Regularization Techniques

  • L1 (Lasso): Pushes some weights to exactly zero — built-in feature selection
  • L2 (Ridge): Shrinks weights smoothly, discourages large coefficients
  • Dropout: Randomly disables neurons during training (deep learning)
  • Early stopping: Stop training once validation loss stops improving
  • Data augmentation: Artificially expand training data (flips, crops, noise)
Key Insight: The gap between training and validation performance is your best diagnostic — a large gap means overfitting, and consistently poor performance on both means underfitting.

16 Cross-Validation & Hyperparameter Tuning

Getting a reliable estimate of real-world performance

python
from sklearn.model_selection import cross_val_score, GridSearchCV from sklearn.ensemble import RandomForestClassifier

K-Fold cross-validation.

python
scores = cross_val_score(RandomForestClassifier(), X, y, cv=5) print("Mean accuracy:", scores.mean())

Grid search over hyperparameters.

python
param_grid = { 'n_estimators': [100, 200, 300], 'max_depth': [None, 10, 20] } grid = GridSearchCV(RandomForestClassifier(), param_grid, cv=5) grid.fit(X_train, y_train) print(grid.best_params_)
Tip: Use RandomizedSearchCV or Bayesian optimization (Optuna) instead of full grid search when the hyperparameter space is large — it finds near-optimal settings far faster.

17 Neural Network Fundamentals

The building blocks of deep learning

A neural network is composed of layers of neurons, each computing a weighted sum of its inputs followed by a non-linear activation function. Stacking layers lets the network approximate arbitrarily complex functions.

Key Components

ComponentRole
Weights & BiasesLearnable parameters adjusted during training
Activation FunctionIntroduces non-linearity (ReLU, Sigmoid, Tanh, Softmax)
Loss FunctionMeasures prediction error (Cross-Entropy, MSE)
OptimizerUpdates weights to reduce loss (SGD, Adam, RMSprop)
BackpropagationComputes gradients of the loss w.r.t. every weight via the chain rule
Key Insight: Training is a loop: forward pass (predict) → compute loss → backward pass (gradients) → optimizer step (update weights) — repeated over many epochs until the loss converges.

18 Deep Learning with TensorFlow & Keras

Google's production-grade deep learning framework

Keras, now TensorFlow's high-level API, makes building and training neural networks concise and readable while TensorFlow handles graph optimization, GPU/TPU execution, and deployment.

python
import tensorflow as tf from tensorflow import keras
python
model = keras.Sequential([ keras.layers.Dense(64, activation='relu', input_shape=(20,)), keras.layers.Dropout(0.3), keras.layers.Dense(32, activation='relu'), keras.layers.Dense(1, activation='sigmoid') ])
python
model.compile( optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'] )
python
history = model.fit( X_train, y_train, validation_split=0.2, epochs=20, batch_size=32 )
python
model.evaluate(X_test, y_test)
Tip: Use keras.callbacks.EarlyStopping and ModelCheckpoint to automatically stop training and save the best-performing weights once validation loss stops improving.

19 Deep Learning with PyTorch

The dominant framework in AI research

PyTorch favors an imperative, "define-by-run" style that feels like regular Python, making it especially popular in research where models change often.

python
import torch import torch.nn as nn
python
class Net(nn.Module): def __init__(self): super().__init__() self.fc1 = nn.Linear(20, 64) self.fc2 = nn.Linear(64, 1) self.relu = nn.ReLU() self.sigmoid = nn.Sigmoid() def forward(self, x): x = self.relu(self.fc1(x)) return self.sigmoid(self.fc2(x))
python
model = Net() criterion = nn.BCELoss() optimizer = torch.optim.Adam(model.parameters(), lr=0.001)
python
for epoch in range(20): optimizer.zero_grad() outputs = model(X_train) loss = criterion(outputs, y_train) loss.backward() optimizer.step()
Note: TensorFlow and PyTorch are functionally comparable today — PyTorch dominates research and LLM development, while TensorFlow (via TF Serving/Lite) remains strong in production and mobile/edge deployment.

20 Convolutional Neural Networks (CNNs)

The architecture behind modern computer vision

Convolutional layers slide small learnable filters across an image to detect local patterns like edges, textures, and shapes — with far fewer parameters than a fully-connected layer would need.

Core Layers

LayerRole
ConvolutionExtracts local spatial features using learnable filters
Pooling (Max/Avg)Downsamples feature maps, adds translation invariance
FlattenConverts 2D feature maps into a 1D vector
Fully ConnectedCombines features for final classification
python
model = keras.Sequential([ keras.layers.Conv2D(32, (3,3), activation='relu', input_shape=(64,64,3)), keras.layers.MaxPooling2D((2,2)), keras.layers.Conv2D(64, (3,3), activation='relu'), keras.layers.MaxPooling2D((2,2)), keras.layers.Flatten(), keras.layers.Dense(128, activation='relu'), keras.layers.Dense(10, activation='softmax') ])
Tip: Well-known CNN architectures — ResNet, EfficientNet, VGG — are pretrained on ImageNet and can be fine-tuned on your own image data instead of training from scratch.

21 Recurrent Networks & Sequence Models

Modeling sequences: text, time series, audio

RNNs process sequences step-by-step, carrying a hidden state forward so earlier inputs can influence later predictions. LSTM and GRU cells add gating mechanisms that solve the vanishing-gradient problem of plain RNNs, letting them retain information over longer sequences.

ArchitectureBest For
Simple RNNShort sequences, educational baselines
LSTMLonger sequences, time series, language modeling
GRUSimilar to LSTM, fewer parameters, faster to train
Seq2Seq (Encoder-Decoder)Translation, summarization

A simple LSTM-based classifier: an embedding layer, an LSTM layer, and a dense output layer.

python
model = keras.Sequential([ keras.layers.Embedding(input_dim=10000, output_dim=64), keras.layers.LSTM(64, return_sequences=False), keras.layers.Dense(1, activation='sigmoid') ])
Note: RNNs/LSTMs have largely been superseded by Transformers for most NLP tasks, but remain relevant for streaming and resource-constrained time-series applications.

22 Transformers & Attention Mechanism

The architecture behind modern LLMs

Introduced in "Attention Is All You Need" (2017), the Transformer processes an entire sequence in parallel using self-attention, which lets every token directly weigh how relevant every other token is — removing the sequential bottleneck of RNNs and enabling massive-scale training.

Core Ideas

  • Self-Attention: Computes a weighted representation of each token based on all other tokens
  • Multi-Head Attention: Runs several attention operations in parallel to capture different relationships
  • Positional Encoding: Injects word-order information since attention itself is order-agnostic
  • Encoder-Decoder / Decoder-only: BERT-style encoders for understanding, GPT-style decoders for generation
python
from transformers import AutoTokenizer, AutoModel
python
tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased') model = AutoModel.from_pretrained('bert-base-uncased')
python
inputs = tokenizer("Transformers changed AI.", return_tensors='pt') outputs = model(**inputs)
Key Insight: Every major modern model — BERT, GPT, T5, Vision Transformers, Whisper — is a variant of the same Transformer building block applied to different data types.

23 Natural Language Processing

Teaching machines to understand and generate text

Text Representation Methods

MethodIdea
Bag-of-Words / TF-IDFSparse counts of word frequency and importance
Word2Vec / GloVeDense vectors where similar words are close together
Contextual Embeddings (BERT)Word meaning changes based on surrounding context

Common NLP Tasks

  • Text classification: Sentiment analysis, spam detection, topic labeling
  • Named Entity Recognition (NER): Extracting people, places, organizations
  • Machine translation: Converting text between languages
  • Question answering & summarization: Extractive or generative
python
from transformers import pipeline
python
classifier = pipeline('sentiment-analysis') result = classifier("This documentation is really helpful!") print(result)

24 Computer Vision

Teaching machines to interpret visual data

Common Vision Tasks

TaskOutputExample Models
Image ClassificationSingle label per imageResNet, EfficientNet, ViT
Object DetectionBounding boxes + labelsYOLO, Faster R-CNN
Semantic SegmentationPer-pixel class labelsU-Net, DeepLab
Image GenerationSynthesized imagesStable Diffusion, GANs
python
from tensorflow.keras.applications import ResNet50 from tensorflow.keras.applications.resnet50 import preprocess_input, decode_predictions
python
model = ResNet50(weights='imagenet') preds = model.predict(preprocess_input(image_batch)) print(decode_predictions(preds, top=3)[0])
Tip: Data augmentation — random flips, rotations, crops, and color jitter — is one of the highest-leverage ways to improve vision model generalization, especially with limited data.

25 Generative AI & Large Language Models

Models that create new content rather than just classify it

Major Approaches

FamilyIdeaExamples
LLMsDecoder-only Transformers predicting the next tokenGPT, Claude, Llama, Gemini
Diffusion ModelsLearn to reverse a gradual noising processStable Diffusion, DALL·E
GANsGenerator and discriminator compete adversariallyStyleGAN
VAEsLearn a compressed latent representation of dataImage compression, generation

Working with LLMs

  • Prompt engineering: Crafting inputs to elicit better model outputs
  • Fine-tuning: Further training a pretrained model on task-specific data
  • RAG (Retrieval-Augmented Generation): Grounding responses in retrieved documents to reduce hallucination
  • RLHF: Aligning model behavior using human feedback as a reward signal
Note: Most production LLM applications today are built by calling an existing model's API and combining it with retrieval, tools, and prompt design — rather than training a model from scratch.

26 Transfer Learning & Pretrained Models

Standing on the shoulders of models trained on massive datasets

Transfer learning reuses a model already trained on a large, general dataset and adapts it to a new, often smaller, task-specific dataset — dramatically reducing the data and compute needed.

python
from tensorflow.keras.applications import MobileNetV2
python
base_model = MobileNetV2(weights='imagenet', include_top=False, input_shape=(224,224,3)) base_model.trainable = False
python
model = keras.Sequential([ base_model, keras.layers.GlobalAveragePooling2D(), keras.layers.Dense(10, activation='softmax') ])

Two Common Strategies

  • Feature extraction: Freeze the pretrained base, train only new top layers
  • Fine-tuning: Unfreeze some/all base layers and continue training at a low learning rate

27 Reinforcement Learning

Learning by trial, error, and reward

An agent interacts with an environment, taking actions based on observed states, and receives rewards that guide it toward better behavior over time — without being told the correct action directly.

ConceptMeaning
PolicyThe agent's strategy mapping states to actions
Value functionExpected future reward from a state
Q-LearningLearns the value of state-action pairs
Policy Gradient / PPODirectly optimizes the policy using gradient ascent on expected reward
Key Insight: RL powers game-playing agents (AlphaGo), robotics control, and is also the technique behind RLHF, which aligns LLM outputs with human preferences.

28 Model Deployment & MLOps

Getting models from notebooks into production

Save and load a trained model.

python
import joblib joblib.dump(model, 'model.pkl') loaded = joblib.load('model.pkl')

Serving with FastAPI.

python
from fastapi import FastAPI app = FastAPI()
python
@app.post("/predict") def predict(features: list): return {"prediction": loaded.predict([features]).tolist()}

MLOps Concerns

  • Versioning: Tracking data, code, and model versions together (DVC, MLflow)
  • Monitoring: Watching for data drift and performance decay in production
  • CI/CD for ML: Automated retraining and testing pipelines
  • Scalability: Batching, model quantization, and serving infrastructure (Docker, Kubernetes, TensorFlow Serving, Triton)
Tip: A model's accuracy at deployment time is only the starting point — plan for monitoring and periodic retraining, since real-world data distributions shift over time.

29 Ethics & Responsible AI

Building AI systems responsibly

  • Bias & fairness: Models can inherit and amplify biases present in training data
  • Explainability: Tools like SHAP and LIME help interpret why a model made a decision
  • Privacy: Techniques like differential privacy and federated learning protect sensitive training data
  • Robustness & safety: Guarding against adversarial inputs and unsafe failure modes
  • Environmental cost: Large model training consumes significant compute and energy
Warning: A model that performs well on average can still fail badly for specific subgroups — always evaluate performance across relevant slices of your data, not just in aggregate.