Complete AI & ML Guide
From NumPy arrays to Transformers and Generative AI. This documentation covers modern AI frameworks, deep learning techniques, and practical ML workflows — with practical examples, deep explanations, and best practices used across research and industry.
1 Introduction to AI & Machine Learning
What AI and ML are, and how they relate
Artificial Intelligence (AI) is the broad field of building systems that perform tasks normally requiring human intelligence. Machine Learning (ML) is a subset of AI in which systems learn patterns from data instead of following explicitly programmed rules. Deep Learning is, in turn, a subset of ML built on multi-layered neural networks capable of learning very complex patterns directly from raw data.
The AI Hierarchy
| Layer | Description | Example |
|---|---|---|
| Artificial Intelligence | Any technique that lets machines mimic intelligent behavior | Rule-based chess engines, expert systems |
| Machine Learning | Systems that learn from data rather than fixed rules | Spam filters, recommendation engines |
| Deep Learning | ML using multi-layer neural networks | Image recognition, language models |
Why Python Dominates AI/ML
- Rich ecosystem: NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch all mature and interoperable
- Readable syntax: Lets researchers focus on ideas instead of boilerplate
- Community & research adoption: Most papers ship Python reference implementations
- Interoperability: Easy to bind to fast C/C++/CUDA backends under the hood
A minimal "hello world" of machine learning.
2 Types of Machine Learning
The three core learning paradigms
| Type | Data | Goal | Examples |
|---|---|---|---|
| Supervised Learning | Labeled (input, correct output) pairs | Predict labels for new inputs | Spam detection, price prediction |
| Unsupervised Learning | Unlabeled data | Discover hidden structure | Customer segmentation, anomaly detection |
| Reinforcement Learning | Environment + reward signal | Learn a policy maximizing reward | Game-playing agents, robotics |
| Semi-supervised | Small labeled + large unlabeled set | Combine both signals | Medical imaging with scarce labels |
| Self-supervised | Unlabeled data with generated pseudo-labels | Learn general representations | Pretraining LLMs, contrastive vision models |
Supervised Learning: Two Main Tasks
- Regression: Predict a continuous number (e.g., house price)
- Classification: Predict a discrete category (e.g., spam or not spam)
3 Math Foundations
The minimum math you need to understand what's happening
Linear Algebra
Data in ML is represented as vectors and matrices. A neural network layer is fundamentally a matrix multiplication followed by a non-linear function.
Vector and matrix.
Dot product (core operation in neural network layers).
Matrix transpose and inverse.
Calculus (Gradients)
Training a model means minimizing a loss function. Gradient descent uses the derivative of the loss with respect to each parameter to know which direction reduces error, and steps the parameters that way, repeatedly.
Probability & Statistics
- Distributions: Understanding how data is spread (normal, uniform, etc.)
- Mean, variance, standard deviation: Core descriptive statistics used everywhere in preprocessing
- Bayes' theorem: Foundation of probabilistic models and Naive Bayes classifiers
- Conditional probability: Central to language models predicting the next token
4 Environment Setup
Getting your AI/ML development environment ready
Create an isolated virtual environment.
Linux/Mac.
Windows.
Or with conda (popular in the ML community)
Install the core stack.
Deep learning frameworks.
Jupyter for interactive notebooks.
Checking GPU Availability
5 NumPy Fundamentals
The array library everything else in AI/ML is built on
NumPy provides the ndarray, a fast, memory-efficient, multi-dimensional array type with vectorized operations that avoid slow Python loops. Every major ML/DL framework either uses NumPy directly or mirrors its API.
Creating arrays.
Shape, dtype, reshape.
Vectorized operations (no explicit loops).
Broadcasting: operate on arrays of different shapes.
Indexing and slicing.
Common Operations
| Function | Purpose |
|---|---|
np.mean / np.std | Descriptive statistics along an axis |
np.dot / @ | Matrix multiplication |
np.concatenate / np.stack | Combine arrays |
np.where | Conditional element selection |
np.linalg.norm | Vector/matrix norms |
for loops — they run in optimized C code and can be 10-100x faster on large datasets.
6 Pandas for Data Manipulation
Loading, cleaning, and exploring tabular data
Pandas introduces the DataFrame, a labeled 2D table structure ideal for real-world, messy datasets — CSVs, SQL query results, spreadsheets, and logs.
Loading data.
Quick exploration.
Selecting columns and rows.
Handling missing values.
Grouping and aggregation.
Merging datasets.
7 Data Visualization
Seeing your data before modeling it
Visualization surfaces distributions, outliers, correlations, and class imbalance that summary statistics alone can hide. Matplotlib is the low-level foundation; Seaborn builds statistical plots on top of it with far less code.
Distribution of a single feature.
Relationship between two features.
Correlation heatmap.
Box plot to spot outliers.
Common Plot Types
| Plot | Best For |
|---|---|
| Histogram / KDE | Feature distributions |
| Scatter plot | Relationship between two variables |
| Box / violin plot | Outliers and spread by category |
| Correlation heatmap | Feature interdependence |
| Confusion matrix | Classification error patterns |
8 Data Preprocessing & Feature Engineering
Turning raw data into model-ready features
Scaling & Encoding
Feature scaling (mean 0, std 1) — required for many algorithms.
One-hot encoding categorical variables.
Train/test split.
Common Preprocessing Steps
- Handling missing data: Imputation (mean/median/mode) or removal
- Feature scaling: Standardization or min-max normalization
- Encoding categories: One-hot, ordinal, or target encoding
- Feature engineering: Creating new informative features (ratios, dates split into day/month, text length)
- Outlier handling: Clipping, transformation, or removal
- Class imbalance: Oversampling (SMOTE), undersampling, or class weights
9 Scikit-learn Basics
The standard toolkit for classical machine learning
Scikit-learn provides a consistent API — fit, predict, transform — across dozens of algorithms, plus utilities for preprocessing, model selection, and evaluation.
Pipelines chain preprocessing + model into one object.
Pipeline prevents data leakage automatically and makes cross-validation and deployment far simpler.
10 Regression Algorithms
Predicting continuous values
| Algorithm | Idea | When to Use |
|---|---|---|
| Linear Regression | Fits a straight line/hyperplane minimizing squared error | Simple, interpretable baselines |
| Ridge / Lasso | Linear regression with L2/L1 penalty on coefficients | Many features, need regularization |
| Polynomial Regression | Linear regression on polynomial feature expansions | Non-linear but smooth relationships |
| Decision Tree Regressor | Splits data into regions, predicts region average | Non-linear, interpretable |
| Random Forest / Gradient Boosting | Ensembles of trees | Strong tabular baselines |
| SVR | Support vector machine adapted for regression | Small-to-medium, high-dimensional data |
11 Classification Algorithms
Predicting discrete categories
| Algorithm | Idea | When to Use |
|---|---|---|
| Logistic Regression | Linear decision boundary via sigmoid function | Fast, interpretable baseline |
| k-Nearest Neighbors | Classifies by majority vote of nearest points | Small datasets, simple boundaries |
| Decision Tree | Recursive feature-based splits | Interpretable, non-linear |
| Random Forest | Ensemble of decision trees, majority vote | Strong general-purpose baseline |
| Gradient Boosting (XGBoost/LightGBM) | Sequentially corrects previous trees' errors | Top performer on tabular data |
| SVM | Maximizes margin between classes | High-dimensional, clear-margin data |
| Naive Bayes | Probabilistic, assumes feature independence | Text classification, spam filtering |
12 Clustering & Dimensionality Reduction
Finding structure without labels
Clustering
| Algorithm | Idea |
|---|---|
| K-Means | Partitions data into k clusters by minimizing distance to centroids |
| Hierarchical Clustering | Builds a tree of nested clusters (dendrogram) |
| DBSCAN | Density-based; finds arbitrarily shaped clusters and outliers |
| Gaussian Mixture Models | Soft, probabilistic clustering |
Dimensionality Reduction
PCA: project high-dimensional data onto principal components.
K-Means clustering.
13 Ensemble Methods
Combining multiple models for better performance
- Bagging: Train many models on random data subsets in parallel and average results (e.g., Random Forest)
- Boosting: Train models sequentially, each correcting the previous one's errors (e.g., XGBoost, LightGBM, CatBoost, AdaBoost)
- Stacking: Train a meta-model on the outputs of several base models
- Voting: Combine predictions from different model types by majority vote or averaging
14 Model Evaluation Metrics
Measuring what actually matters for your problem
Classification Metrics
| Metric | Formula / Idea | Best For |
|---|---|---|
| Accuracy | Correct / Total | Balanced classes |
| Precision | TP / (TP + FP) | Cost of false positives is high |
| Recall | TP / (TP + FN) | Cost of false negatives is high |
| F1-score | Harmonic mean of precision/recall | Imbalanced classes |
| ROC-AUC | Area under TPR vs FPR curve | Ranking quality across thresholds |
Regression Metrics
| Metric | Idea |
|---|---|
| MAE | Average absolute error, robust to outliers |
| MSE / RMSE | Penalizes larger errors more heavily |
| R² | Proportion of variance explained by the model |
15 Overfitting, Underfitting & Regularization
Building models that generalize
| Problem | Symptom | Fix |
|---|---|---|
| Underfitting | Poor performance on both train and test data | More complex model, more/better features, less regularization |
| Overfitting | Great on training data, poor on test data | More data, regularization, simpler model, early stopping |
Regularization Techniques
- L1 (Lasso): Pushes some weights to exactly zero — built-in feature selection
- L2 (Ridge): Shrinks weights smoothly, discourages large coefficients
- Dropout: Randomly disables neurons during training (deep learning)
- Early stopping: Stop training once validation loss stops improving
- Data augmentation: Artificially expand training data (flips, crops, noise)
16 Cross-Validation & Hyperparameter Tuning
Getting a reliable estimate of real-world performance
K-Fold cross-validation.
Grid search over hyperparameters.
RandomizedSearchCV or Bayesian optimization (Optuna) instead of full grid search when the hyperparameter space is large — it finds near-optimal settings far faster.
17 Neural Network Fundamentals
The building blocks of deep learning
A neural network is composed of layers of neurons, each computing a weighted sum of its inputs followed by a non-linear activation function. Stacking layers lets the network approximate arbitrarily complex functions.
Key Components
| Component | Role |
|---|---|
| Weights & Biases | Learnable parameters adjusted during training |
| Activation Function | Introduces non-linearity (ReLU, Sigmoid, Tanh, Softmax) |
| Loss Function | Measures prediction error (Cross-Entropy, MSE) |
| Optimizer | Updates weights to reduce loss (SGD, Adam, RMSprop) |
| Backpropagation | Computes gradients of the loss w.r.t. every weight via the chain rule |
18 Deep Learning with TensorFlow & Keras
Google's production-grade deep learning framework
Keras, now TensorFlow's high-level API, makes building and training neural networks concise and readable while TensorFlow handles graph optimization, GPU/TPU execution, and deployment.
keras.callbacks.EarlyStopping and ModelCheckpoint to automatically stop training and save the best-performing weights once validation loss stops improving.
19 Deep Learning with PyTorch
The dominant framework in AI research
PyTorch favors an imperative, "define-by-run" style that feels like regular Python, making it especially popular in research where models change often.
20 Convolutional Neural Networks (CNNs)
The architecture behind modern computer vision
Convolutional layers slide small learnable filters across an image to detect local patterns like edges, textures, and shapes — with far fewer parameters than a fully-connected layer would need.
Core Layers
| Layer | Role |
|---|---|
| Convolution | Extracts local spatial features using learnable filters |
| Pooling (Max/Avg) | Downsamples feature maps, adds translation invariance |
| Flatten | Converts 2D feature maps into a 1D vector |
| Fully Connected | Combines features for final classification |
21 Recurrent Networks & Sequence Models
Modeling sequences: text, time series, audio
RNNs process sequences step-by-step, carrying a hidden state forward so earlier inputs can influence later predictions. LSTM and GRU cells add gating mechanisms that solve the vanishing-gradient problem of plain RNNs, letting them retain information over longer sequences.
| Architecture | Best For |
|---|---|
| Simple RNN | Short sequences, educational baselines |
| LSTM | Longer sequences, time series, language modeling |
| GRU | Similar to LSTM, fewer parameters, faster to train |
| Seq2Seq (Encoder-Decoder) | Translation, summarization |
A simple LSTM-based classifier: an embedding layer, an LSTM layer, and a dense output layer.
22 Transformers & Attention Mechanism
The architecture behind modern LLMs
Introduced in "Attention Is All You Need" (2017), the Transformer processes an entire sequence in parallel using self-attention, which lets every token directly weigh how relevant every other token is — removing the sequential bottleneck of RNNs and enabling massive-scale training.
Core Ideas
- Self-Attention: Computes a weighted representation of each token based on all other tokens
- Multi-Head Attention: Runs several attention operations in parallel to capture different relationships
- Positional Encoding: Injects word-order information since attention itself is order-agnostic
- Encoder-Decoder / Decoder-only: BERT-style encoders for understanding, GPT-style decoders for generation
23 Natural Language Processing
Teaching machines to understand and generate text
Text Representation Methods
| Method | Idea |
|---|---|
| Bag-of-Words / TF-IDF | Sparse counts of word frequency and importance |
| Word2Vec / GloVe | Dense vectors where similar words are close together |
| Contextual Embeddings (BERT) | Word meaning changes based on surrounding context |
Common NLP Tasks
- Text classification: Sentiment analysis, spam detection, topic labeling
- Named Entity Recognition (NER): Extracting people, places, organizations
- Machine translation: Converting text between languages
- Question answering & summarization: Extractive or generative
24 Computer Vision
Teaching machines to interpret visual data
Common Vision Tasks
| Task | Output | Example Models |
|---|---|---|
| Image Classification | Single label per image | ResNet, EfficientNet, ViT |
| Object Detection | Bounding boxes + labels | YOLO, Faster R-CNN |
| Semantic Segmentation | Per-pixel class labels | U-Net, DeepLab |
| Image Generation | Synthesized images | Stable Diffusion, GANs |
25 Generative AI & Large Language Models
Models that create new content rather than just classify it
Major Approaches
| Family | Idea | Examples |
|---|---|---|
| LLMs | Decoder-only Transformers predicting the next token | GPT, Claude, Llama, Gemini |
| Diffusion Models | Learn to reverse a gradual noising process | Stable Diffusion, DALL·E |
| GANs | Generator and discriminator compete adversarially | StyleGAN |
| VAEs | Learn a compressed latent representation of data | Image compression, generation |
Working with LLMs
- Prompt engineering: Crafting inputs to elicit better model outputs
- Fine-tuning: Further training a pretrained model on task-specific data
- RAG (Retrieval-Augmented Generation): Grounding responses in retrieved documents to reduce hallucination
- RLHF: Aligning model behavior using human feedback as a reward signal
26 Transfer Learning & Pretrained Models
Standing on the shoulders of models trained on massive datasets
Transfer learning reuses a model already trained on a large, general dataset and adapts it to a new, often smaller, task-specific dataset — dramatically reducing the data and compute needed.
Two Common Strategies
- Feature extraction: Freeze the pretrained base, train only new top layers
- Fine-tuning: Unfreeze some/all base layers and continue training at a low learning rate
27 Reinforcement Learning
Learning by trial, error, and reward
An agent interacts with an environment, taking actions based on observed states, and receives rewards that guide it toward better behavior over time — without being told the correct action directly.
| Concept | Meaning |
|---|---|
| Policy | The agent's strategy mapping states to actions |
| Value function | Expected future reward from a state |
| Q-Learning | Learns the value of state-action pairs |
| Policy Gradient / PPO | Directly optimizes the policy using gradient ascent on expected reward |
28 Model Deployment & MLOps
Getting models from notebooks into production
Save and load a trained model.
Serving with FastAPI.
MLOps Concerns
- Versioning: Tracking data, code, and model versions together (DVC, MLflow)
- Monitoring: Watching for data drift and performance decay in production
- CI/CD for ML: Automated retraining and testing pipelines
- Scalability: Batching, model quantization, and serving infrastructure (Docker, Kubernetes, TensorFlow Serving, Triton)
29 Ethics & Responsible AI
Building AI systems responsibly
- Bias & fairness: Models can inherit and amplify biases present in training data
- Explainability: Tools like SHAP and LIME help interpret why a model made a decision
- Privacy: Techniques like differential privacy and federated learning protect sensitive training data
- Robustness & safety: Guarding against adversarial inputs and unsafe failure modes
- Environmental cost: Large model training consumes significant compute and energy