Get in Touch

Course Outline

Introduction

This section offers a foundational overview of when to apply 'machine learning,' key considerations, and the underlying concepts, including advantages and limitations. It explores datatypes (structured/unstructured/static/streamed), data validity and volume, the distinction between data-driven and user-driven analytics, statistical models versus machine learning models, the challenges of unsupervised learning, the bias-variance trade-off, iteration and evaluation, cross-validation approaches, and the paradigms of supervised, unsupervised, and reinforcement learning.

MAJOR TOPICS

1. Understanding Naive Bayes

  • Basic concepts of Bayesian methods
  • Probability
  • Joint probability
  • Conditional probability via Bayes' theorem
  • The Naive Bayes algorithm
  • Naive Bayes classification
  • The Laplace estimator
  • Incorporating numeric features into Naive Bayes

2. Understanding Decision Trees

  • The divide-and-conquer strategy
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Understanding Neural Networks

  • Transitioning from biological to artificial neurons
  • Activation functions
  • Network topology
  • Determining the number of layers
  • Direction of information flow
  • Number of nodes per layer
  • Training neural networks using backpropagation
  • Deep Learning

4. Understanding Support Vector Machines

  • Classification using hyperplanes
  • Maximizing the margin
  • Scenarios with linearly separable data
  • Scenarios with non-linearly separable data
  • Applying kernels for non-linear spaces

5. Understanding Clustering

  • Clustering as a machine learning task
  • The k-means clustering algorithm
  • Assigning and updating clusters based on distance
  • Selecting the optimal number of clusters

6. Measuring Performance for Classification

  • Handling classification prediction data
  • An in-depth look at confusion matrices
  • Evaluating performance using confusion matrices
  • Metric options beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Tuning Standard Models for Enhanced Performance

  • Leveraging caret for automated parameter tuning
  • Constructing a simple tuned model
  • Customizing the tuning workflow
  • Boosting model performance via meta-learning
  • Concepts of ensembles
  • Bagging
  • Boosting
  • Random forests
  • Training random forests
  • Evaluating random forest performance

MINOR TOPICS

8. Understanding Classification using Nearest Neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting an appropriate k
  • Data preparation for kNN
  • Why the kNN algorithm is considered lazy

9. Understanding Classification Rules

  • The separate-and-conquer approach
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Understanding Regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlations
  • Multiple linear regression

11. Understanding Regression Trees and Model Trees

  • Incorporating regression into trees

12. Understanding Association Rules

  • The Apriori algorithm for association rule learning
  • Measuring rule significance via support and confidence
  • Constructing rule sets using the Apriori principle

Extras

  • Spark/PySpark/MLlib and Multi-armed bandits

Requirements

Python Knowledge

 21 Hours

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories