About this course
Pandas, NumPy and Matplotlib applied to real datasets, ending with a portfolio analysis project.
This is a 6-week short course - enough depth to build real projects and a portfolio piece, without a long commitment. It runs online in the August 2026 cohort (starting 1 August 2026) and is taught the EchoLens way: you learn by doing real, gradeable work rather than just watching lectures.
What's included
- Live, instructor-led online sessions across 6 weeks (24 hours total).
- Hands-on coding quests you solve inside the EchoLens browser compiler - nothing to install.
- Gems, stages and a leaderboard that keep you moving instead of grade anxiety.
- A verified certificate with a scannable QR code, ready to share on LinkedIn, when you finish.
- The first week is open free so you can try the course before you pay.
What you will learn
Course learning outcomes
By the end of this course, you will be able to:
- CLO 1. Write Python programs using core control flow, data structures and NumPy arrays to process data.
- CLO 2. Load, clean, transform, group and visualise real datasets using pandas, Matplotlib and Seaborn.
- CLO 3. Perform exploratory data analysis and build a first scikit-learn model into a portfolio-ready analysis notebook.
Course outline - level by level
12 leveles, each with hands-on quests you clear in the portal.
- Level 1. Setup and Python basics · Control flow and functions - Set up Jupyter and Colab, then master variables, data types, operators, reading code, if/else, loops, functions and list comprehensions.
- Level 2. Data structures for analysis · Files, libraries, and environments - Work with lists, tuples, sets and dictionaries, choose the right structure, read and write files, import libraries, and set up virtual environments with pip.
- Level 3. NumPy arrays · NumPy operations - Create and shape arrays, index and slice, and use vectorised math, broadcasting and aggregations across axes - and see why arrays beat lists for data.
- Level 4. Meet pandas · Loading real data - Load real CSV and Excel data into Series and DataFrames, inspect rows, columns and the index, and read a messy dataset with head, info and describe.
- Level 5. Selecting and filtering · Cleaning data - Select the rows you need with loc, iloc and boolean masks, then handle missing values, duplicates and wrong types to build consistent, trustworthy tables.
- Level 6. Transforming columns · Grouping and joining - Transform columns with apply, map, string and date operations, then group, aggregate, pivot and merge across tables.
- Level 7. Matplotlib basics · Customising charts - Build line, bar and scatter charts with figures and axes, add titles, labels, legends, styles and subplots, and read a plot critically.
- Level 8. Seaborn for statistics · Telling a story with charts - Use Seaborn for distributions, categorical plots and heatmaps, choose the right visual, remove clutter, and land one clear message per chart.
- Level 9. Descriptive statistics · The EDA workflow - Summarise a dataset honestly with mean, median, spread and correlation, spot outliers, ask questions of the data, and turn it into findings.
- Level 10. A messy real dataset · Probability and distributions - Clean and explore a messy real dataset end to end, document decisions, and get a first feel for probability and common distributions.
- Level 11. First taste of scikit-learn · Preparing data for models - Meet the fit and predict pattern, split train and test sets, encode and scale features, and train a simple model end to end.
- Level 12. Communicating results · Bringing it together - Turn raw data into insight in clean, reproducible notebooks, and communicate findings a non-technical reader gets.
How you submit: Coding quests in the built-in compiler with an AI copilot beside the editor - exactly the Cursor/Copilot workflow the course teaches.
Your production end project
An End-to-End Data Analysis Report
A real dataset, cleaned, analysed, visualised, and explained - the portfolio piece every junior data analyst needs.
Take one real, messy dataset and carry it all the way from raw file to a clear, decision-ready analysis. You clean it, explore it, visualise it, and write up the findings in a notebook a hiring manager or a client could read and act on.
- ✓A raw dataset loaded and inspected, with its problems documented
- ✓Full cleaning: missing values, duplicates, and wrong types handled and explained
- ✓At least three exploratory questions answered with pandas
- ✓At least four clear, labelled charts that support the findings
- ✓Descriptive statistics and one correlation insight
- ✓A written summary of five key findings in plain English
- ✓A short recommendation for a decision maker
- ✓A clean, commented, reproducible notebook
Shipped when: The notebook runs top to bottom without errors, every chart supports a stated finding, and a non-technical reader understands the conclusion in two minutes.
Who it's for
Python for Data Science suits learners at a beginner to intermediate level who want a practical, project-based route into Python for Data Science. You need only a browser and an internet connection - all coding runs inside the EchoLens compiler, so there is nothing to set up.
Certificate
Finish every stage and EchoLens issues a verified certificate carrying a QR code anyone can scan to confirm it on our site. You can add it to your CV or share it to LinkedIn in one click.