> For the complete documentation index, see [llms.txt](https://opencampus.gitbook.io/opencampus-machine-learning-program/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://opencampus.gitbook.io/opencampus-machine-learning-program/courses/archive/intermediate-machine-learning-legacy/week-3.md).

# Week 3 - Intro Kaggle competition - EDA and baseline models with PyTorch

Learning and testing - a.k.a. don't do Bullshit Machine Learning

## Course session

**Kaggle**

* Introduction
* Titanic
* Paddy

{% embed url="<https://www.kaggle.com/competitions/paddy-disease-classification/overview>" %}

* Exploratory Data Analysis(EDA) for Paddy Disease Classification

{% embed url="<https://www.kaggle.com/code/henrikho/opencampus-paddy-eda>" %}

**Solutions exercise MLP**

Presentation from the participants of the MLP from Coursera

**Walk-through**

PyTorchLightning

PyTorch 303 (Lab 03)

{% embed url="<https://colab.research.google.com/drive/1v1Ts2kC91coyZHJYicjBw7EwS5C9lD9q?usp=sharing>" %}

## **To-do**

😊

Go for your own through the Colab Notebook above (PyTorch303) and try to understand and repeat the steps for your own.

Do Week 3 of the Coursera Course

{% embed url="<https://www.coursera.org/learn/machine-learning-duke>" %}

Please register at kaggle.com and join the competition. Go through the Exploratory Data Analysis Notebook session and then train a Logistic regression as baseline model!

{% embed url="<https://www.kaggle.com/competitions/paddy-disease-classification/overview>" %}

The main objective of this Kaggle competition is to develop a machine or deep learning-based model to classify the given paddy leaf images accurately. A training dataset of 10,407 (75%) labeled images across ten classes (nine disease categories and normal leaf) is provided. Moreover, the competition host also provides additional metadata for each image, such as the paddy variety and age. Your task is to classify each paddy image in the given test dataset of 3,469 (25%) images into one of the nine disease categories or a normal leaf.

So that is where we will be heading in the next session trying different tools and techniques.

EDA Notebook

{% embed url="<https://www.kaggle.com/code/henrikho/opencampus-paddy-eda>" %}

Logistic regression (try first on your own but if your stuck look at the notebook below):

{% embed url="<https://www.kaggle.com/henrikho/opencampus-paddy-pytorch-logistic-regression>" %}

😊😊

Build an MLP in PyTorchLightning for Paddy Challenge on Kaggle

😊😊😊

Do your own EDA on the Paddy Challenge and/or look at other EDA notebooks from competitors. Make a final presentable EDA notebook

Transfer the CNN from the Coursera assignment to our Kaggle competition

Familiarize yourself with this PyTorch Tutorials:

{% embed url="<https://uvadlc-notebooks.readthedocs.io/en/latest/tutorial_notebooks/tutorial5/Inception_ResNet_DenseNet.html>" %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://opencampus.gitbook.io/opencampus-machine-learning-program/courses/archive/intermediate-machine-learning-legacy/week-3.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
