SUNNY SOLUTION

Machine Learning Approaches for Loan Approval Prediction

By Anh Nguyen • 20/07/2026 22:17 • šŸ‘ 6 views

Share

Machine Learning Approaches for Loan Approval Prediction

Overview

This document explains how four machine learning algorithms work using the same loan approval dataset:

FeatureDescription
AgeApplicant age
IncomeAnnual income
ApprovedLoan approval result (1=Approved, 0=Rejected)

Example:

AgeIncomeApproved
2530,0000
3570,0001
4590,0001
2325,0000
4060,0001

Problem Statement

Given:

Age = 30
Income = 1,000,000

Predict:

Approved = ?

We will compare:

  1. Decision Tree
  2. Logistic Regression
  3. Random Forest
  4. CatBoost

1. Decision Tree

How It Works

Decision Tree learns a series of IF-THEN rules.

Example:

Income > 60,000?
        |
    +---+---+
    |       |
   No      Yes
    |       |
 Reject   Approve

The model starts at the root node and follows a path until reaching a leaf.


Training Process

Step 1

Find the best feature to split.

Example:

Age
Income

The algorithm determines:

Income > 60,000

creates the purest separation.


Step 2

Split records.

Income <= 60,000
      -> Mostly Rejected

Income > 60,000
      -> Mostly Approved

Step 3

Repeat recursively.

Income > 60,000 ?
          |
       Yes
          |
     Age > 30 ?
      /      \
   No         Yes
Approve     Approve

Example Prediction

Customer:

Age = 30
Income = 1,000,000

Path:

Income > 60,000
     ↓
Approve

Prediction:

Approved

Advantages

  • Easy to understand
  • Easy to visualize
  • Fast training

Disadvantages

  • Overfitting
  • Can ignore useful features
  • Small changes in data can change the tree

2. Logistic Regression

How It Works

Unlike Decision Tree, Logistic Regression does not create rules.

Instead it learns a mathematical formula:

Score =
b0
+ b1 Ɨ Age
+ b2 Ɨ Income

Then converts the score into a probability.

Probability =
1 / (1 + e^-Score)

Training Process

The model learns coefficients.

Example:

Score =
-5
+ 0.02 Ɨ Age
+ 0.00005 Ɨ Income

Example Prediction

Customer:

Age = 30
Income = 1,000,000

Formula:

Score

=
-5
+ (0.02 Ɨ 30)
+ (0.00005 Ɨ 1,000,000)

=
45.6

Probability:

ā‰ˆ 100%

Prediction:

Approved

Key Characteristic

Logistic Regression always considers:

Age
AND
Income

simultaneously.

It cannot completely ignore a feature like a Decision Tree often does.


Advantages

  • Very fast
  • Easy to explain
  • Works well on small datasets
  • Uses all features

Disadvantages

  • Struggles with complex nonlinear relationships
  • Requires feature scaling

3. Random Forest

How It Works

Random Forest is a collection of many Decision Trees.

Instead of:

1 tree

we build:

100 trees
200 trees
500 trees

Architecture

                Dataset
                    |
    ---------------------------------
    |       |       |       |       |
  Tree1   Tree2   Tree3   Tree4   Tree5
    |       |       |       |       |
    ---------------------------------
                    |
             Final Decision

Voting Mechanism

Suppose:

Tree 1 -> Approve
Tree 2 -> Approve
Tree 3 -> Reject
Tree 4 -> Approve
Tree 5 -> Approve

Final result:

Approve

because:

4 votes > 1 vote

Example Prediction

Customer:

Age = 30
Income = 1,000,000

Predictions:

Tree 1 -> Approve
Tree 2 -> Approve
Tree 3 -> Approve
Tree 4 -> Reject
Tree 5 -> Approve
...

Final:

Approved

Why Better Than One Tree

Single tree:

May learn wrong rule

Random Forest:

Many trees reduce mistakes

Advantages

  • Less overfitting
  • High accuracy
  • Works well on small datasets

Disadvantages

  • Harder to explain
  • Larger model size

4. CatBoost

How It Works

CatBoost is a Gradient Boosting algorithm.

Instead of many trees voting independently:

Tree1
Tree2
Tree3

each tree learns from previous mistakes.


Training Process

Tree 1

Initial prediction:

Customer A -> Reject

Actual:

Approve

Error detected.


Tree 2

Learns:

How can I fix Tree 1?

Tree 3

Learns:

How can I fix Tree 1 + Tree 2?

Final Model

Tree1
+
Tree2
+
Tree3
+
...
+
Tree200

Illustration

Tree 1
Accuracy = 70%

Tree 2
Fixes mistakes

Accuracy = 75%

Tree 3
Fixes remaining mistakes

Accuracy = 80%

...

Tree 200

Accuracy = 90%+

Example Prediction

Customer:

Age = 30
Income = 1,000,000

Tree contributions:

Tree 1 -> +0.30

Tree 2 -> +0.20

Tree 3 -> +0.10

...

Final Probability
= 99%

Prediction:

Approved

Why CatBoost Is Powerful

CatBoost:

  • Learns complex patterns
  • Reduces overfitting
  • Works very well on tabular data
  • Handles categorical features automatically

Advantages

  • Usually highest accuracy
  • Minimal feature engineering
  • Strong performance on small datasets

Disadvantages

  • Slower than Logistic Regression
  • More complex to understand

Comparison

Learning Strategy

AlgorithmStrategy
Decision TreeOne set of rules
Logistic RegressionOne mathematical formula
Random ForestMany trees vote
CatBoostTrees correct previous mistakes

Uses Age and Income Together?

AlgorithmUses Both Features?
Decision TreeNot always
Logistic RegressionAlways
Random ForestUsually
CatBoostUsually

Interpretability

AlgorithmInterpretability
Decision TreeExcellent
Logistic RegressionExcellent
Random ForestMedium
CatBoostLow

Expected Performance on ~1000 Records

AlgorithmTypical Performance
Decision TreeGood
Logistic RegressionGood
Random ForestVery Good
CatBoostExcellent

Summary

For the loan approval problem:

Decision Tree

Learns IF-THEN rules

Example:

Income > 60k -> Approve

Logistic Regression

Learns a probability formula

Example:

Age and Income both influence approval

Random Forest

Many Decision Trees vote

Example:

200 trees decide together

CatBoost

Many trees learn sequentially
and fix previous mistakes

Example:

Tree 2 improves Tree 1
Tree 3 improves Tree 1 + Tree 2
...

Recommendation for This Dataset (~1000 Rows)

A practical experimentation order:

1. Logistic Regression
2. Decision Tree
3. Random Forest
4. CatBoost

Then compare using 5-fold cross-validation and choose the model with the best balance between accuracy, interpretability, and maintainability.