Anomaly Detection
Our project created an intelligent tool that sifts through data to find anything out of the ordinary, enabling banks to swiftly detect and investigate anomalies for improved decision-making and risk management.
MLAutomata
Contents
  • Introduction
  • Objective
  • Solution
  • Data Analysis
  • Model
  • Explainability
  • Conclusion
For Anomaly Detection

Introduction

Often the challenge associated with tasks like fraud and spam detection is the lack of all likely patterns needed to train suitable supervised learning models. This problem occurs when the fraudulent patterns are not only rare, they also change over time. Change in fraudulent pattern is because fraudsters continue to innovate new ways to circumvent measures put in place to prevent fraud. Limited data and continuously changing patterns makes learning significantly difficult.

Objective

The objective of the Anomaly Detection Model is to integrate machine learning methods into the existing fraud detection system to enhance its accuracy and efficiency. By leveraging advanced algorithms, the model aims to autonomously identify patterns indicative of fraudulent credit card transactions, thereby reducing false positives and improving overall fraud detection performance. Operating independently from the traditional rules engine, the model will classify transactions into genuine or suspicious categories, facilitating faster and more accurate decision-making in fraud prevention efforts.

Solution

Our comprehensive solution to detect and prevent fraud in the banking sector encompasses a meticulous approach integrating exploratory data analysis, statistical examination, and advanced machine learning techniques. Leveraging our proprietary ML Automata platform, we have developed, fitted, and rigorously evaluated multiple models, including Isolation Forest, Local Outlier Factor (LOF), and One-Class Support Vector Machine (One-Class SVM). Following rigorous technical and business validation, the selected model has been seamlessly deployed into production. This solution endeavors to significantly reduce false positives while enhancing the accuracy of fraudulent transaction detection, thereby fortifying the bank's security infrastructure and safeguarding against financial risks.
Technologies Used :
MLAutomata

Data Analysis

The dataset Provided is comprised of 46,407 transactions, our analysis reveals that the vast majority, specifically 45,867 transactions, correspond to non-fraudulent transactions, whereas only a small subset, totalling 540 transactions, are identified as fraudulent transactions. The dataset encompasses diverse customer and transactional attributes, including account balance, transaction location, type, and amount. Prior to importing the data into ML Automata, pre-processing steps were undertaken based on a combination of exploratory data analysis (EDA) and expert guidance.
image
The categorical columns are all one-hot encoded and numerical columns are normalized. The median income for a person is around 10k but many of the income levels were below the 2*Inter Quartile Range, many of the transaction amount withdrawn were more than the 2*IQR.
image
More than half the transactions were made by individual customers and 1/3rd of the customers is Business. The correlation among all the features were fairly low.
image

Model

We fitted several different unsupervised models including Isolation Forest, Local Outlier Factor (LOF), and One-Class Support Vector Machine (One-Class SVM).
image
The Isolation forest with the contamination rate of 0.02 gave us the better model than others with the testing F1 score of 72.6. most of the anomalies are clustered together. This can be observed by 3d TSNE plots
image

Explainability

We considered predicted anomalies from the Isolation Forest as labels and trained a random forest model to get the Explainable plots. The random forest model showed very high accuracy in predicting the anomalies, we calculated SHAP values to interpret the results from the random forest model.
image
By looking at the summary plots we can observe that transaction_count, fired_rule and income are the top 3 features which helps us in determining if a transaction is fraudulent or not. The top two features are directly proportional which means as transaction count increases and fired rule confidence increases the chances of a fraud is also increases but income is inversely proportional which means lower income level increases the chance of a fraudulent transactions.

Conclusion

The Isolation Forest performed well compared to other models and we were able to significantly reduce and detect the fraudulent transactions. This model will help us in reducing the manual power in detecting the fraud and we will be able to automate the process of detecting frauds. Training a random forest model on top of anomaly model helped us in understanding the model results and how are each feature are related with fraudulent transactions, based on this we can derive some hypothesis and implement it in our rules engine.
Contact Us