Below: the dataset and its one known trap, the build steps, the metrics that match the business problem, common mistakes, and the scope for Diploma, B.Tech and M.Tech.
- Dataset: IBM’s Telco Customer Churn sample, with 7,043 customers and 1,869 who left (26.5%).
- The output is a decision, not a score: a ranked list of who to call, with a reason for each.
- Judge the model on the leavers: recall, precision, and how many leavers fall in the top tenth of the list.
- Start simple. Logistic regression is the baseline and the easiest model to explain in a viva.
- Set the threshold from costs: what an offer costs against what a lost customer costs.
Which dataset should you use for customer churn prediction?
The Telco Customer Churn dataset on Kaggle is an IBM sample. Each row is one customer, and the Churn column says whether they left.
| What | Detail | What to do with it |
|---|---|---|
| Customers | 7,043 rows, one per customer | Use cross-validation; the set is too small to trust one split |
| Customers who left | 1,869, which is 26.5% | Stratify every split and use class weights |
| Fields | Contract type, services taken, tenure, payment method, monthly and total charges | One-hot encode the categories; scale the numbers for linear models |
| Known trap | Total charges load as text, with blanks for brand-new customers | Convert to numbers and decide how to fill the blanks, then say so in the report |
How do you build a churn prediction model step by step?
- Clean the table. Drop the customer ID, fix the total-charges column and check every category for odd values.
- Explore before you model. Plot churn against contract type, tenure and monthly charges. These charts go straight into the report.
- Split, then encode. Hold out a stratified test set, and fit encoders and scalers on the training part only.
- Train a baseline. Logistic regression with class weights.
- Compare two more models. A random forest and a gradient-boosted model, tuned with cross-validation.
- Explain the scores. Use the logistic coefficients or SHAP values to show what drives churn for one customer.
- Turn scores into a call list. Sort customers by score, pick a cut-off from the retention budget and show the list in a Streamlit app.
Which metrics matter in a churn prediction project?
Predicting that nobody leaves is 73.5% accurate on this data and helps no one. Report what the retention team would care about:
- Recall on leavers: how many of the customers who left you would have flagged.
- Precision on leavers: how many of your calls would reach someone who was really leaving.
- Lift in the top tenth: how many more leavers your top 10% holds than a random 10% would.
- ROC and precision-recall curves, to compare models across thresholds.
What mistakes should you avoid in a churn project?
- Leaking the answer. A field that is only known after a customer leaves must not be a feature.
- Encoding before splitting. Fit encoders on training data, or test information slips in.
- Stopping at the model. Without a threshold and a list, the project has no output a company could act on.
- Calling correlation a cause. Month-to-month contracts go with churn; that does not prove the contract causes it.
- Hiding the baseline. Show how much each model adds over logistic regression.
How does a churn project change for Diploma, B.Tech and M.Tech?
| Level | Scope | What to show |
|---|---|---|
| Diploma | Cleaning, charts, and logistic regression against a decision tree | A dashboard of churn by contract and tenure, and a confusion matrix |
| B.Tech / B.E. | Three models compared, a cost-based threshold and SHAP explanations | A call-list app with a reason beside each customer |
| M.Tech / M.E. | Survival analysis for when a customer will leave, uplift modelling for who responds to an offer, or deep tabular models against boosted trees | A base paper reproduced, one measured extension and a paper in IEEE format |
Diploma students will find more small builds on our Diploma projects page. For M.Tech, see how an M.Tech project is built from a base paper.
What goes in the report, and which viva questions come up?
Lead the report with the business problem, then the data, the method, the results and what the company should do. For the black book and slides, see our report and PPT help. Expect these questions:
- Why did you not use accuracy to pick the model?
- How did you fill the blank total charges, and why?
- Which three features matter most, and does that make business sense?
- How did you choose the cut-off for the call list?
- How would you retrain the model as customers and prices change?
Related guides: credit card fraud detection, where the rare class is far rarer, and demand forecasting. The machine learning ideas on tabular data list more topics of this kind.
This guide is a plan to discuss with your guide, and the figures in it come from the linked dataset pages and papers, not from our own work. Our 8 delivered case studies are M.Tech and M.E. projects, each shown with its real paper pages, screens and results.
Frequently asked questions
Which dataset is best for a customer churn prediction project?
IBM’s Telco Customer Churn sample on Kaggle is the usual choice. It has 7,043 customers, 1,869 of whom left, with their contract, services, tenure and charges. It is small, clean and well known, so examiners can check your results against published ones.
Which algorithm is used for churn prediction?
Start with logistic regression, because its coefficients are easy to explain. Then compare a random forest and a gradient-boosted model such as XGBoost. On tables of this size the boosted model usually scores a little higher, and the logistic model is easier to defend.
Is churn prediction a good project for final year?
It suits a Diploma or B.Tech mini project as it stands. For a B.Tech major project, add a threshold chosen from retention costs, an explanation for each customer and a working call-list app. For M.Tech, add survival analysis or uplift modelling.
What accuracy should a churn model reach?
Do not aim at an accuracy figure. About 26.5% of the Telco customers left, so predicting that nobody leaves is already 73.5% accurate. Report recall and precision on the customers who left, and how many leavers appear in the top tenth of your list.
Can The Ultimate Project World help with a churn prediction project?
Yes. Tell us your level, your review dates and what your guide expects in the free consultation. We will say exactly which parts we can take on, such as the code, the app, the report, the PPT and viva preparation.
