AI NIFTY 50 Stock Scanner with Python

 

AI NIFTY 50 Stock Scanner with Python – Complete Beginner's Guide

Build Your Own AI Stock Scanner in Google Colab

Imagine opening your computer every morning and asking:

Which NIFTY 50 stocks currently have the strongest technical and AI-based setup?

Instead of manually checking 50 stocks one by one, Python can do much of this work automatically.

The AI NIFTY 50 Stock Scanner presented in this tutorial downloads historical stock-market data, calculates technical indicators, trains machine-learning models, predicts the probability of a positive 5-day return, applies trend and volume filters, calculates entry/stop-loss/target levels, ranks stocks and finally exports the results to Excel.

The complete program is designed to answer a question such as:

"Which NIFTY 50 stocks currently have the strongest combination of AI probability, expected return, trend, momentum and volume?"

This article explains the uploaded Python program from the ground up so that even someone with very little Python knowledge can understand and use it.


1. What Are We Building?

The program is an:

AI-Based NIFTY 50 Stock Scanner

It performs approximately this workflow:

NIFTY 50 Stocks
       ↓
Download Historical Data
       ↓
Calculate Technical Indicators
       ↓
Create Machine Learning Features
       ↓
Create 5-Day Prediction Target
       ↓
Train AI Models
       ↓
Test Models
       ↓
Predict Today's Probability
       ↓
Estimate 5-Day Return
       ↓
Apply Trend + Volume Filters
       ↓
Calculate Stop Loss & Target
       ↓
Calculate AI Score
       ↓
Rank Stocks
       ↓
Generate BUY / STRONG BUY / SELL / NO TRADE
       ↓
Export Results to Excel

This is considerably more advanced than a simple RSI or moving-average scanner.


2. Who Can Use This Program?

You don't need to be an experienced programmer.

This project is suitable for someone who knows:

  • Basic stock-market terminology

  • What a stock price is

  • What BUY and SELL mean

  • Basic computer usage

Python knowledge is helpful, but you can initially use the notebook by simply running the cells.

Later, you can learn Python by modifying individual sections.


3. What Is Google Colab?

The program is designed to run in Google Colab.

Google Colab is an online Python environment.

Normally, if you want to run Python programs, you might need to install:

  • Python

  • Pandas

  • NumPy

  • XGBoost

  • Scikit-learn

  • Matplotlib

  • other packages

With Colab, much of this setup can be done directly inside the notebook.

Think of Colab as:

A Python computer running inside your web browser.


4. Understanding the Notebook

The program is divided into 24 code cells.

Don't be frightened by the number.

You don't need to understand all 24 cells before running the program.

The easiest approach is:

Run
 ↓
Observe
 ↓
Understand
 ↓
Modify
 ↓
Experiment

Always start from the first cell and work downward.


5. Cell 1 – Installing Python Libraries

The first cell contains:

!pip -q install yfinance xgboost ta scikit-learn pandas numpy matplotlib seaborn

This installs the libraries required by the program.

Let's understand them.

yfinance

Used to download market data.

XGBoost

Used for the machine-learning models.

ta

Provides technical-analysis indicators.

scikit-learn

Provides machine-learning evaluation tools.

pandas

Used for tables and time-series data.

numpy

Used for numerical calculations.

matplotlib

Used for charts.

seaborn

Used for visualization.

The important point for a beginner is:

Don't manually install these one by one. Run Cell 1 first.


6. Cell 2 – Importing the Libraries

The next cell contains imports such as:

import numpy as np
import pandas as pd
import yfinance as yf
import matplotlib.pyplot as plt
import seaborn as sns

Importing means:

"Make these tools available to my Python program."

For example:

import pandas as pd

allows the program to use:

pd.DataFrame()

7. What Is Pandas?

Pandas is extremely important for this project.

Think of Pandas as:

Excel inside Python.

A stock-data table might look like:

DateOpenHighLowCloseVolume
1 Jan1,0001,0209901,015500,000
2 Jan1,0151,0401,0051,035650,000
3 Jan1,0351,0501,0251,045720,000

Pandas stores this kind of information in something called a DataFrame.


8. Cell 3 – Defining the NIFTY 50 Stocks

The program creates a list:

NIFTY50 = [
    'ADANIENT.NS',
    'ADANIPORTS.NS',
    'APOLLOHOSP.NS',
    ...
]

These are Yahoo Finance ticker symbols for Indian stocks.

The .NS suffix generally identifies NSE-listed instruments in the Yahoo Finance symbol format.

For example:

TCS.NS
INFY.NS
RELIANCE.NS
ONGC.NS

The program then loops through these symbols.


9. Important Issue in the Stock List

There is a small coding issue in the uploaded notebook.

This part appears as:

'POWERGRID.NS''RELIANCE.NS'

There is no comma between the two strings.

Python automatically joins adjacent strings, so this can effectively become one incorrect ticker:

POWERGRID.NSRELIANCE.NS

Correct version

It should be:

'POWERGRID.NS',
'RELIANCE.NS',

Why is this important?

Because the program may fail to download data correctly for those entries.

I strongly recommend correcting this before using the scanner.


10. Cell 4 – Downloading Historical Data

The program sets:

START_DATE = "2015-01-01"

This means the program attempts to obtain historical data starting from January 1, 2015.

It downloads data for:

NIFTY 50 stocks
+
NIFTY index
+
India VIX

The program uses:

yf.download()

to obtain the market data.


11. What Data Is Downloaded?

For every stock, the program keeps:

Open
High
Low
Close
Volume

These are known as OHLCV.

O – Open

Opening price.

H – High

Highest price during the session.

L – Low

Lowest price during the session.

C – Close

Closing price.

V – Volume

Number of shares traded.


12. Why Is Historical Data Necessary?

The AI cannot learn from today's price alone.

It needs historical examples.

Imagine giving the AI 10 years of information:

2015
2016
2017
...
2025
2026

It can study relationships between:

  • Price momentum

  • RSI

  • MACD

  • Trend

  • Volume

  • Volatility

  • NIFTY market behaviour

  • India VIX

and subsequent stock performance.


13. What Does data = {} Mean?

The program creates:

data = {}

A Python dictionary is like a labelled storage cabinet.

For example:

data["TCS.NS"]
data["INFY.NS"]
data["RELIANCE.NS"]

Each label contains the corresponding stock's historical DataFrame.


14. Cell 5 – Creating Technical Features

This is one of the most important sections of the entire program.

The function:

create_features()

takes raw stock prices and creates additional information.

These additional variables are called:

Features

A feature is simply information that the machine-learning model can use.


15. Return Features

The program calculates:

RET_1
RET_3
RET_5
RET_10
RET_20

These represent price changes over different periods.

For example:

df["RET_5"] = df["Close"].pct_change(5)

means:

Calculate the percentage change in closing price over approximately five trading observations.

This tells the AI whether the stock has recently been moving upward or downward.


16. Moving Averages

The program calculates:

SMA20
SMA50
SMA100
SMA200

SMA means:

Simple Moving Average

For example, SMA50 represents the average closing price over the previous 50 observations.

If:

Price > SMA50

the stock is trading above its 50-period average.

If:

Price < SMA50

the stock is trading below it.


17. Why Are Multiple Moving Averages Used?

The program doesn't rely on just one moving average.

It uses:

20-day
50-day
100-day
200-day

This gives the model information about both short-term and long-term trends.

For example:

Price > SMA20
Price > SMA50
Price > SMA200

suggests a stronger upward trend than if the price were below all three.


18. Distance from Moving Averages

The program also calculates:

DIST_SMA20
DIST_SMA50
DIST_SMA200

For example:

Close / SMA50 - 1

Suppose:

Close = ₹1,100
SMA50 = ₹1,000

Then:

1,100 / 1,000 - 1
= 0.10
= 10%

So the stock is:

10% above its 50-day moving average.


19. RSI – Relative Strength Index

The program calculates:

RSI

using a 14-period RSI.

RSI is a momentum indicator.

It helps answer:

How strong has recent price movement been?

The program later uses:

RSI > 50

as part of its trend filter for strong setups.


20. MACD

The program calculates:

MACD
MACD_SIGNAL
MACD_HIST

MACD is another momentum/trend indicator.

Instead of using only the closing price, the AI gets information about the relationship between different moving averages.

The histogram provides additional information about the difference between MACD and its signal line.


21. ADX

The program calculates:

ADX
DI_PLUS
DI_MINUS

ADX is commonly used to measure trend strength.

In this scanner, the strong-buy filter requires:

ADX > 18

So the program wants a reasonably established trend rather than a completely directionless market.


22. Bollinger Bands

The program calculates:

BB_HIGH
BB_LOW
BB_MID
BB_WIDTH
BB_POSITION

Bollinger Bands provide information about:

  • Price location

  • Volatility

  • Expansion/contraction of price movement

The program doesn't simply ask:

"Is price above the upper band?"

Instead, it converts Bollinger information into numerical features that the AI can use.


23. ATR – Average True Range

The program calculates:

ATR
ATR_PCT

ATR is useful for understanding how much a stock normally moves.

This becomes particularly important later because the program uses ATR to calculate:

Stop Loss

and:

Target


24. Example of ATR

Suppose:

Stock Price = ₹1,000
ATR = ₹20

The program calculates:

Stop Loss

₹1,000 - 1.5 × ₹20

= ₹970

Target

₹1,000 + 2.5 × ₹20

= ₹1,050

So the volatility of the stock influences the proposed risk levels.


25. Volume Ratio

The program calculates:

VOL20
VOLUME_RATIO

The volume ratio compares today's volume with its recent average.

For example:

Today's Volume = 2,000,000
20-day average = 1,000,000

Then:

Volume Ratio = 2.0

This means today's volume is approximately:

2 times the recent average.

The scanner later requires:

VOLUME_RATIO >= 1.0

for a strong trend-based setup.


26. 52-Week High

The program calculates:

HIGH_52W
LOW_52W
DIST_52W_HIGH

The 52-week high is the highest price over approximately 252 trading observations.

This is particularly interesting for momentum traders.

For example:

52-week high = ₹1,000
Current price = ₹950

Then:

Distance from high = -5%

The stock is approximately 5% below its yearly high.


27. Trend Score

The program creates:

TREND_SCORE

The score checks four conditions:

Close > SMA20
Close > SMA50
Close > SMA200
EMA20 > EMA50

Each true condition contributes one point.

Therefore the score can range approximately from:

0 to 4

For example:

Stock A

Close > SMA20     ✓
Close > SMA50     ✓
Close > SMA200    ✓
EMA20 > EMA50     ✓

Trend Score:

4

This indicates a strong alignment of the chosen trend conditions.


28. NIFTY Market Regime

The program doesn't look at individual stocks in isolation.

It also looks at the broader NIFTY market.

It calculates:

NIFTY_RET_1
NIFTY_RET_5
NIFTY_RET_20
NIFTY_TREND

This is important because:

A stock can behave differently when the overall market is strong versus weak.

The AI therefore receives some information about the broader market environment.


29. India VIX

The program also incorporates:

VIX
VIX_CHANGE

India VIX is a volatility indicator associated with expected market volatility.

The AI therefore receives information about:

Stock
+
Market
+
Volatility

rather than only looking at the stock's own price.


30. Cell 6 – Creating the Prediction Target

This is the key step that turns the program into a prediction model.

The code calculates:

FUTURE_RETURN_5D

The formula is essentially:

Price 5 days later
------------------
Current price
- 1

For example:

Current price = ₹1,000
Price after 5 days = ₹1,040

Then:

Future return = 4%

31. The AI Classification Target

The program then creates:

TARGET

using:

df["FUTURE_RETURN_5D"] > 0.02

This means:

Did the stock gain more than 2% over the next five trading observations?

If yes:

TARGET = 1

If not:

TARGET = 0

Therefore the classifier is essentially learning:

Is there a greater-than-2% five-day opportunity?


32. Two AI Models Are Used

This is one of the most interesting parts of the program.

The notebook trains two different XGBoost models.

Model 1 – Classifier

Answers:

What is the probability that the stock will achieve the desired positive-return condition?

Model 2 – Regressor

Answers:

What 5-day return does the model expect?

So instead of asking only:

BUY or SELL?

the program asks two questions:

Question 1:
How likely is the positive outcome?

Question 2:
How large might the 5-day return be?

33. What Is XGBoost?

XGBoost is a machine-learning algorithm based on decision-tree boosting.

You don't need to understand its mathematics to use the notebook.

A simple way to think about it is:

The model looks at many relationships in the historical data and builds a collection of decision rules to make predictions.

The program uses:

XGBClassifier

for classification.

And:

XGBRegressor

for return prediction.


34. What Information Does the AI See?

The notebook defines a list called:

FEATURES

It contains information such as:

1-day return
3-day return
5-day return
10-day return
20-day return

Distance from SMA20
Distance from SMA50
Distance from SMA200

RSI

MACD
MACD Signal
MACD Histogram

ADX
DI+
DI-

Bollinger Band Width
Bollinger Position

ATR percentage

Volume Ratio

Distance from 52-week high

Trend Score

NIFTY returns
NIFTY trend

India VIX
VIX change

This is a fairly rich feature set.


35. Think of Features Like a Doctor's Report

Imagine a doctor trying to understand a patient.

The doctor doesn't look at only one measurement.

They may consider:

Temperature
Blood pressure
Heart rate
Age
Symptoms
History

Similarly, this AI doesn't look only at price.

It considers:

Momentum
Trend
Volatility
Volume
Market trend
VIX
Moving averages
Technical indicators

Together, these become the model's "information set."


36. Cell 8 – Building the Dataset

The program loops through the NIFTY 50 stocks:

for symbol in NIFTY50:

For every stock, it:

  1. Gets historical data

  2. Creates technical features

  3. Creates the future-return target

  4. Removes incomplete rows

  5. Stores the finished dataset

This creates a separate dataset for each stock.


37. Why Are Missing Rows Removed?

Many technical indicators require historical observations.

For example:

SMA200

cannot be calculated properly until enough historical data exists.

Therefore the beginning of the dataset can contain missing values.

The program uses:

dropna()

to remove rows where required information isn't available.


38. Cell 9 – Training the AI

Now the machine learning begins.

The program requires:

len(df) >= 1000

before training a stock model.

Then it divides the historical data:

80% → Training
20% → Testing

39. Why Is the Data Split Chronologically?

This is extremely important for financial data.

The program does:

split = int(len(df) * 0.80)

train = df.iloc[:split]
test = df.iloc[split:]

So the earlier observations are used for training.

Later observations are used for testing.

Conceptually:

PAST
───────────────────────
Training Data
───────────────────────
Future
───────────────────────
Testing Data
───────────────────────

This is much more appropriate for time-series research than randomly shuffling stock-market observations.


40. Training the Classifier

The classifier is trained using:

clf.fit(
    X_train,
    y_train
)

It learns the relationship between:

FEATURES
   ↓
TARGET

where TARGET represents the future 5-day outcome condition.


41. Training the Return Predictor

The second model learns:

FEATURES
   ↓
FUTURE_RETURN_5D

This is the regression model.

So the two models have different jobs.

ModelJob
ClassifierProbability of positive outcome
RegressorExpected 5-day return

42. Measuring Model Performance

The program calculates:

Accuracy

How often the predicted class matched the actual class.

Precision

When the model predicts a positive signal, how often was it correct?

AUC

A measure of how well the model distinguishes between the two classes across probability thresholds.

The program stores these results in:

performance_df

43. Why Accuracy Alone Is Not Enough

Suppose a model has:

Accuracy = 80%

That sounds excellent.

But accuracy alone can be misleading.

For example, if only 20% of observations are positive and the model predicts "negative" almost all the time, it could still obtain high accuracy.

Therefore the notebook also calculates:

Precision
AUC

This is a good practice.


44. Cell 11 – Generating Today's Signals

This is where everything becomes interesting.

The program takes the latest available data for every trained stock.

It creates:

X_latest

This is essentially:

Today's feature information for the stock.

The AI then receives this information.


45. AI Probability

The classifier produces:

probability

For example:

0.78

The program converts this to:

78%

This can be interpreted as the model's estimated probability for the positive class under the model's definition.

It is not the same thing as a guaranteed 78% chance of making money.


46. Expected 5-Day Return

The regression model produces:

expected_return

Suppose the output is:

0.035

That corresponds to:

3.5%

The program therefore has two important outputs:

AI Probability = 78%

Expected 5-Day Return = 3.5%

47. Entry Price

The program uses the current/latest price as:

Entry = price

For example:

Current price = ₹1,250

Then:

Entry = ₹1,250

This is a model-generated reference level, not a guarantee that the actual execution price will be the same.


48. Stop Loss Calculation

The program uses:

stop_loss = price - 1.5 * atr

So the stop is:

1.5 ATR below the current price.

Example:

Price = ₹1,000
ATR = ₹20

Then:

Stop Loss
= 1000 - 1.5 × 20
= ₹970

49. Target Calculation

The program uses:

target = price + 2.5 * atr

So:

Price = ₹1,000
ATR = ₹20

Then:

Target
= 1000 + 2.5 × 20
= ₹1,050

However, the program also says:

target = max(target, price * 1.05)

This means the target cannot be below approximately:

5% above the current price.


50. Trend Filter

The scanner requires several conditions for a strong trend.

The program checks:

Close > SMA50

SMA50 > SMA200

RSI > 50

ADX > 18

In simple language:

The stock should be above its 50-day average, the 50-day average should be above its 200-day average, momentum should be positive, and the trend should have reasonable strength.


51. Volume Filter

The scanner also checks:

VOLUME_RATIO >= 1.0

That means:

Current volume should be at least around its 20-period average.

This attempts to avoid situations where the price is moving without adequate trading activity.


52. How Does the Program Decide STRONG BUY?

The program requires:

AI Probability >= 70%

Expected 5-day return >= 2%

Trend conditions satisfied

Volume condition satisfied

If all these conditions are met:

STRONG BUY

53. How Does It Decide BUY?

The next level is:

AI Probability >= 60%

Expected 5-day return > 1%

Trend condition satisfied

Then:

BUY

54. When Does It Say SELL?

If:

AI Probability < 40%

the program labels the stock:

SELL

Otherwise, if none of the conditions are strong enough:

NO TRADE

55. Understanding the Four Signals

The final signal categories are:

SignalMeaning
STRONG BUYStrongest combination according to the programmed rules
BUYPositive setup but doesn't meet all STRONG BUY requirements
SELLAI probability below 40%
NO TRADEConditions are not strong enough

These labels are generated by the code's rules.

They should not be interpreted as guaranteed trading recommendations.


56. Cell 13 – Showing Buy Candidates

The program extracts:

STRONG BUY
BUY

from the complete signal table.

It displays:

Symbol
Price
AI Probability
Expected 5D Return
RSI
ADX
Volume Ratio
Entry
Stop Loss
Target
Signal

This is one of the most useful tables for a trader.


57. AI Score

The notebook then calculates:

AI_SCORE

This creates a ranking system.

The score combines:

AI Probability
+
Expected Return
+
Trend Score
+
ADX
+
Volume Ratio

with different weights.

The largest weight in the formula is given to:

AI Probability

followed by:

Expected Return
Trend Score
ADX
Volume

The purpose is to rank the stocks rather than simply classify them.


58. Why Do We Need an AI Score?

Suppose the scanner finds:

Stock A → BUY
Stock B → BUY
Stock C → STRONG BUY
Stock D → STRONG BUY

You may still want to know:

Which stock has the strongest overall score?

The AI Score attempts to answer that.

The program sorts stocks:

sort_values(
    "AI_SCORE",
    ascending=False
)

Therefore the highest-scoring stocks appear first.


59. Cell 15 – Backtesting

Now comes another important section.

The program tests the strategy on historical data.

It takes the 20% test portion and calculates:

AI Probability
Expected Return
Signal
Future 5-day return

The historical signal rule is:

Probability >= 70%
AND
Expected return >= 2%
AND
Close > SMA50
AND
SMA50 > SMA200

When these conditions are satisfied, the strategy takes the future 5-day return as the strategy return.


60. What Is a Backtest?

Suppose the program finds that a historical signal occurred on:

1 March

It then looks at what happened over the following five trading observations.

For example:

Entry = ₹1,000

Price after 5 trading days = ₹1,035

Return = +3.5%

The program records that as a successful historical outcome.

Repeat this across many historical signals and you can evaluate the strategy.


61. Backtest Metrics

The program calculates:

Number of Trades

How many historical signals occurred.

Win Rate

Percentage of trades with positive returns.

Profit Factor

A comparison between gross profits and gross losses.

Total Return

Compounded result of the recorded strategy returns.

Maximum Drawdown

Largest decline in the simulated equity curve.


62. Example

Suppose historical trades were:

+4%
+3%
-2%
+5%
-1%

The program can calculate:

Number of trades = 5
Winning trades = 3
Win rate = 60%

It can also calculate profit factor and cumulative performance.


63. Maximum Drawdown

Suppose your simulated account goes:

₹1,00,000
₹1,10,000
₹1,20,000
₹1,05,000

The peak was:

₹1,20,000

and it subsequently declined to:

₹1,05,000

That decline is important because it tells us about the risk experienced by the strategy.

A strategy producing high returns with very large drawdowns may be difficult to follow in real life.


64. Cell 18 – Building a Portfolio

The notebook goes one step further.

Instead of considering only individual stocks, it creates a portfolio approach.

For each date, it identifies stocks generating signals.

Then it sorts them by:

AI Probability

and selects up to:

Top 5

stocks.

The portfolio return is calculated using the average future return of those selected stocks.


65. Portfolio Equity Curve

The program calculates:

Equity
Peak
Drawdown

and then plots the equity curve.

The equity curve helps answer:

How would the simulated portfolio have grown or declined over time?

The chart is titled:

AI NIFTY 50 Strategy Equity Curve

66. Feature Importance

The notebook also examines:

model.feature_importances_

This produces a ranking of features that the XGBoost classifier considered useful in making its predictions.

The program displays the top 20 features.

For example, depending on the trained model, important features might include things such as:

RSI
Volume Ratio
NIFTY Trend
Distance from SMA
ADX
Momentum

The exact ranking should be taken from the actual model output rather than assumed in advance.


67. Why Is Feature Importance Useful?

Imagine the AI says:

BUY

You may want to know:

What information was useful to the model?

Feature importance can provide a clue.

However, feature importance should not be interpreted as:

"This indicator caused the stock to rise."

It tells us about the model's use of the feature, not a guaranteed causal relationship.


68. Final Scan

Cell 22 creates:

final_scan

This becomes the main final table.

It contains:

Symbol
Price
AI Probability
Expected 5D Return
RSI
ADX
Volume Ratio
Trend Score
Entry
Stop Loss
Target
AI Score
Signal

This is essentially the scanner's final dashboard.


69. STRONG BUY List

Cell 23 filters the final table:

final_scan["Signal"] == "STRONG BUY"

and displays all stocks that satisfy the strong-buy conditions.

It also prints:

Number of STRONG BUY signals

This provides a quick summary.


70. Exporting the Results to Excel

The final cell creates:

NIFTY50_AI_Scanner.xlsx

The Excel workbook contains three sheets.

Sheet 1 – AI Signals

Contains the current scan.

Sheet 2 – Model Performance

Contains:

Accuracy
Precision
AUC

for the individual stock models.

Sheet 3 – Backtest

Contains:

Trades
Win Rate
Profit Factor
Total Return
Maximum Drawdown

This makes the project much easier to analyse outside Colab.


71. Complete Program in Simple English

The entire notebook can be understood as:

1. Install Python packages

2. Import packages

3. Create NIFTY 50 stock list

4. Download historical prices

5. Calculate technical indicators

6. Calculate market information

7. Calculate India VIX information

8. Create 5-day future-return target

9. Train an AI classifier

10. Train an AI return predictor

11. Test the models

12. Calculate today's probability

13. Calculate expected 5-day return

14. Check trend

15. Check volume

16. Generate BUY/SELL/NO TRADE

17. Calculate stop loss

18. Calculate target

19. Calculate AI score

20. Rank stocks

21. Backtest the strategy

22. Calculate portfolio performance

23. Plot equity curve

24. Export everything to Excel

72. What Makes This Different from a Normal Stock Screener?

A traditional screener might say:

RSI > 50
AND
Price > SMA50
AND
Volume > Average Volume

This program goes further.

It uses:

Technical Indicators
        +
Market Regime
        +
India VIX
        +
Machine Learning
        +
Expected Return
        +
Trend Filters
        +
Volume Filter
        +
Risk Levels
        +
Backtesting

Therefore it is better described as an:

AI-Assisted Quantitative Stock Research System

rather than simply a technical scanner.


73. Important: AI Does Not Mean Guaranteed Prediction

This is perhaps the most important lesson for a beginner.

Suppose the program says:

AI Probability = 75%

It does not mean:

"There is a guaranteed 75% chance that I will make money."

It means that the trained model assigns a probability of 75% to its defined positive class based on the historical relationships it learned.

Real markets can behave differently.


74. Important Limitation #1 – Historical Data

The AI learns from historical data.

Markets evolve.

A relationship that worked between:

2015–2022

may behave differently later.

Therefore:

Past performance does not guarantee future performance.


75. Important Limitation #2 – No Transaction Costs in the Backtest

The current backtest logic does not explicitly subtract:

  • Brokerage

  • STT

  • Exchange charges

  • GST

  • Slippage

  • Bid-ask spread

  • Other trading costs

Therefore the backtested return can be more optimistic than actual trading performance.

A production version should incorporate realistic costs.


76. Important Limitation #3 – Five-Day Trades Can Overlap

The strategy uses:

Future 5-day return

for each signal.

If signals occur on consecutive days, their five-day holding periods can overlap.

The current backtest does not explicitly maintain a real position ledger showing:

Entry date
Exit date
Capital allocated
Open positions
Position overlap

Therefore the backtest should be treated as a research approximation, not a broker-ready trading simulator.


77. Important Limitation #4 – Portfolio Backtest Simplification

The portfolio section averages the future returns of up to five selected stocks.

A real portfolio needs to consider:

  • Capital

  • Position size

  • Entry price

  • Exit date

  • Position overlap

  • Transaction costs

  • Slippage

  • Cash

  • Rebalancing

  • Maximum exposure

The current code simplifies these factors.


78. Important Limitation #5 – Model Probability

The classifier's probability is generated by XGBoost.

A probability such as:

72%

should not automatically be assumed to be perfectly calibrated.

For serious quantitative research, probability calibration should also be investigated.


79. Important Limitation #6 – The AI Score Is a Custom Formula

The:

AI_SCORE

is not produced directly by XGBoost.

It is a manually designed formula combining:

AI probability
Expected return
Trend score
ADX
Volume

with selected weights.

Therefore:

AI Probability and AI Score are two different things.

The first comes from the classifier.

The second is a ranking formula created by the programmer.


80. Important Improvement – Use Decimal Prices

The uploaded code converts OHLC values to integers:

df["Close"] = df["Close"].astype(int)

This removes decimal precision.

For example:

₹1,234.75

becomes approximately:

₹1,234

For research, it would generally be preferable to retain the original floating-point price values.

This is a technical improvement worth making.


81. Important Improvement – Verify the NIFTY 50 List

The notebook contains a manually maintained list.

NIFTY 50 constituents change over time.

Therefore, a production-quality scanner should periodically update its universe rather than assuming that a hard-coded list remains current forever.


82. How a Complete Beginner Should Run It

Follow these steps.

Step 1

Open the notebook in Google Colab.

Step 2

Connect the runtime.

Step 3

Run Cell 1.

Step 4

Run Cell 2.

Step 5

Check the NIFTY 50 list.

Step 6

Fix the missing comma between:

POWERGRID.NS
RELIANCE.NS

Step 7

Run the data-download cell.

Step 8

Wait for the historical data to download.

Step 9

Run the feature-generation cells.

Step 10

Wait for the AI models to train.

Step 11

Look at model performance.

Step 12

Look at:

signals_df

Step 13

Look at:

final_scan

Step 14

Review:

STRONG BUY

Step 15

Review the backtest.

Step 16

Download:

NIFTY50_AI_Scanner.xlsx

83. How to Read the Final Table

Suppose your final table looks like:

SymbolPriceAI ProbabilityExpected 5D ReturnRSIADXVolumeStop LossTargetSignal
ABC₹1,00078%3.4%62251.4₹970₹1,050STRONG BUY
XYZ₹80064%1.8%57221.2₹770₹840BUY

Don't just look at the word:

STRONG BUY

Look at the entire row.

Ask:

  1. Is AI probability high?

  2. Is expected return attractive?

  3. Is the trend strong?

  4. Is volume supportive?

  5. Is the stop loss reasonable?

  6. Is the target realistic?

  7. What did the historical backtest show?

  8. How many trades generated the backtest?

  9. What was maximum drawdown?


84. A Better Way to Use the Scanner

Don't use this program as:

"The computer said BUY, therefore I will buy."

Instead use it as:

"The computer found a candidate. Now I will investigate it."

A better workflow is:

AI Scanner
    ↓
Shortlist
    ↓
Check Chart
    ↓
Check Market Trend
    ↓
Check Fundamental Context
    ↓
Check Risk/Reward
    ↓
Check Position Size
    ↓
Make Decision

85. Beginner Experiment #1

Change nothing.

Run the complete notebook.

Record:

Number of STRONG BUY signals
Number of BUY signals
Number of SELL signals
Top 10 AI Score stocks

This establishes your baseline.


86. Beginner Experiment #2

Change the probability threshold.

Currently STRONG BUY requires:

70%

Experiment with:

65%
70%
75%
80%

Then compare:

Number of trades
Win rate
Profit factor
Total return
Maximum drawdown

This teaches you how changing a threshold changes a strategy.


87. Beginner Experiment #3

Compare Stocks

Run the scanner and identify five stocks.

Then compare:

AI Probability
Expected Return
RSI
ADX
Volume Ratio
Trend Score
Backtest

Ask:

Does the highest AI probability always produce the highest historical return?

This is an excellent quantitative-research question.


88. Beginner Experiment #4 – Build Your Own Ranking

Create your own ranking rules.

For example:

AI Probability > 65%
RSI > 50
ADX > 20
Volume Ratio > 1
Price > SMA50
SMA50 > SMA200

Then compare the results with the original scanner.

You have now started modifying an AI trading system.


89. What You Can Add in the Future

This project can become much more powerful.

Possible upgrades include:

Risk Management

  • Position sizing

  • Maximum capital per stock

  • Maximum portfolio risk

  • ATR-based sizing

Better Backtesting

  • Transaction costs

  • Slippage

  • Holding-period simulation

  • Portfolio capital tracking

  • Overlapping trade handling

Advanced Validation

  • Walk-forward testing

  • Out-of-sample testing

  • Cross-validation for time series

  • Probability calibration

More Market Data

  • Sector indices

  • Market breadth

  • FII/DII activity

  • Advance/decline ratio

  • Sector relative strength

Options

  • Option chain

  • Open interest

  • PCR

  • IV

  • IV Rank

  • Max Pain

AI

  • Random Forest

  • LightGBM

  • Logistic Regression

  • Neural Networks

  • Ensemble models

  • Regime classification


90. The Ideal Future Version

A more advanced version of this project could eventually look like:

             NIFTY 50
                 ↓
        Market Regime
                 ↓
      Stock Feature Engine
                 ↓
       ┌─────────┴─────────┐
       ↓                   ↓
 Classification       Regression
       ↓                   ↓
 Probability         Expected Return
       └─────────┬─────────┘
                 ↓
          Trend Filter
                 ↓
          Volume Filter
                 ↓
          Risk Engine
                 ↓
       Position Sizing
                 ↓
          AI Ranking
                 ↓
       Portfolio Builder
                 ↓
        Backtest Engine
                 ↓
       Trading Dashboard

That would transform the notebook into a much more complete quantitative research platform.


91. Final Takeaway

The uploaded program is much more than a simple stock screener.

It combines:

Python
+
Historical Market Data
+
Technical Analysis
+
NIFTY Market Data
+
India VIX
+
Machine Learning
+
Return Prediction
+
Trend Filtering
+
Volume Analysis
+
Risk Levels
+
Backtesting
+
Portfolio Analysis

The central idea is:

Use historical market information to train an AI model, use the model to estimate the probability and expected return of a future 5-day move, then combine that information with technical filters and risk rules to rank NIFTY 50 stocks.

The most important thing for a beginner to remember is:

AI Prediction
       ≠
Guaranteed Profit

The scanner is best used as a research and decision-support tool.

The real objective is not to find a magical BUY signal.

It is to build a systematic process:

Research → Test → Validate → Manage Risk → Improve → Repeat


92. Quick Reference – What Each Cell Does

CellPurpose
1Install Python packages
2Import libraries
3Define NIFTY 50 universe
4Download historical data
5Create technical and market features
6Create 5-day prediction target
7Define AI features
8Build datasets
9Train and test AI models
10Compare model performance
11Generate current AI signals
12Validate feature data
13Show BUY candidates
14Calculate AI Score
15Backtest individual stocks
16Run backtests for all stocks
17Calculate backtest statistics
18Build top-5 portfolio simulation
19Calculate portfolio return/drawdown
20Plot equity curve
21Display feature importance
22Create final scanner
23Show STRONG BUY stocks
24Export results to Excel

Final Warning

Before using this scanner with real money, the following should be improved:

  • Fix the POWERGRID/RELIANCE ticker-list issue

  • Keep decimal price precision

  • Update the NIFTY 50 universe automatically

  • Add brokerage and slippage

  • Build a proper trade/position-based backtester

  • Handle overlapping 5-day positions correctly

  • Perform walk-forward/out-of-sample testing

  • Validate probability calibration

  • Test across different market regimes

  • Add proper position sizing and portfolio risk controls

Only after these improvements should the system be considered for serious paper-trading evaluation.

This article is for educational and research purposes. Machine-learning predictions and historical backtests do not guarantee future market performance.

Comments

  1. if you have idea, i can automate for you. Feel free to connect me.

    ReplyDelete

Post a Comment

Popular posts from this blog

📈 Opening Range Breakout Strategy Explained with Python Code

🐢 Turtle Trading – Classic Breakout and Trend-Following Method (with Python)

100 trading strategies