AI NIFTY 50 Stock Scanner with Python
AI NIFTY 50 Stock Scanner with Python – Complete Beginner's Guide
Build Your Own AI Stock Scanner in Google Colab
Imagine opening your computer every morning and asking:
Which NIFTY 50 stocks currently have the strongest technical and AI-based setup?
Instead of manually checking 50 stocks one by one, Python can do much of this work automatically.
The AI NIFTY 50 Stock Scanner presented in this tutorial downloads historical stock-market data, calculates technical indicators, trains machine-learning models, predicts the probability of a positive 5-day return, applies trend and volume filters, calculates entry/stop-loss/target levels, ranks stocks and finally exports the results to Excel.
The complete program is designed to answer a question such as:
"Which NIFTY 50 stocks currently have the strongest combination of AI probability, expected return, trend, momentum and volume?"
This article explains the uploaded Python program from the ground up so that even someone with very little Python knowledge can understand and use it.
1. What Are We Building?
The program is an:
AI-Based NIFTY 50 Stock Scanner
It performs approximately this workflow:
NIFTY 50 Stocks
↓
Download Historical Data
↓
Calculate Technical Indicators
↓
Create Machine Learning Features
↓
Create 5-Day Prediction Target
↓
Train AI Models
↓
Test Models
↓
Predict Today's Probability
↓
Estimate 5-Day Return
↓
Apply Trend + Volume Filters
↓
Calculate Stop Loss & Target
↓
Calculate AI Score
↓
Rank Stocks
↓
Generate BUY / STRONG BUY / SELL / NO TRADE
↓
Export Results to Excel
This is considerably more advanced than a simple RSI or moving-average scanner.
2. Who Can Use This Program?
You don't need to be an experienced programmer.
This project is suitable for someone who knows:
Basic stock-market terminology
What a stock price is
What BUY and SELL mean
Basic computer usage
Python knowledge is helpful, but you can initially use the notebook by simply running the cells.
Later, you can learn Python by modifying individual sections.
3. What Is Google Colab?
The program is designed to run in Google Colab.
Google Colab is an online Python environment.
Normally, if you want to run Python programs, you might need to install:
Python
Pandas
NumPy
XGBoost
Scikit-learn
Matplotlib
other packages
With Colab, much of this setup can be done directly inside the notebook.
Think of Colab as:
A Python computer running inside your web browser.
4. Understanding the Notebook
The program is divided into 24 code cells.
Don't be frightened by the number.
You don't need to understand all 24 cells before running the program.
The easiest approach is:
Run
↓
Observe
↓
Understand
↓
Modify
↓
Experiment
Always start from the first cell and work downward.
5. Cell 1 – Installing Python Libraries
The first cell contains:
!pip -q install yfinance xgboost ta scikit-learn pandas numpy matplotlib seaborn
This installs the libraries required by the program.
Let's understand them.
yfinance
Used to download market data.
XGBoost
Used for the machine-learning models.
ta
Provides technical-analysis indicators.
scikit-learn
Provides machine-learning evaluation tools.
pandas
Used for tables and time-series data.
numpy
Used for numerical calculations.
matplotlib
Used for charts.
seaborn
Used for visualization.
The important point for a beginner is:
Don't manually install these one by one. Run Cell 1 first.
6. Cell 2 – Importing the Libraries
The next cell contains imports such as:
import numpy as np
import pandas as pd
import yfinance as yf
import matplotlib.pyplot as plt
import seaborn as sns
Importing means:
"Make these tools available to my Python program."
For example:
import pandas as pd
allows the program to use:
pd.DataFrame()
7. What Is Pandas?
Pandas is extremely important for this project.
Think of Pandas as:
Excel inside Python.
A stock-data table might look like:
| Date | Open | High | Low | Close | Volume |
|---|---|---|---|---|---|
| 1 Jan | 1,000 | 1,020 | 990 | 1,015 | 500,000 |
| 2 Jan | 1,015 | 1,040 | 1,005 | 1,035 | 650,000 |
| 3 Jan | 1,035 | 1,050 | 1,025 | 1,045 | 720,000 |
Pandas stores this kind of information in something called a DataFrame.
8. Cell 3 – Defining the NIFTY 50 Stocks
The program creates a list:
NIFTY50 = [
'ADANIENT.NS',
'ADANIPORTS.NS',
'APOLLOHOSP.NS',
...
]
These are Yahoo Finance ticker symbols for Indian stocks.
The .NS suffix generally identifies NSE-listed instruments in the Yahoo Finance symbol format.
For example:
TCS.NS
INFY.NS
RELIANCE.NS
ONGC.NS
The program then loops through these symbols.
9. Important Issue in the Stock List
There is a small coding issue in the uploaded notebook.
This part appears as:
'POWERGRID.NS''RELIANCE.NS'
There is no comma between the two strings.
Python automatically joins adjacent strings, so this can effectively become one incorrect ticker:
POWERGRID.NSRELIANCE.NS
Correct version
It should be:
'POWERGRID.NS',
'RELIANCE.NS',
Why is this important?
Because the program may fail to download data correctly for those entries.
I strongly recommend correcting this before using the scanner.
10. Cell 4 – Downloading Historical Data
The program sets:
START_DATE = "2015-01-01"
This means the program attempts to obtain historical data starting from January 1, 2015.
It downloads data for:
NIFTY 50 stocks
+
NIFTY index
+
India VIX
The program uses:
yf.download()
to obtain the market data.
11. What Data Is Downloaded?
For every stock, the program keeps:
Open
High
Low
Close
Volume
These are known as OHLCV.
O – Open
Opening price.
H – High
Highest price during the session.
L – Low
Lowest price during the session.
C – Close
Closing price.
V – Volume
Number of shares traded.
12. Why Is Historical Data Necessary?
The AI cannot learn from today's price alone.
It needs historical examples.
Imagine giving the AI 10 years of information:
2015
2016
2017
...
2025
2026
It can study relationships between:
Price momentum
RSI
MACD
Trend
Volume
Volatility
NIFTY market behaviour
India VIX
and subsequent stock performance.
13. What Does data = {} Mean?
The program creates:
data = {}
A Python dictionary is like a labelled storage cabinet.
For example:
data["TCS.NS"]
data["INFY.NS"]
data["RELIANCE.NS"]
Each label contains the corresponding stock's historical DataFrame.
14. Cell 5 – Creating Technical Features
This is one of the most important sections of the entire program.
The function:
create_features()
takes raw stock prices and creates additional information.
These additional variables are called:
Features
A feature is simply information that the machine-learning model can use.
15. Return Features
The program calculates:
RET_1
RET_3
RET_5
RET_10
RET_20
These represent price changes over different periods.
For example:
df["RET_5"] = df["Close"].pct_change(5)
means:
Calculate the percentage change in closing price over approximately five trading observations.
This tells the AI whether the stock has recently been moving upward or downward.
16. Moving Averages
The program calculates:
SMA20
SMA50
SMA100
SMA200
SMA means:
Simple Moving Average
For example, SMA50 represents the average closing price over the previous 50 observations.
If:
Price > SMA50
the stock is trading above its 50-period average.
If:
Price < SMA50
the stock is trading below it.
17. Why Are Multiple Moving Averages Used?
The program doesn't rely on just one moving average.
It uses:
20-day
50-day
100-day
200-day
This gives the model information about both short-term and long-term trends.
For example:
Price > SMA20
Price > SMA50
Price > SMA200
suggests a stronger upward trend than if the price were below all three.
18. Distance from Moving Averages
The program also calculates:
DIST_SMA20
DIST_SMA50
DIST_SMA200
For example:
Close / SMA50 - 1
Suppose:
Close = ₹1,100
SMA50 = ₹1,000
Then:
1,100 / 1,000 - 1
= 0.10
= 10%
So the stock is:
10% above its 50-day moving average.
19. RSI – Relative Strength Index
The program calculates:
RSI
using a 14-period RSI.
RSI is a momentum indicator.
It helps answer:
How strong has recent price movement been?
The program later uses:
RSI > 50
as part of its trend filter for strong setups.
20. MACD
The program calculates:
MACD
MACD_SIGNAL
MACD_HIST
MACD is another momentum/trend indicator.
Instead of using only the closing price, the AI gets information about the relationship between different moving averages.
The histogram provides additional information about the difference between MACD and its signal line.
21. ADX
The program calculates:
ADX
DI_PLUS
DI_MINUS
ADX is commonly used to measure trend strength.
In this scanner, the strong-buy filter requires:
ADX > 18
So the program wants a reasonably established trend rather than a completely directionless market.
22. Bollinger Bands
The program calculates:
BB_HIGH
BB_LOW
BB_MID
BB_WIDTH
BB_POSITION
Bollinger Bands provide information about:
Price location
Volatility
Expansion/contraction of price movement
The program doesn't simply ask:
"Is price above the upper band?"
Instead, it converts Bollinger information into numerical features that the AI can use.
23. ATR – Average True Range
The program calculates:
ATR
ATR_PCT
ATR is useful for understanding how much a stock normally moves.
This becomes particularly important later because the program uses ATR to calculate:
Stop Loss
and:
Target
24. Example of ATR
Suppose:
Stock Price = ₹1,000
ATR = ₹20
The program calculates:
Stop Loss
₹1,000 - 1.5 × ₹20
= ₹970
Target
₹1,000 + 2.5 × ₹20
= ₹1,050
So the volatility of the stock influences the proposed risk levels.
25. Volume Ratio
The program calculates:
VOL20
VOLUME_RATIO
The volume ratio compares today's volume with its recent average.
For example:
Today's Volume = 2,000,000
20-day average = 1,000,000
Then:
Volume Ratio = 2.0
This means today's volume is approximately:
2 times the recent average.
The scanner later requires:
VOLUME_RATIO >= 1.0
for a strong trend-based setup.
26. 52-Week High
The program calculates:
HIGH_52W
LOW_52W
DIST_52W_HIGH
The 52-week high is the highest price over approximately 252 trading observations.
This is particularly interesting for momentum traders.
For example:
52-week high = ₹1,000
Current price = ₹950
Then:
Distance from high = -5%
The stock is approximately 5% below its yearly high.
27. Trend Score
The program creates:
TREND_SCORE
The score checks four conditions:
Close > SMA20
Close > SMA50
Close > SMA200
EMA20 > EMA50
Each true condition contributes one point.
Therefore the score can range approximately from:
0 to 4
For example:
Stock A
Close > SMA20 ✓
Close > SMA50 ✓
Close > SMA200 ✓
EMA20 > EMA50 ✓
Trend Score:
4
This indicates a strong alignment of the chosen trend conditions.
28. NIFTY Market Regime
The program doesn't look at individual stocks in isolation.
It also looks at the broader NIFTY market.
It calculates:
NIFTY_RET_1
NIFTY_RET_5
NIFTY_RET_20
NIFTY_TREND
This is important because:
A stock can behave differently when the overall market is strong versus weak.
The AI therefore receives some information about the broader market environment.
29. India VIX
The program also incorporates:
VIX
VIX_CHANGE
India VIX is a volatility indicator associated with expected market volatility.
The AI therefore receives information about:
Stock
+
Market
+
Volatility
rather than only looking at the stock's own price.
30. Cell 6 – Creating the Prediction Target
This is the key step that turns the program into a prediction model.
The code calculates:
FUTURE_RETURN_5D
The formula is essentially:
Price 5 days later
------------------
Current price
- 1
For example:
Current price = ₹1,000
Price after 5 days = ₹1,040
Then:
Future return = 4%
31. The AI Classification Target
The program then creates:
TARGET
using:
df["FUTURE_RETURN_5D"] > 0.02
This means:
Did the stock gain more than 2% over the next five trading observations?
If yes:
TARGET = 1
If not:
TARGET = 0
Therefore the classifier is essentially learning:
Is there a greater-than-2% five-day opportunity?
32. Two AI Models Are Used
This is one of the most interesting parts of the program.
The notebook trains two different XGBoost models.
Model 1 – Classifier
Answers:
What is the probability that the stock will achieve the desired positive-return condition?
Model 2 – Regressor
Answers:
What 5-day return does the model expect?
So instead of asking only:
BUY or SELL?
the program asks two questions:
Question 1:
How likely is the positive outcome?
Question 2:
How large might the 5-day return be?
33. What Is XGBoost?
XGBoost is a machine-learning algorithm based on decision-tree boosting.
You don't need to understand its mathematics to use the notebook.
A simple way to think about it is:
The model looks at many relationships in the historical data and builds a collection of decision rules to make predictions.
The program uses:
XGBClassifier
for classification.
And:
XGBRegressor
for return prediction.
34. What Information Does the AI See?
The notebook defines a list called:
FEATURES
It contains information such as:
1-day return
3-day return
5-day return
10-day return
20-day return
Distance from SMA20
Distance from SMA50
Distance from SMA200
RSI
MACD
MACD Signal
MACD Histogram
ADX
DI+
DI-
Bollinger Band Width
Bollinger Position
ATR percentage
Volume Ratio
Distance from 52-week high
Trend Score
NIFTY returns
NIFTY trend
India VIX
VIX change
This is a fairly rich feature set.
35. Think of Features Like a Doctor's Report
Imagine a doctor trying to understand a patient.
The doctor doesn't look at only one measurement.
They may consider:
Temperature
Blood pressure
Heart rate
Age
Symptoms
History
Similarly, this AI doesn't look only at price.
It considers:
Momentum
Trend
Volatility
Volume
Market trend
VIX
Moving averages
Technical indicators
Together, these become the model's "information set."
36. Cell 8 – Building the Dataset
The program loops through the NIFTY 50 stocks:
for symbol in NIFTY50:
For every stock, it:
Gets historical data
Creates technical features
Creates the future-return target
Removes incomplete rows
Stores the finished dataset
This creates a separate dataset for each stock.
37. Why Are Missing Rows Removed?
Many technical indicators require historical observations.
For example:
SMA200
cannot be calculated properly until enough historical data exists.
Therefore the beginning of the dataset can contain missing values.
The program uses:
dropna()
to remove rows where required information isn't available.
38. Cell 9 – Training the AI
Now the machine learning begins.
The program requires:
len(df) >= 1000
before training a stock model.
Then it divides the historical data:
80% → Training
20% → Testing
39. Why Is the Data Split Chronologically?
This is extremely important for financial data.
The program does:
split = int(len(df) * 0.80)
train = df.iloc[:split]
test = df.iloc[split:]
So the earlier observations are used for training.
Later observations are used for testing.
Conceptually:
PAST
───────────────────────
Training Data
───────────────────────
Future
───────────────────────
Testing Data
───────────────────────
This is much more appropriate for time-series research than randomly shuffling stock-market observations.
40. Training the Classifier
The classifier is trained using:
clf.fit(
X_train,
y_train
)
It learns the relationship between:
FEATURES
↓
TARGET
where TARGET represents the future 5-day outcome condition.
41. Training the Return Predictor
The second model learns:
FEATURES
↓
FUTURE_RETURN_5D
This is the regression model.
So the two models have different jobs.
| Model | Job |
|---|---|
| Classifier | Probability of positive outcome |
| Regressor | Expected 5-day return |
42. Measuring Model Performance
The program calculates:
Accuracy
How often the predicted class matched the actual class.
Precision
When the model predicts a positive signal, how often was it correct?
AUC
A measure of how well the model distinguishes between the two classes across probability thresholds.
The program stores these results in:
performance_df
43. Why Accuracy Alone Is Not Enough
Suppose a model has:
Accuracy = 80%
That sounds excellent.
But accuracy alone can be misleading.
For example, if only 20% of observations are positive and the model predicts "negative" almost all the time, it could still obtain high accuracy.
Therefore the notebook also calculates:
Precision
AUC
This is a good practice.
44. Cell 11 – Generating Today's Signals
This is where everything becomes interesting.
The program takes the latest available data for every trained stock.
It creates:
X_latest
This is essentially:
Today's feature information for the stock.
The AI then receives this information.
45. AI Probability
The classifier produces:
probability
For example:
0.78
The program converts this to:
78%
This can be interpreted as the model's estimated probability for the positive class under the model's definition.
It is not the same thing as a guaranteed 78% chance of making money.
46. Expected 5-Day Return
The regression model produces:
expected_return
Suppose the output is:
0.035
That corresponds to:
3.5%
The program therefore has two important outputs:
AI Probability = 78%
Expected 5-Day Return = 3.5%
47. Entry Price
The program uses the current/latest price as:
Entry = price
For example:
Current price = ₹1,250
Then:
Entry = ₹1,250
This is a model-generated reference level, not a guarantee that the actual execution price will be the same.
48. Stop Loss Calculation
The program uses:
stop_loss = price - 1.5 * atr
So the stop is:
1.5 ATR below the current price.
Example:
Price = ₹1,000
ATR = ₹20
Then:
Stop Loss
= 1000 - 1.5 × 20
= ₹970
49. Target Calculation
The program uses:
target = price + 2.5 * atr
So:
Price = ₹1,000
ATR = ₹20
Then:
Target
= 1000 + 2.5 × 20
= ₹1,050
However, the program also says:
target = max(target, price * 1.05)
This means the target cannot be below approximately:
5% above the current price.
50. Trend Filter
The scanner requires several conditions for a strong trend.
The program checks:
Close > SMA50
SMA50 > SMA200
RSI > 50
ADX > 18
In simple language:
The stock should be above its 50-day average, the 50-day average should be above its 200-day average, momentum should be positive, and the trend should have reasonable strength.
51. Volume Filter
The scanner also checks:
VOLUME_RATIO >= 1.0
That means:
Current volume should be at least around its 20-period average.
This attempts to avoid situations where the price is moving without adequate trading activity.
52. How Does the Program Decide STRONG BUY?
The program requires:
AI Probability >= 70%
Expected 5-day return >= 2%
Trend conditions satisfied
Volume condition satisfied
If all these conditions are met:
STRONG BUY
53. How Does It Decide BUY?
The next level is:
AI Probability >= 60%
Expected 5-day return > 1%
Trend condition satisfied
Then:
BUY
54. When Does It Say SELL?
If:
AI Probability < 40%
the program labels the stock:
SELL
Otherwise, if none of the conditions are strong enough:
NO TRADE
55. Understanding the Four Signals
The final signal categories are:
| Signal | Meaning |
|---|---|
| STRONG BUY | Strongest combination according to the programmed rules |
| BUY | Positive setup but doesn't meet all STRONG BUY requirements |
| SELL | AI probability below 40% |
| NO TRADE | Conditions are not strong enough |
These labels are generated by the code's rules.
They should not be interpreted as guaranteed trading recommendations.
56. Cell 13 – Showing Buy Candidates
The program extracts:
STRONG BUY
BUY
from the complete signal table.
It displays:
Symbol
Price
AI Probability
Expected 5D Return
RSI
ADX
Volume Ratio
Entry
Stop Loss
Target
Signal
This is one of the most useful tables for a trader.
57. AI Score
The notebook then calculates:
AI_SCORE
This creates a ranking system.
The score combines:
AI Probability
+
Expected Return
+
Trend Score
+
ADX
+
Volume Ratio
with different weights.
The largest weight in the formula is given to:
AI Probability
followed by:
Expected Return
Trend Score
ADX
Volume
The purpose is to rank the stocks rather than simply classify them.
58. Why Do We Need an AI Score?
Suppose the scanner finds:
Stock A → BUY
Stock B → BUY
Stock C → STRONG BUY
Stock D → STRONG BUY
You may still want to know:
Which stock has the strongest overall score?
The AI Score attempts to answer that.
The program sorts stocks:
sort_values(
"AI_SCORE",
ascending=False
)
Therefore the highest-scoring stocks appear first.
59. Cell 15 – Backtesting
Now comes another important section.
The program tests the strategy on historical data.
It takes the 20% test portion and calculates:
AI Probability
Expected Return
Signal
Future 5-day return
The historical signal rule is:
Probability >= 70%
AND
Expected return >= 2%
AND
Close > SMA50
AND
SMA50 > SMA200
When these conditions are satisfied, the strategy takes the future 5-day return as the strategy return.
60. What Is a Backtest?
Suppose the program finds that a historical signal occurred on:
1 March
It then looks at what happened over the following five trading observations.
For example:
Entry = ₹1,000
Price after 5 trading days = ₹1,035
Return = +3.5%
The program records that as a successful historical outcome.
Repeat this across many historical signals and you can evaluate the strategy.
61. Backtest Metrics
The program calculates:
Number of Trades
How many historical signals occurred.
Win Rate
Percentage of trades with positive returns.
Profit Factor
A comparison between gross profits and gross losses.
Total Return
Compounded result of the recorded strategy returns.
Maximum Drawdown
Largest decline in the simulated equity curve.
62. Example
Suppose historical trades were:
+4%
+3%
-2%
+5%
-1%
The program can calculate:
Number of trades = 5
Winning trades = 3
Win rate = 60%
It can also calculate profit factor and cumulative performance.
63. Maximum Drawdown
Suppose your simulated account goes:
₹1,00,000
₹1,10,000
₹1,20,000
₹1,05,000
The peak was:
₹1,20,000
and it subsequently declined to:
₹1,05,000
That decline is important because it tells us about the risk experienced by the strategy.
A strategy producing high returns with very large drawdowns may be difficult to follow in real life.
64. Cell 18 – Building a Portfolio
The notebook goes one step further.
Instead of considering only individual stocks, it creates a portfolio approach.
For each date, it identifies stocks generating signals.
Then it sorts them by:
AI Probability
and selects up to:
Top 5
stocks.
The portfolio return is calculated using the average future return of those selected stocks.
65. Portfolio Equity Curve
The program calculates:
Equity
Peak
Drawdown
and then plots the equity curve.
The equity curve helps answer:
How would the simulated portfolio have grown or declined over time?
The chart is titled:
AI NIFTY 50 Strategy Equity Curve
66. Feature Importance
The notebook also examines:
model.feature_importances_
This produces a ranking of features that the XGBoost classifier considered useful in making its predictions.
The program displays the top 20 features.
For example, depending on the trained model, important features might include things such as:
RSI
Volume Ratio
NIFTY Trend
Distance from SMA
ADX
Momentum
The exact ranking should be taken from the actual model output rather than assumed in advance.
67. Why Is Feature Importance Useful?
Imagine the AI says:
BUY
You may want to know:
What information was useful to the model?
Feature importance can provide a clue.
However, feature importance should not be interpreted as:
"This indicator caused the stock to rise."
It tells us about the model's use of the feature, not a guaranteed causal relationship.
68. Final Scan
Cell 22 creates:
final_scan
This becomes the main final table.
It contains:
Symbol
Price
AI Probability
Expected 5D Return
RSI
ADX
Volume Ratio
Trend Score
Entry
Stop Loss
Target
AI Score
Signal
This is essentially the scanner's final dashboard.
69. STRONG BUY List
Cell 23 filters the final table:
final_scan["Signal"] == "STRONG BUY"
and displays all stocks that satisfy the strong-buy conditions.
It also prints:
Number of STRONG BUY signals
This provides a quick summary.
70. Exporting the Results to Excel
The final cell creates:
NIFTY50_AI_Scanner.xlsx
The Excel workbook contains three sheets.
Sheet 1 – AI Signals
Contains the current scan.
Sheet 2 – Model Performance
Contains:
Accuracy
Precision
AUC
for the individual stock models.
Sheet 3 – Backtest
Contains:
Trades
Win Rate
Profit Factor
Total Return
Maximum Drawdown
This makes the project much easier to analyse outside Colab.
71. Complete Program in Simple English
The entire notebook can be understood as:
1. Install Python packages
2. Import packages
3. Create NIFTY 50 stock list
4. Download historical prices
5. Calculate technical indicators
6. Calculate market information
7. Calculate India VIX information
8. Create 5-day future-return target
9. Train an AI classifier
10. Train an AI return predictor
11. Test the models
12. Calculate today's probability
13. Calculate expected 5-day return
14. Check trend
15. Check volume
16. Generate BUY/SELL/NO TRADE
17. Calculate stop loss
18. Calculate target
19. Calculate AI score
20. Rank stocks
21. Backtest the strategy
22. Calculate portfolio performance
23. Plot equity curve
24. Export everything to Excel
72. What Makes This Different from a Normal Stock Screener?
A traditional screener might say:
RSI > 50
AND
Price > SMA50
AND
Volume > Average Volume
This program goes further.
It uses:
Technical Indicators
+
Market Regime
+
India VIX
+
Machine Learning
+
Expected Return
+
Trend Filters
+
Volume Filter
+
Risk Levels
+
Backtesting
Therefore it is better described as an:
AI-Assisted Quantitative Stock Research System
rather than simply a technical scanner.
73. Important: AI Does Not Mean Guaranteed Prediction
This is perhaps the most important lesson for a beginner.
Suppose the program says:
AI Probability = 75%
It does not mean:
"There is a guaranteed 75% chance that I will make money."
It means that the trained model assigns a probability of 75% to its defined positive class based on the historical relationships it learned.
Real markets can behave differently.
74. Important Limitation #1 – Historical Data
The AI learns from historical data.
Markets evolve.
A relationship that worked between:
2015–2022
may behave differently later.
Therefore:
Past performance does not guarantee future performance.
75. Important Limitation #2 – No Transaction Costs in the Backtest
The current backtest logic does not explicitly subtract:
Brokerage
STT
Exchange charges
GST
Slippage
Bid-ask spread
Other trading costs
Therefore the backtested return can be more optimistic than actual trading performance.
A production version should incorporate realistic costs.
76. Important Limitation #3 – Five-Day Trades Can Overlap
The strategy uses:
Future 5-day return
for each signal.
If signals occur on consecutive days, their five-day holding periods can overlap.
The current backtest does not explicitly maintain a real position ledger showing:
Entry date
Exit date
Capital allocated
Open positions
Position overlap
Therefore the backtest should be treated as a research approximation, not a broker-ready trading simulator.
77. Important Limitation #4 – Portfolio Backtest Simplification
The portfolio section averages the future returns of up to five selected stocks.
A real portfolio needs to consider:
Capital
Position size
Entry price
Exit date
Position overlap
Transaction costs
Slippage
Cash
Rebalancing
Maximum exposure
The current code simplifies these factors.
78. Important Limitation #5 – Model Probability
The classifier's probability is generated by XGBoost.
A probability such as:
72%
should not automatically be assumed to be perfectly calibrated.
For serious quantitative research, probability calibration should also be investigated.
79. Important Limitation #6 – The AI Score Is a Custom Formula
The:
AI_SCORE
is not produced directly by XGBoost.
It is a manually designed formula combining:
AI probability
Expected return
Trend score
ADX
Volume
with selected weights.
Therefore:
AI Probability and AI Score are two different things.
The first comes from the classifier.
The second is a ranking formula created by the programmer.
80. Important Improvement – Use Decimal Prices
The uploaded code converts OHLC values to integers:
df["Close"] = df["Close"].astype(int)
This removes decimal precision.
For example:
₹1,234.75
becomes approximately:
₹1,234
For research, it would generally be preferable to retain the original floating-point price values.
This is a technical improvement worth making.
81. Important Improvement – Verify the NIFTY 50 List
The notebook contains a manually maintained list.
NIFTY 50 constituents change over time.
Therefore, a production-quality scanner should periodically update its universe rather than assuming that a hard-coded list remains current forever.
82. How a Complete Beginner Should Run It
Follow these steps.
Step 1
Open the notebook in Google Colab.
Step 2
Connect the runtime.
Step 3
Run Cell 1.
Step 4
Run Cell 2.
Step 5
Check the NIFTY 50 list.
Step 6
Fix the missing comma between:
POWERGRID.NS
RELIANCE.NS
Step 7
Run the data-download cell.
Step 8
Wait for the historical data to download.
Step 9
Run the feature-generation cells.
Step 10
Wait for the AI models to train.
Step 11
Look at model performance.
Step 12
Look at:
signals_df
Step 13
Look at:
final_scan
Step 14
Review:
STRONG BUY
Step 15
Review the backtest.
Step 16
Download:
NIFTY50_AI_Scanner.xlsx
83. How to Read the Final Table
Suppose your final table looks like:
| Symbol | Price | AI Probability | Expected 5D Return | RSI | ADX | Volume | Stop Loss | Target | Signal |
|---|---|---|---|---|---|---|---|---|---|
| ABC | ₹1,000 | 78% | 3.4% | 62 | 25 | 1.4 | ₹970 | ₹1,050 | STRONG BUY |
| XYZ | ₹800 | 64% | 1.8% | 57 | 22 | 1.2 | ₹770 | ₹840 | BUY |
Don't just look at the word:
STRONG BUY
Look at the entire row.
Ask:
Is AI probability high?
Is expected return attractive?
Is the trend strong?
Is volume supportive?
Is the stop loss reasonable?
Is the target realistic?
What did the historical backtest show?
How many trades generated the backtest?
What was maximum drawdown?
84. A Better Way to Use the Scanner
Don't use this program as:
"The computer said BUY, therefore I will buy."
Instead use it as:
"The computer found a candidate. Now I will investigate it."
A better workflow is:
AI Scanner
↓
Shortlist
↓
Check Chart
↓
Check Market Trend
↓
Check Fundamental Context
↓
Check Risk/Reward
↓
Check Position Size
↓
Make Decision
85. Beginner Experiment #1
Change nothing.
Run the complete notebook.
Record:
Number of STRONG BUY signals
Number of BUY signals
Number of SELL signals
Top 10 AI Score stocks
This establishes your baseline.
86. Beginner Experiment #2
Change the probability threshold.
Currently STRONG BUY requires:
70%
Experiment with:
65%
70%
75%
80%
Then compare:
Number of trades
Win rate
Profit factor
Total return
Maximum drawdown
This teaches you how changing a threshold changes a strategy.
87. Beginner Experiment #3
Compare Stocks
Run the scanner and identify five stocks.
Then compare:
AI Probability
Expected Return
RSI
ADX
Volume Ratio
Trend Score
Backtest
Ask:
Does the highest AI probability always produce the highest historical return?
This is an excellent quantitative-research question.
88. Beginner Experiment #4 – Build Your Own Ranking
Create your own ranking rules.
For example:
AI Probability > 65%
RSI > 50
ADX > 20
Volume Ratio > 1
Price > SMA50
SMA50 > SMA200
Then compare the results with the original scanner.
You have now started modifying an AI trading system.
89. What You Can Add in the Future
This project can become much more powerful.
Possible upgrades include:
Risk Management
Position sizing
Maximum capital per stock
Maximum portfolio risk
ATR-based sizing
Better Backtesting
Transaction costs
Slippage
Holding-period simulation
Portfolio capital tracking
Overlapping trade handling
Advanced Validation
Walk-forward testing
Out-of-sample testing
Cross-validation for time series
Probability calibration
More Market Data
Sector indices
Market breadth
FII/DII activity
Advance/decline ratio
Sector relative strength
Options
Option chain
Open interest
PCR
IV
IV Rank
Max Pain
AI
Random Forest
LightGBM
Logistic Regression
Neural Networks
Ensemble models
Regime classification
90. The Ideal Future Version
A more advanced version of this project could eventually look like:
NIFTY 50
↓
Market Regime
↓
Stock Feature Engine
↓
┌─────────┴─────────┐
↓ ↓
Classification Regression
↓ ↓
Probability Expected Return
└─────────┬─────────┘
↓
Trend Filter
↓
Volume Filter
↓
Risk Engine
↓
Position Sizing
↓
AI Ranking
↓
Portfolio Builder
↓
Backtest Engine
↓
Trading Dashboard
That would transform the notebook into a much more complete quantitative research platform.
91. Final Takeaway
The uploaded program is much more than a simple stock screener.
It combines:
Python
+
Historical Market Data
+
Technical Analysis
+
NIFTY Market Data
+
India VIX
+
Machine Learning
+
Return Prediction
+
Trend Filtering
+
Volume Analysis
+
Risk Levels
+
Backtesting
+
Portfolio Analysis
The central idea is:
Use historical market information to train an AI model, use the model to estimate the probability and expected return of a future 5-day move, then combine that information with technical filters and risk rules to rank NIFTY 50 stocks.
The most important thing for a beginner to remember is:
AI Prediction
≠
Guaranteed Profit
The scanner is best used as a research and decision-support tool.
The real objective is not to find a magical BUY signal.
It is to build a systematic process:
Research → Test → Validate → Manage Risk → Improve → Repeat
92. Quick Reference – What Each Cell Does
| Cell | Purpose |
|---|---|
| 1 | Install Python packages |
| 2 | Import libraries |
| 3 | Define NIFTY 50 universe |
| 4 | Download historical data |
| 5 | Create technical and market features |
| 6 | Create 5-day prediction target |
| 7 | Define AI features |
| 8 | Build datasets |
| 9 | Train and test AI models |
| 10 | Compare model performance |
| 11 | Generate current AI signals |
| 12 | Validate feature data |
| 13 | Show BUY candidates |
| 14 | Calculate AI Score |
| 15 | Backtest individual stocks |
| 16 | Run backtests for all stocks |
| 17 | Calculate backtest statistics |
| 18 | Build top-5 portfolio simulation |
| 19 | Calculate portfolio return/drawdown |
| 20 | Plot equity curve |
| 21 | Display feature importance |
| 22 | Create final scanner |
| 23 | Show STRONG BUY stocks |
| 24 | Export results to Excel |
Final Warning
Before using this scanner with real money, the following should be improved:
Fix the POWERGRID/RELIANCE ticker-list issue
Keep decimal price precision
Update the NIFTY 50 universe automatically
Add brokerage and slippage
Build a proper trade/position-based backtester
Handle overlapping 5-day positions correctly
Perform walk-forward/out-of-sample testing
Validate probability calibration
Test across different market regimes
Add proper position sizing and portfolio risk controls
Only after these improvements should the system be considered for serious paper-trading evaluation.
This article is for educational and research purposes. Machine-learning predictions and historical backtests do not guarantee future market performance.
if you have idea, i can automate for you. Feel free to connect me.
ReplyDelete