🤖 Machine Learning Regression
Forecast Stock Returns with Python
Can Machine Learning predict tomorrow's stock return?
Traditional technical analysis uses indicators such as Moving Averages, RSI, MACD, ADX and Bollinger Bands to generate trading signals.
Machine Learning gives us another approach.
Instead of saying:
"Buy when RSI is below 30."
we can ask:
"Based on today's market conditions, what return might the stock generate tomorrow?"
This is where Machine Learning Regression becomes useful.
In this article, we will build a simple regression model in Python that uses historical market data and technical features to forecast the next trading day's return.
⚠️ Important: This is an educational example, not a guaranteed prediction system or financial advice. Financial markets are noisy, and historical relationships can disappear.
1. What Is Regression?
Regression is a supervised Machine Learning technique used to predict a continuous numerical value.
For example:
| Problem | Prediction |
|---|---|
| House price prediction | ₹75 lakh |
| Temperature prediction | 32.5°C |
| Sales prediction | ₹1,25,000 |
| Stock return prediction | 0.85% |
For trading, our target could be:
Tomorrow's percentage return.
For example:
Today's Close = ₹1,000
Tomorrow's Close = ₹1,015
Return = (1015 - 1000) / 1000
= 0.015
= 1.5%
Our ML model attempts to learn relationships between today's market features and tomorrow's return.
2. Regression vs Classification
This distinction is important.
Classification
Classification predicts a category.
UP
DOWN
Example:
Will NIFTY go up tomorrow?
YES / NO
Regression
Regression predicts a numerical value.
Expected return = 0.72%
For example:
RSI = 61
MACD = 15.2
ATR = 220
Volume = 1.2 million
Predicted return = 0.68%
So:
Classification → What will happen?
Regression → How much might it happen?
3. How Can Regression Be Used in Trading?
A typical ML trading workflow looks like this:
Historical Market Data
↓
Feature Engineering
↓
Technical Indicators
↓
Create Target
↓
Train Regression Model
↓
Evaluate Model
↓
Predict Future Return
↓
Trading Decision
For example:
Predicted Return > +0.50%
↓
BUY
Predicted Return between -0.50% and +0.50%
↓
HOLD
Predicted Return < -0.50%
↓
SELL / AVOID
The thresholds are only examples. They must be tested rather than assumed to work.
4. What Data Do We Need?
For this example, we will use historical OHLCV data.
OHLCV
O = Open
H = High
L = Low
C = Close
V = Volume
We can obtain historical data using yfinance.
Install the required libraries:
!pip install yfinance scikit-learn pandas numpy matplotlib
Import them:
import yfinance as yf
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
5. Download Historical Data
Let's use Apple as an example.
data = yf.download(
"AAPL",
start="2018-01-01",
end="2026-01-01",
auto_adjust=True
)
data.head()
You should see columns similar to:
Open
High
Low
Close
Volume
If you want to experiment with an Indian stock, you can replace the ticker with a Yahoo Finance symbol such as:
data = yf.download(
"RELIANCE.NS",
start="2018-01-01",
end="2026-01-01",
auto_adjust=True
)
6. Calculate Daily Returns
The first feature we will create is the daily return.
data["Return"] = data["Close"].pct_change()
For example:
Day 1 ₹100
Day 2 ₹102
Return:
(102 - 100) / 100 = 0.02
Therefore:
Return = 2%
7. Create Technical Features
Machine Learning models work with numerical features.
Let's create some commonly used indicators.
Moving Averages
data["SMA_10"] = data["Close"].rolling(10).mean()
data["SMA_20"] = data["Close"].rolling(20).mean()
These provide information about the recent price trend.
RSI
We can calculate a simple RSI:
delta = data["Close"].diff()
gain = delta.clip(lower=0)
loss = -delta.clip(upper=0)
avg_gain = gain.rolling(14).mean()
avg_loss = loss.rolling(14).mean()
rs = avg_gain / avg_loss
data["RSI"] = 100 - (100 / (1 + rs))
RSI provides information about recent price momentum.
Volatility
Let's calculate 20-day rolling volatility.
data["Volatility"] = data["Return"].rolling(20).std()
Higher volatility means returns have been fluctuating more strongly.
Price Momentum
We can also calculate 5-day momentum:
data["Momentum_5"] = data["Close"].pct_change(5)
And 10-day momentum:
data["Momentum_10"] = data["Close"].pct_change(10)
8. Create the Target Variable
This is the most important step.
We don't want to predict today's return.
We want to predict tomorrow's return.
Therefore:
data["Target"] = data["Close"].pct_change().shift(-1)
Conceptually:
Today's features
↓
Tomorrow's return
For example:
Date RSI Momentum Target
---------------------------------------
Monday 55.2 0.012 0.008
Tuesday 61.4 0.009 -0.004
Wednesday 48.7 -0.003 0.011
The model learns:
Today's information → Tomorrow's return
9. Select Features
Let's define our feature set.
features = [
"Return",
"SMA_10",
"SMA_20",
"RSI",
"Volatility",
"Momentum_5",
"Momentum_10"
]
Our target is:
target = "Target"
Remove rows containing missing values:
df = data[features + [target]].dropna()
10. Why We Must Be Careful With Time-Series Data
This is one of the most important concepts in Machine Learning for trading.
Normally, we might randomly split data:
80% Training
20% Testing
But randomly shuffling financial time-series data can introduce look-ahead bias.
Imagine:
2018 ─────── 2022 ─────── 2026
↑
Training
↑
Testing
The model should learn from the past and then be tested on the future.
Therefore, we will use a chronological split.
11. Train/Test Split
split = int(len(df) * 0.80)
train = df.iloc[:split]
test = df.iloc[split:]
Create X and y:
X_train = train[features]
y_train = train[target]
X_test = test[features]
y_test = test[target]
Notice that we have not shuffled the data.
This preserves the time sequence.
12. Train a Linear Regression Model
Let's start with a simple model.
model = LinearRegression()
model.fit(X_train, y_train)
That's it!
Our first regression model has been trained.
Conceptually, the model is trying to learn something like:
Predicted Return =
intercept
+ coefficient₁ × Return
+ coefficient₂ × SMA_10
+ coefficient₃ × SMA_20
+ coefficient₄ × RSI
+ coefficient₅ × Volatility
+ ...
The coefficients are learned from historical data.
13. Make Predictions
Now let's predict the returns for our unseen test data.
predictions = model.predict(X_test)
Create a DataFrame:
results = pd.DataFrame({
"Actual": y_test,
"Predicted": predictions
})
results.head(10)
You might get results such as:
Actual Predicted
---------------------
0.0082 0.0037
-0.0045 0.0012
0.0061 0.0048
-0.0090 -0.0021
0.0025 0.0017
Remember:
The model is forecasting a noisy numerical quantity.
It is not predicting the future with certainty.
14. Evaluate the Model
A regression model should not be judged simply by looking at a few predictions.
Let's calculate some standard metrics.
Mean Absolute Error
mae = mean_absolute_error(y_test, predictions)
print("MAE:", mae)
MAE tells us the average absolute prediction error.
Root Mean Squared Error
rmse = np.sqrt(mean_squared_error(y_test, predictions))
print("RMSE:", rmse)
RMSE penalizes larger errors more heavily.
R² Score
r2 = r2_score(y_test, predictions)
print("R² Score:", r2)
R² measures how much variation in the target is explained by the model under the assumptions of this metric.
Important
A low or even negative R² does not automatically mean the model is useless for trading.
Trading performance depends on:
direction
magnitude
transaction costs
position sizing
turnover
risk
drawdown
execution
A model with modest predictive accuracy can sometimes be useful, while a model with an impressive statistical score can still lose money after costs.
15. Visualize Actual vs Predicted Returns
Let's compare actual and predicted returns.
plt.figure(figsize=(12, 5))
plt.plot(y_test.values, label="Actual Return")
plt.plot(predictions, label="Predicted Return")
plt.title("Actual vs Predicted Returns")
plt.xlabel("Test Period")
plt.ylabel("Return")
plt.legend()
plt.show()
The closer the predictions are to the actual movements, the better the model is performing.
However, visual similarity alone is not enough.
16. Add a Trading Signal
Now comes the interesting part.
Suppose we use a simple rule:
Predicted return > +0.50%
→ BUY
Predicted return < -0.50%
→ SELL
Otherwise
→ HOLD
In Python:
results["Signal"] = np.where(
results["Predicted"] > 0.005,
1,
np.where(
results["Predicted"] < -0.005,
-1,
0
)
)
Here:
1 = BUY
0 = HOLD
-1 = SELL
17. Build a Simple Strategy Return
For a basic long/short illustration, we can multiply the signal by the next-day actual return:
results["Strategy_Return"] = (
results["Signal"] * results["Actual"]
)
Calculate cumulative returns:
results["Strategy_Cumulative"] = (
1 + results["Strategy_Return"]
).cumprod()
Calculate buy-and-hold cumulative return:
results["BuyHold_Cumulative"] = (
1 + results["Actual"]
).cumprod()
Plot both:
plt.figure(figsize=(12, 5))
plt.plot(
results["Strategy_Cumulative"],
label="ML Strategy"
)
plt.plot(
results["BuyHold_Cumulative"],
label="Buy & Hold"
)
plt.title("ML Regression Strategy vs Buy & Hold")
plt.xlabel("Test Period")
plt.ylabel("Growth of ₹1")
plt.legend()
plt.show()
This comparison is much more useful than simply asking whether the regression model has a high R².
18. Calculate Strategy Statistics
Let's calculate some basic metrics.
strategy_returns = results["Strategy_Return"]
total_return = (
results["Strategy_Cumulative"].iloc[-1] - 1
)
win_rate = (
(strategy_returns > 0).sum()
/ (strategy_returns != 0).sum()
)
Print them:
print("Total Strategy Return:", total_return)
print("Win Rate:", win_rate)
You can also calculate maximum drawdown.
equity = results["Strategy_Cumulative"]
peak = equity.cummax()
drawdown = (equity - peak) / peak
max_drawdown = drawdown.min()
print("Maximum Drawdown:", max_drawdown)
Now we have three important measurements:
Return
Win Rate
Maximum Drawdown
But a serious backtest should go much further.
19. Add Transaction Costs
This is where many beginner backtests become unrealistic.
Suppose our strategy changes position frequently.
Every trade can involve costs such as:
Brokerage
STT
Exchange transaction charges
GST
Stamp duty
Slippage
Bid-ask spread
For a simple demonstration, assume a transaction cost:
cost = 0.0005
We can identify signal changes:
results["Position_Change"] = (
results["Signal"].diff().abs()
)
Then subtract a simple cost:
results["Net_Return"] = (
results["Strategy_Return"]
- results["Position_Change"] * cost
)
And calculate cumulative net return:
results["Net_Cumulative"] = (
1 + results["Net_Return"]
).cumprod()
This is still a simplified cost model, but it demonstrates an important principle:
A strategy that looks profitable before costs may not be profitable after costs.
20. Predict the Next Trading Day
Once the model has been trained, we can use the latest available feature row.
latest_features = data[features].dropna().iloc[-1:]
next_return_prediction = model.predict(
latest_features
)[0]
print(
"Predicted next-day return:",
next_return_prediction
)
Convert it to percentage:
print(
f"Predicted return: {next_return_prediction * 100:.2f}%"
)
For example:
Predicted return: 0.64%
Again, this is a model estimate—not a guaranteed market outcome.
21. Generate a Simple Decision
threshold = 0.005
if next_return_prediction > threshold:
signal = "BUY"
elif next_return_prediction < -threshold:
signal = "SELL"
else:
signal = "HOLD"
print("ML Signal:", signal)
Possible output:
Predicted return: 0.64%
ML Signal: BUY
22. Complete Python Program
Here is the complete simplified example.
import yfinance as yf
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.metrics import (
mean_absolute_error,
mean_squared_error,
r2_score
)
# ---------------------------------------
# 1. Download Data
# ---------------------------------------
data = yf.download(
"AAPL",
start="2018-01-01",
end="2026-01-01",
auto_adjust=True
)
# ---------------------------------------
# 2. Daily Return
# ---------------------------------------
data["Return"] = data["Close"].pct_change()
# ---------------------------------------
# 3. Moving Averages
# ---------------------------------------
data["SMA_10"] = data["Close"].rolling(10).mean()
data["SMA_20"] = data["Close"].rolling(20).mean()
# ---------------------------------------
# 4. RSI
# ---------------------------------------
delta = data["Close"].diff()
gain = delta.clip(lower=0)
loss = -delta.clip(upper=0)
avg_gain = gain.rolling(14).mean()
avg_loss = loss.rolling(14).mean()
rs = avg_gain / avg_loss
data["RSI"] = 100 - (100 / (1 + rs))
# ---------------------------------------
# 5. Volatility
# ---------------------------------------
data["Volatility"] = (
data["Return"].rolling(20).std()
)
# ---------------------------------------
# 6. Momentum
# ---------------------------------------
data["Momentum_5"] = (
data["Close"].pct_change(5)
)
data["Momentum_10"] = (
data["Close"].pct_change(10)
)
# ---------------------------------------
# 7. Target = Next-Day Return
# ---------------------------------------
data["Target"] = (
data["Close"].pct_change().shift(-1)
)
# ---------------------------------------
# 8. Features
# ---------------------------------------
features = [
"Return",
"SMA_10",
"SMA_20",
"RSI",
"Volatility",
"Momentum_5",
"Momentum_10"
]
df = data[features + ["Target"]].dropna()
# ---------------------------------------
# 9. Time-Based Train/Test Split
# ---------------------------------------
split = int(len(df) * 0.80)
train = df.iloc[:split]
test = df.iloc[split:]
X_train = train[features]
y_train = train["Target"]
X_test = test[features]
y_test = test["Target"]
# ---------------------------------------
# 10. Train Model
# ---------------------------------------
model = LinearRegression()
model.fit(X_train, y_train)
# ---------------------------------------
# 11. Prediction
# ---------------------------------------
predictions = model.predict(X_test)
# ---------------------------------------
# 12. Evaluation
# ---------------------------------------
mae = mean_absolute_error(
y_test,
predictions
)
rmse = np.sqrt(
mean_squared_error(
y_test,
predictions
)
)
r2 = r2_score(
y_test,
predictions
)
print("MAE :", mae)
print("RMSE:", rmse)
print("R2 :", r2)
# ---------------------------------------
# 13. Results
# ---------------------------------------
results = pd.DataFrame({
"Actual": y_test.values,
"Predicted": predictions
}, index=y_test.index)
# ---------------------------------------
# 14. Trading Signal
# ---------------------------------------
threshold = 0.005
results["Signal"] = np.where(
results["Predicted"] > threshold,
1,
np.where(
results["Predicted"] < -threshold,
-1,
0
)
)
# ---------------------------------------
# 15. Strategy Return
# ---------------------------------------
results["Strategy_Return"] = (
results["Signal"] *
results["Actual"]
)
results["Equity"] = (
1 + results["Strategy_Return"]
).cumprod()
# ---------------------------------------
# 16. Buy & Hold
# ---------------------------------------
results["BuyHold"] = (
1 + results["Actual"]
).cumprod()
# ---------------------------------------
# 17. Maximum Drawdown
# ---------------------------------------
peak = results["Equity"].cummax()
drawdown = (
results["Equity"] - peak
) / peak
max_drawdown = drawdown.min()
print(
"Maximum Drawdown:",
max_drawdown
)
# ---------------------------------------
# 18. Plot
# ---------------------------------------
plt.figure(figsize=(12, 5))
plt.plot(
results["Equity"],
label="ML Strategy"
)
plt.plot(
results["BuyHold"],
label="Buy & Hold"
)
plt.title(
"Machine Learning Regression Strategy"
)
plt.xlabel("Date")
plt.ylabel("Growth of ₹1")
plt.legend()
plt.show()
23. Why Linear Regression Is Only the Beginning
Linear Regression assumes a relatively simple relationship between the input features and target.
Financial markets are rarely that simple.
The relationship may look more like:
RSI + Momentum + Volatility
↓
Market Regime
↓
Return
The relationship can also change over time.
Therefore, we can experiment with more powerful models.
Possible models
Linear Regression
↓
Ridge Regression
↓
Lasso Regression
↓
Random Forest
↓
Gradient Boosting
↓
XGBoost
↓
Neural Networks
However:
More complicated does not automatically mean more profitable.
A simple model with robust testing can be more useful than a complex model that is overfit.
24. Regression Features for a Better Trading Model
A more advanced system could include:
Price Features
Daily Return
5-Day Return
10-Day Return
20-Day Return
Gap %
High-Low %
Close-Open %
Trend Features
SMA 20
SMA 50
SMA 200
EMA 20
EMA 50
ADX
Supertrend
Momentum Features
RSI
MACD
ROC
Stochastic
Volatility Features
ATR
Historical Volatility
Bollinger Band Width
India VIX
Volume Features
Volume
Volume Change
OBV
Relative Volume
This creates a much richer feature set.
25. Predicting Returns vs Predicting Price
A common beginner mistake is trying to predict the exact future price.
For example:
Today's price = ₹1,000
Prediction = ₹1,037
Predicting returns can be more useful for a trading model:
Expected return = +3.7%
Returns also make it easier to compare different stocks.
For example:
Stock A → +1.2%
Stock B → +2.4%
Stock C → -0.8%
A model can rank opportunities based on expected return.
26. From Prediction to Stock Ranking
This leads to an interesting quantitative strategy.
Suppose our model produces:
| Stock | Predicted Return |
|---|---|
| Stock A | +1.40% |
| Stock B | +0.85% |
| Stock C | +0.62% |
| Stock D | +0.21% |
| Stock E | -0.55% |
Instead of trading every stock, we could investigate only the stocks with the strongest predicted returns.
For example:
Predicted Return > 0.50%
↓
Candidate Stocks
↓
Additional Filters
↓
Risk Management
↓
Trade
This becomes a Machine Learning Stock Ranking System.
27. A Better ML Trading Architecture
A production-style system could look like this:
MARKET DATA
↓
┌───────────────────┐
│ Feature Engineering│
└───────────────────┘
↓
┌──────────────────────┐
│ Technical Indicators │
└──────────────────────┘
↓
MACHINE LEARNING
↓
Predicted Return
↓
Signal Generation
↓
Risk Management Layer
↓
Position Size / SL
↓
Backtest Engine
↓
Paper Trading / Alerts
↓
Live Execution
Notice that Machine Learning is only one component.
A complete trading system also needs:
Data validation
Signal rules
Risk management
Position sizing
Execution logic
Transaction-cost modelling
Monitoring
Logging
Failure handling
28. Avoiding Look-Ahead Bias
This deserves special attention.
Suppose you calculate:
data["Future_Return"] = data["Close"].pct_change().shift(-1)
This is acceptable as a target for training.
But you must never accidentally use future information as a feature.
For example, this would be problematic:
Today's features
+
Tomorrow's closing price
↓
Predict tomorrow's return
The model would effectively be given the answer.
This is called look-ahead bias.
29. Avoiding Overfitting
Imagine a model produces:
Training accuracy: 95%
Testing accuracy: 51%
That is a warning sign.
The model may have memorized historical patterns rather than learning something that generalizes.
A robust workflow should use:
Training Data
↓
Validation / Walk-Forward Testing
↓
Out-of-Sample Testing
↓
Paper Trading
↓
Live Trading
And avoid repeatedly tuning a model against the same test set.
30. Walk-Forward Testing
A more realistic approach is walk-forward testing.
For example:
Train: 2018–2021
Test: 2022
Train: 2018–2022
Test: 2023
Train: 2018–2023
Test: 2024
Train: 2018–2024
Test: 2025
This simulates how a model would actually operate as new information arrives.
For serious trading research, walk-forward testing is much more informative than a single train/test split.
31. What Should We Measure?
Do not evaluate an ML trading model using only:
R²
MAE
RMSE
Also examine trading metrics such as:
| Metric | Why it matters |
|---|---|
| Total Return | Overall performance |
| CAGR | Annualized growth |
| Win Rate | Percentage of winning trades |
| Average Win | Typical winning trade |
| Average Loss | Typical losing trade |
| Profit Factor | Gross profit / gross loss |
| Maximum Drawdown | Worst decline |
| Sharpe Ratio | Return relative to volatility |
| Sortino Ratio | Downside-risk adjusted return |
| Expectancy | Average expected result per trade |
| Number of Trades | Strategy activity |
| Turnover | Trading frequency |
A model should be judged as a trading system, not just as a statistical model.
32. 🚀 Advanced Project: ML Return Predictor
Now let's turn the example into a complete project.
Project Goal
Build an ML system that predicts the next-day return of NIFTY 50 stocks.
Input
OHLCV
RSI
MACD
ADX
ATR
SMA 20
SMA 50
SMA 200
Momentum
Volatility
Volume
ML Models
Compare:
Linear Regression
Random Forest
Gradient Boosting
XGBoost
Output
Symbol
Current Price
Predicted Return
Signal
Confidence/Model Score
Example:
RELIANCE +1.25% BUY
ICICIBANK +0.94% BUY
INFY +0.72% BUY
TCS +0.31% HOLD
ITC -0.44% HOLD
Then apply:
ML Prediction
+
Technical Filter
+
Trend Filter
+
Risk Management
↓
Final Trading Signal
33. 🧠 Python Pebble Challenge
Modify the program to use:
SMA 50
SMA 200
ADX
MACD
ATR
RSI
Volume Ratio
Then predict the next-day return.
Create the following output:
------------------------------------
ML TRADING PREDICTION
------------------------------------
Symbol : RELIANCE
Current Price : ₹______
Predicted Return: ____%
Signal : BUY / HOLD / SELL
Model : Linear Regression
------------------------------------
34. 🔥 Master Challenge
Build a NIFTY 50 ML Return Prediction System.
Your program should:
Download NIFTY 50 stock data.
Calculate technical indicators.
Create next-day return as the target.
Train multiple regression models.
Compare MAE and RMSE.
Compare out-of-sample trading performance.
Include transaction costs.
Calculate maximum drawdown.
Rank stocks by predicted return.
Generate the TOP 5 candidates.
Apply a trend filter.
Apply a volatility filter.
Generate final BUY/HOLD/AVOID signals.
Plot the equity curve.
Perform walk-forward testing.
The final architecture could be:
NIFTY 50
↓
Market Data
↓
Feature Engineering
↓
┌─────────────────┐
│ ML Regression │
└─────────────────┘
↓
Predicted Returns
↓
Stock Ranking
↓
Technical Filters
↓
Risk Management
↓
Backtesting
↓
Paper Trading
↓
Live System
35. Final Takeaway
Machine Learning Regression provides a different way of thinking about trading.
Instead of creating a rigid rule such as:
IF RSI < 30
THEN BUY
we can allow a model to learn relationships between multiple market features:
Price
+
Momentum
+
Trend
+
Volatility
+
Volume
↓
Machine Learning
↓
Expected Return
But remember:
Prediction is not the same as profitability.
A useful ML trading system requires much more than a good prediction score.
It needs:
Good Data
+
Good Features
+
Proper Time-Series Validation
+
Robust Backtesting
+
Transaction Costs
+
Risk Management
+
Position Sizing
+
Execution Discipline
The ultimate objective is not:
"Can Machine Learning predict tomorrow?"
It is:
"Can a robust, out-of-sample tested prediction provide enough edge to build a risk-controlled trading system?"
That is the real challenge of Machine Learning for Trading. 🚀
code provided may not work properly. please comment if you want proper running code.
ReplyDelete