Returns, moving averages, and a first signal in pandas
Compute simple and log returns, measure volatility, build moving averages with rolling windows, and turn them into a signal that only acts on information available at the time.
In this lesson you will
- Calculate simple and log returns from a price series and explain when each is useful.
- Measure daily volatility and scale it to an annual figure.
- Build moving averages with rolling windows and handle the missing values they create.
- Turn a condition into a signal and shift it so that trades use only past information.
- Show with code how using same-bar information inflates results.
Prices on their own are hard to compare. A move from 100 to 101 and a move from 10 to 11 are both one point, but very different changes. Almost all quantitative work therefore starts by converting prices into returns. This lesson covers returns and volatility, then moving averages, then the step that turns a calculation into a trading decision: a signal, timed correctly.
Simple and log returns
The simple return from one close to the next is the percentage change:
where is today's close and is yesterday's. The log return uses the natural logarithm of the same ratio:
For small moves the two are almost identical. The difference is in how they combine over time. Simple returns compound: to get a total return you multiply across days. Log returns add: the total log return is the sum of the daily ones. That is why the random walk in the previous lesson summed random numbers and then applied np.exp, which undoes a logarithm.
import numpy as np
import pandas as pd
def make_prices(n_days=260, seed=42):
"""Synthetic daily OHLCV bars from a random walk. Not real market data."""
rng = np.random.default_rng(seed)
dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
prev_close = np.concatenate([[100.0], close[:-1]])
open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
volume = rng.integers(800_000, 1_500_000, n_days)
return pd.DataFrame(
{"open": open_, "high": high, "low": low, "close": close, "volume": volume},
index=dates,
)
prices = make_prices()
close = prices["close"]
simple = close.pct_change() # (today / yesterday) - 1
log_ret = np.log(close / close.shift(1)) # ln(today / yesterday)
table = pd.DataFrame({"close": close, "simple": simple, "log": log_ret})
print(table.head(4).round(4))
total = close.iloc[-1] / close.iloc[0] - 1
print(f"Total return, first to last close: {total:.2%}")
print(f"Sum of log returns: {log_ret.sum():.4f} -> {np.exp(log_ret.sum()) - 1:.2%}")
print(f"Daily volatility: {simple.std():.2%} annualised: {simple.std() * np.sqrt(252):.1%}") close simple log
date
2024-01-01 100.3964 NaN NaN
2024-01-02 99.1811 -0.0121 -0.0122
2024-01-03 100.1083 0.0093 0.0093
2024-01-04 101.2750 0.0117 0.0116
Total return, first to last close: -7.16%
Sum of log returns: -0.0743 -> -7.16%
Daily volatility: 1.13% annualised: 17.9%.pct_change() computes simple returns in one step. close.shift(1) moves every value down one row, so each row sees the previous day's close next to its own; dividing the two and taking np.log gives log returns. pandas does arithmetic on whole columns at once and lines rows up by date, so there is no loop to write.
The first row is NaN, short for "not a number", which is how pandas marks a missing value. There is no return on the first day because there is no previous close. pandas skips NaN in sums and averages.
.iloc[...] selects by position rather than by label, so close.iloc[0] is the first close and close.iloc[-1] the last. The output confirms the key property: the sum of the log returns, converted back with np.exp(...) - 1, gives exactly the total simple return.
Volatility and annualisation
Volatility is usually measured as the standard deviation of returns: how widely daily returns spread around their average. .std() computes it. Here it is 1.13% per day, close to the 1.2% the generator used.
Daily figures are often converted to annual ones, called annualisation, so they can be compared across instruments and timeframes. If daily returns were independent, variance would grow in proportion to time, and volatility with its square root. With about 252 trading days a year:
That gives 17.9% here. This scaling is a convention resting on an assumption, independent returns, that real markets only roughly satisfy. Module 03 looks at when it breaks down.
Moving averages and a signal
A moving average is the average of the last n closes, recalculated each day. In pandas, .rolling(n) creates a window that slides along the column, and .mean() averages each window. Comparing a fast average with a slow one gives a simple trend condition: when the 20-day average is above the 50-day, recent prices are higher than older ones.
To make that condition into a signal, store it as 1 (long) or 0 (flat). Then comes the step that matters most: deciding when the signal can be acted on.
import numpy as np
import pandas as pd
def make_prices(n_days=260, seed=42):
"""Synthetic daily OHLCV bars from a random walk. Not real market data."""
rng = np.random.default_rng(seed)
dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
prev_close = np.concatenate([[100.0], close[:-1]])
open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
volume = rng.integers(800_000, 1_500_000, n_days)
return pd.DataFrame(
{"open": open_, "high": high, "low": low, "close": close, "volume": volume},
index=dates,
)
prices = make_prices()
df = pd.DataFrame({"close": prices["close"]})
df["ma_fast"] = df["close"].rolling(20).mean() # average of the last 20 closes
df["ma_slow"] = df["close"].rolling(50).mean() # average of the last 50 closes
# Signal: 1 (long) when the fast average is above the slow one, else 0 (flat)
df["signal"] = (df["ma_fast"] > df["ma_slow"]).astype(int)
# Position: act on the NEXT bar, because today's signal needs today's close
df["position"] = df["signal"].shift(1).fillna(0).astype(int)
print(df.iloc[47:52].round(2))
print("Position changes:", int(df["position"].diff().abs().sum()))
print(f"Days in the market: {df['position'].mean():.0%}") close ma_fast ma_slow signal position
date
2024-03-06 106.20 103.70 NaN 0 0
2024-03-07 107.10 104.01 NaN 0 0
2024-03-08 107.22 104.29 101.29 1 0
2024-03-11 107.63 104.47 101.44 1 1
2024-03-12 108.48 104.70 101.63 1 1
Position changes: 8
Days in the market: 28%Assigning to df["ma_fast"] adds a new column. The rolling averages are NaN until their window is full: the 50-day average first appears on the 50th bar, 8 March. Comparisons with NaN are False, so the signal is 0 until then, which is the safe default. .astype(int) converts True/False into 1/0.
Now look at 8 March. The signal turns to 1 that day, but the position stays 0 and only becomes 1 on 11 March, the next bar. That is the effect of the highlighted line. The signal on 8 March uses 8 March's close, which is only known once the day is over. The earliest you could act on it is the next session. .shift(1) moves each signal one row later, and .fillna(0) replaces the NaN the shift leaves in the first row with "flat".
The last two lines summarise the position. .diff() takes the change from the previous row, so each entry or exit shows up as +1 or −1; taking the absolute value and summing counts the changes. The average of a 0/1 column is the share of days spent long.
Why the shift matters: a demonstration
Forgetting that shift is one of the most common errors in backtesting, and its effect can be enormous. This example uses a deliberately naive rule, "be long when today closed higher than yesterday", and computes its result twice: once trading on the next bar, and once "trading" on the same bar that produced the signal.
import numpy as np
import pandas as pd
def make_prices(n_days=260, seed=42):
"""Synthetic daily OHLCV bars from a random walk. Not real market data."""
rng = np.random.default_rng(seed)
dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
prev_close = np.concatenate([[100.0], close[:-1]])
open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
volume = rng.integers(800_000, 1_500_000, n_days)
return pd.DataFrame(
{"open": open_, "high": high, "low": low, "close": close, "volume": volume},
index=dates,
)
prices = make_prices()
df = pd.DataFrame({"close": prices["close"]})
df["ret"] = df["close"].pct_change().fillna(0)
# A deliberately naive rule: be long whenever today closed higher than yesterday.
df["signal"] = (df["ret"] > 0).astype(int)
honest = df["signal"].shift(1).fillna(0) * df["ret"] # trade on the next bar
leaky = df["signal"] * df["ret"] # "trade" on the bar that created the signal
print(f"Honest version: {(1 + honest).prod() - 1:+.1%}")
print(f"Leaky version: {(1 + leaky).prod() - 1:+.1%}")Honest version: +24.3%
Leaky version: +208.3%Multiplying the position by the day's return gives the strategy's return for that day: the full return when long, zero when flat. .prod() multiplies the values together to compound them.
The leaky version is long only on days that closed up, because it decides after seeing the close. It captures every up day and avoids every down day, which no real trader can do. That is look-ahead bias in its plainest form.
The honest number deserves just as much suspicion. This is a random walk, so there is nothing to find, and +24.3% is luck: a different seed would give a different number, quite possibly negative. A single test on a single path proves nothing, which is why Modules 05 and 06 exist.
What you have learned
You can now turn prices into returns, measure their volatility, build rolling indicators, and convert a condition into a correctly timed position. The final lesson of this module puts these pieces together into an equity curve and compares it with simply holding.
Key takeaways
- Simple returns are the percentage change from one close to the next; log returns add up over time, which makes them convenient for calculations.
- Volatility is the standard deviation of returns; daily volatility is often scaled by the square root of 252 to give an annual figure.
- A rolling window has no value until it is full. The first rows of a 50-day average are missing, and that is correct.
- A signal computed from today's close can only be acted on tomorrow. Shifting the signal by one bar is the simplest defence against look-ahead bias.
Compare signals with and without the shift
Using make_prices with three different seeds, build the 20/50 moving-average signal and compute the strategy's total return twice, once with the position shifted by one bar and once without. Record the six numbers in a table. Then explain in two sentences why the unshifted version is not a result you could ever have achieved.
Self-check
Answer in your own words first, then reveal the answer.
A price goes from 100 to 110 and back to 100. What are the two simple returns, and what do the two log returns add up to?
Show answerHide answer
The simple returns are +10% and about −9.09%, which do not add to zero even though the price ended where it started. The log returns are ln(1.1) ≈ +0.0953 and ln(100/110) ≈ −0.0953, which add to exactly zero, matching the total change.
Why are the first 49 values of a 50-day rolling mean missing?
Show answerHide answer
A 50-day average needs 50 closes. Until the 50th bar, there are not enough values, so pandas returns
NaNrather than an average of fewer days. Filling those values with something else would invent data.What does
df["signal"].shift(1)do, and why does the strategy use it?Show answerHide answer
It moves every signal one row later, so the position on each day is the signal from the previous day's close. The signal needs the day's close to be calculated, so it cannot be traded until the next bar. Without the shift, the test trades on information it would not have had.
