Skip to content

Module 01 · Lesson 4 of 5 (01.4)

Returns, moving averages, and a first signal in pandas

Beginner18 minDraft — under review

Compute simple and log returns, measure volatility, build moving averages with rolling windows, and turn them into a signal that only acts on information available at the time.

Before this lesson: pandas: price data as a table

In this lesson you will

  • Calculate simple and log returns from a price series and explain when each is useful.
  • Measure daily volatility and scale it to an annual figure.
  • Build moving averages with rolling windows and handle the missing values they create.
  • Turn a condition into a signal and shift it so that trades use only past information.
  • Show with code how using same-bar information inflates results.

Prices on their own are hard to compare. A move from 100 to 101 and a move from 10 to 11 are both one point, but very different changes. Almost all quantitative work therefore starts by converting prices into returns. This lesson covers returns and volatility, then moving averages, then the step that turns a calculation into a trading decision: a signal, timed correctly.

Simple and log returns

The simple return from one close to the next is the percentage change:

rt=PtPt−1−1r_t = \frac{P_t}{P_{t-1}} - 1

where PtP_t is today's close and Pt−1P_{t-1} is yesterday's. The log return uses the natural logarithm of the same ratio:

ℓt=ln⁡(PtPt−1)\ell_t = \ln\left(\frac{P_t}{P_{t-1}}\right)

For small moves the two are almost identical. The difference is in how they combine over time. Simple returns compound: to get a total return you multiply (1+rt)(1 + r_t) across days. Log returns add: the total log return is the sum of the daily ones. That is why the random walk in the previous lesson summed random numbers and then applied np.exp, which undoes a logarithm.

returns.pyPython
import numpy as np
import pandas as pd
 
 
def make_prices(n_days=260, seed=42):
    """Synthetic daily OHLCV bars from a random walk. Not real market data."""
    rng = np.random.default_rng(seed)
    dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
    close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
    prev_close = np.concatenate([[100.0], close[:-1]])
    open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
    high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
    low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
    volume = rng.integers(800_000, 1_500_000, n_days)
    return pd.DataFrame(
        {"open": open_, "high": high, "low": low, "close": close, "volume": volume},
        index=dates,
    )
 
 
prices = make_prices()
close = prices["close"]
 
simple = close.pct_change()                 # (today / yesterday) - 1
log_ret = np.log(close / close.shift(1))    # ln(today / yesterday)
 
table = pd.DataFrame({"close": close, "simple": simple, "log": log_ret})
print(table.head(4).round(4))
 
total = close.iloc[-1] / close.iloc[0] - 1
print(f"Total return, first to last close: {total:.2%}")
print(f"Sum of log returns: {log_ret.sum():.4f} -> {np.exp(log_ret.sum()) - 1:.2%}")
print(f"Daily volatility: {simple.std():.2%}  annualised: {simple.std() * np.sqrt(252):.1%}")
Output
               close  simple     log
date                                
2024-01-01  100.3964     NaN     NaN
2024-01-02   99.1811 -0.0121 -0.0122
2024-01-03  100.1083  0.0093  0.0093
2024-01-04  101.2750  0.0117  0.0116
Total return, first to last close: -7.16%
Sum of log returns: -0.0743 -> -7.16%
Daily volatility: 1.13%  annualised: 17.9%

.pct_change() computes simple returns in one step. close.shift(1) moves every value down one row, so each row sees the previous day's close next to its own; dividing the two and taking np.log gives log returns. pandas does arithmetic on whole columns at once and lines rows up by date, so there is no loop to write.

The first row is NaN, short for "not a number", which is how pandas marks a missing value. There is no return on the first day because there is no previous close. pandas skips NaN in sums and averages.

.iloc[...] selects by position rather than by label, so close.iloc[0] is the first close and close.iloc[-1] the last. The output confirms the key property: the sum of the log returns, converted back with np.exp(...) - 1, gives exactly the total simple return.

Volatility and annualisation

Volatility is usually measured as the standard deviation of returns: how widely daily returns spread around their average. .std() computes it. Here it is 1.13% per day, close to the 1.2% the generator used.

Daily figures are often converted to annual ones, called annualisation, so they can be compared across instruments and timeframes. If daily returns were independent, variance would grow in proportion to time, and volatility with its square root. With about 252 trading days a year:

σannual=σdaily×252\sigma_{\text{annual}} = \sigma_{\text{daily}} \times \sqrt{252}

That gives 17.9% here. This scaling is a convention resting on an assumption, independent returns, that real markets only roughly satisfy. Module 03 looks at when it breaks down.

Moving averages and a signal

A moving average is the average of the last n closes, recalculated each day. In pandas, .rolling(n) creates a window that slides along the column, and .mean() averages each window. Comparing a fast average with a slow one gives a simple trend condition: when the 20-day average is above the 50-day, recent prices are higher than older ones.

To make that condition into a signal, store it as 1 (long) or 0 (flat). Then comes the step that matters most: deciding when the signal can be acted on.

signal.pyPython
import numpy as np
import pandas as pd
 
 
def make_prices(n_days=260, seed=42):
    """Synthetic daily OHLCV bars from a random walk. Not real market data."""
    rng = np.random.default_rng(seed)
    dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
    close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
    prev_close = np.concatenate([[100.0], close[:-1]])
    open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
    high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
    low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
    volume = rng.integers(800_000, 1_500_000, n_days)
    return pd.DataFrame(
        {"open": open_, "high": high, "low": low, "close": close, "volume": volume},
        index=dates,
    )
 
 
prices = make_prices()
df = pd.DataFrame({"close": prices["close"]})
 
df["ma_fast"] = df["close"].rolling(20).mean()   # average of the last 20 closes
df["ma_slow"] = df["close"].rolling(50).mean()   # average of the last 50 closes
 
# Signal: 1 (long) when the fast average is above the slow one, else 0 (flat)
df["signal"] = (df["ma_fast"] > df["ma_slow"]).astype(int)
 
# Position: act on the NEXT bar, because today's signal needs today's close
df["position"] = df["signal"].shift(1).fillna(0).astype(int)
 
print(df.iloc[47:52].round(2))
print("Position changes:", int(df["position"].diff().abs().sum()))
print(f"Days in the market: {df['position'].mean():.0%}")
Output
             close  ma_fast  ma_slow  signal  position
date                                                  
2024-03-06  106.20   103.70      NaN       0         0
2024-03-07  107.10   104.01      NaN       0         0
2024-03-08  107.22   104.29   101.29       1         0
2024-03-11  107.63   104.47   101.44       1         1
2024-03-12  108.48   104.70   101.63       1         1
Position changes: 8
Days in the market: 28%

Assigning to df["ma_fast"] adds a new column. The rolling averages are NaN until their window is full: the 50-day average first appears on the 50th bar, 8 March. Comparisons with NaN are False, so the signal is 0 until then, which is the safe default. .astype(int) converts True/False into 1/0.

Now look at 8 March. The signal turns to 1 that day, but the position stays 0 and only becomes 1 on 11 March, the next bar. That is the effect of the highlighted line. The signal on 8 March uses 8 March's close, which is only known once the day is over. The earliest you could act on it is the next session. .shift(1) moves each signal one row later, and .fillna(0) replaces the NaN the shift leaves in the first row with "flat".

The last two lines summarise the position. .diff() takes the change from the previous row, so each entry or exit shows up as +1 or −1; taking the absolute value and summing counts the changes. The average of a 0/1 column is the share of days spent long.

Why the shift matters: a demonstration

Forgetting that shift is one of the most common errors in backtesting, and its effect can be enormous. This example uses a deliberately naive rule, "be long when today closed higher than yesterday", and computes its result twice: once trading on the next bar, and once "trading" on the same bar that produced the signal.

lookahead.pyPython
import numpy as np
import pandas as pd
 
 
def make_prices(n_days=260, seed=42):
    """Synthetic daily OHLCV bars from a random walk. Not real market data."""
    rng = np.random.default_rng(seed)
    dates = pd.bdate_range("2024-01-01", periods=n_days, name="date")
    close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.012, n_days)))
    prev_close = np.concatenate([[100.0], close[:-1]])
    open_ = prev_close * (1 + rng.normal(0, 0.002, n_days))
    high = np.maximum(open_, close) * (1 + np.abs(rng.normal(0, 0.004, n_days)))
    low = np.minimum(open_, close) * (1 - np.abs(rng.normal(0, 0.004, n_days)))
    volume = rng.integers(800_000, 1_500_000, n_days)
    return pd.DataFrame(
        {"open": open_, "high": high, "low": low, "close": close, "volume": volume},
        index=dates,
    )
 
 
prices = make_prices()
df = pd.DataFrame({"close": prices["close"]})
df["ret"] = df["close"].pct_change().fillna(0)
 
# A deliberately naive rule: be long whenever today closed higher than yesterday.
df["signal"] = (df["ret"] > 0).astype(int)
 
honest = df["signal"].shift(1).fillna(0) * df["ret"]   # trade on the next bar
leaky = df["signal"] * df["ret"]                       # "trade" on the bar that created the signal
 
print(f"Honest version:  {(1 + honest).prod() - 1:+.1%}")
print(f"Leaky version:   {(1 + leaky).prod() - 1:+.1%}")
Output
Honest version:  +24.3%
Leaky version:   +208.3%

Multiplying the position by the day's return gives the strategy's return for that day: the full return when long, zero when flat. .prod() multiplies the (1+r)(1 + r) values together to compound them.

The leaky version is long only on days that closed up, because it decides after seeing the close. It captures every up day and avoids every down day, which no real trader can do. That is look-ahead bias in its plainest form.

The honest number deserves just as much suspicion. This is a random walk, so there is nothing to find, and +24.3% is luck: a different seed would give a different number, quite possibly negative. A single test on a single path proves nothing, which is why Modules 05 and 06 exist.

What you have learned

You can now turn prices into returns, measure their volatility, build rolling indicators, and convert a condition into a correctly timed position. The final lesson of this module puts these pieces together into an equity curve and compares it with simply holding.

Key takeaways

  • Simple returns are the percentage change from one close to the next; log returns add up over time, which makes them convenient for calculations.
  • Volatility is the standard deviation of returns; daily volatility is often scaled by the square root of 252 to give an annual figure.
  • A rolling window has no value until it is full. The first rows of a 50-day average are missing, and that is correct.
  • A signal computed from today's close can only be acted on tomorrow. Shifting the signal by one bar is the simplest defence against look-ahead bias.

Lab exercise

Compare signals with and without the shift

Using make_prices with three different seeds, build the 20/50 moving-average signal and compute the strategy's total return twice, once with the position shifted by one bar and once without. Record the six numbers in a table. Then explain in two sentences why the unshifted version is not a result you could ever have achieved.

Lab coming soon

Until the Lab opens, run the exercise in your own environment from Module 01.

Self-check

Answer in your own words first, then reveal the answer.

  1. 01A price goes from 100 to 110 and back to 100. What are the two simple returns, and what do the two log returns add up to?

    Show answer

    The simple returns are +10% and about −9.09%, which do not add to zero even though the price ended where it started. The log returns are ln(1.1) ≈ +0.0953 and ln(100/110) ≈ −0.0953, which add to exactly zero, matching the total change.

  2. 02Why are the first 49 values of a 50-day rolling mean missing?

    Show answer

    A 50-day average needs 50 closes. Until the 50th bar, there are not enough values, so pandas returns NaN rather than an average of fewer days. Filling those values with something else would invent data.

  3. 03What does df["signal"].shift(1) do, and why does the strategy use it?

    Show answer

    It moves every signal one row later, so the position on each day is the signal from the previous day's close. The signal needs the day's close to be calculated, so it cannot be traded until the next bar. Without the shift, the test trades on information it would not have had.

Finished the lesson, the Lab exercise and the self-check?

Glossary terms in this lesson

Related Academy articles

Educational material only. Examples use synthetic data and are not investment advice or a forecast of results.