You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
Mean-Variance Portfolio Optimization:
Challenging the role of traditional
covariance estimation

ZAKARIA MARAKBI

Master of Science Thesis
Stockholm, Sweden 2016

Effektiv portföljförvaltning: en
utvärdering av metoder för
kovariansskattning

ZAKARIA MARAKBI

Examensarbete
Stockholm, Sverige 2016

Mean-Variance Portfolio Optimization:
Challenging the role of traditional covariance estimation
Zakaria Marakbi

Master of Science Thesis INDEK 2016:63
KTH Industrial Engineering and Management
Industrial Management
SE-100 44 STOCKHOLM

Effektiv portföljförvaltning:
en utvärdering av metoder för kovariansskattning
Zakaria Marakbi

Examensarbete INDEK 2016:63
KTH Industriell teknik och management
Industriell ekonomi och organisation
SE-100 44 STOCKHOLM

.
Master of Science Thesis INDEK 2016:63
ME211X
Mean-Variance Portfolio Optimization: Challenging the
role of traditional covariance estimation
Zakaria Marakbi
Approved:

Examiner:

Supervisor:

2016-06-06

Gustav Martinsson

Tomas Sörensson

Abstract
Ever since its introduction in 1952, the Mean-Variance (MV) portfolio selection theory has remained a centerpiece within the realm of efficient asset allocation. However, in scientific circles,
the theory has stirred controversy. A strand of criticism has emerged that points to the phenomenon that Mean-Variance Optimization suffers from the severe drawback of estimation errors
contained in the expected return vector and the covariance matrix, resulting in portfolios that
may significantly deviate from the true optimal portfolio.
While a substantial amount of effort has been devoted to estimating the expected return
vector in this context, much less is written about the covariance matrix input. In recent times,
however, research that points to the importance of the covariance matrix in MV optimization
has emerged. As a result, there has been a growing interest whether MV optimization can be
enhanced by improving the estimate of the covariance matrix.
Hence, this thesis was set forth by the purpose to investigate whether financial practitioners and institutions can allocate portfolios consisting of assets in a more efficient manner by
changing the covariance matrix input in mean-variance optimization. In the quest of achieving
this purpose, an out-of-sample analysis of MV optimized portfolios was performed, where the
performance of five prominent covariance matrix estimators were compared, holding all other
things equal in the MV optimization. The optimization was performed under realistic investment constraints, taking incurred transaction costs into account, and for an investment asset
universe ranging from equity to bonds.
The empirical findings in this study suggest one dominant estimator: the covariance matrix
estimator implied by the Gerber Statistic (GS). Specifically, by using this covariance matrix estimator in lieu of the traditional sample covariance matrix, the MV optimization rendered more
efficient portfolios in terms of higher Sharpe ratios, higher risk-adjusted returns and lower maximum drawdowns. The outperformance was protruding during recessionary times. This suggests
that an investor that employs traditional MVO in quantitative asset allocation can improve their
asset picking abilities by changing to the, in theory, more robust GS covariance matrix estimator
in times of volatile financial markets.
Keywords: portfolio allocation, mean-variance optimization, efficient frontier, covariance matrix, estimation error, optimization enigma, random matrix theory, shrinking, robust statistics

Mean-Variance Portfolio Optimization:
Challenging the role of traditional covariance estimation
A Master’s Thesis
KTH Royal Institute of Technology
Stockholm, Sweden
Zakaria Marakbi
marakbi@kth.se
June 6, 2016

Preface
This study was conducted during the spring of 2016 at the Royal Institute of Technology (KTH).
I would like to express my sincere gratitude to Tomas Sörensson at the Royal Institute of
Technology for providing valuable insights and support during the process of writing this thesis.
In addition, I would like to thank my fellow students at KTH Royal Institute of Technology
who have reviewed my work and contributed to the process of writing this thesis.

i

Contents
1 Introduction
1.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.2 Problematization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.3 Purpose, research questions and expected contribution . . . . . . . . . . . . . . .
1.4 Disposition of the thesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1
1
3
4
4

2 Literature Review
2.1 Portfolio allocation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2.2 Parameter estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2.2.1 Expected returns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2.2.2 The covariance matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

5
5
8
10
14

3 Theoretical Framework
3.1 Basic preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.1.1 Return . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.1.2 Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.1.3 Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.2 Portfolio optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.2.1 MVO - an equivalent deterministic formulation . . . . . . . . . . . . . . .
3.2.2 Matching the real world - short selling and transaction costs . . . . . . .
3.3 Estimating expected returns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.4 Estimating the covariance matrix . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.4.1 The sample covariance matrix . . . . . . . . . . . . . . . . . . . . . . . . .
3.4.2 The single-index market model . . . . . . . . . . . . . . . . . . . . . . . .
3.4.3 Shrinkage towards the single-index model . . . . . . . . . . . . . . . . . .
3.4.4 Random matrix theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.4.5 Gerber statistic based covariance matrix . . . . . . . . . . . . . . . . . . .
3.4.6 Nearest positive semidefinite covariance matrix . . . . . . . . . . . . . . .
3.5 Performance measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.5.1 Sharpe ratio . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.5.2 Maximum drawdown . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.5.3 Portfolio weight turnover . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.5.4 Risk-adjusted return . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

19
19
19
20
20
23
24
27
29
30
30
31
33
35
36
40
40
40
40
41
41

4 Methodology
4.1 Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4.1.1 Sub-periods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4.2 Portfolio performance evaluation methodology . . . . . . . . . . . . . . . . . . . .

42
42
43
44

ii

Contents

4.2.1 MVO Specifics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4.2.2 Backtesting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
Covariance prediction accuracy . . . . . . . . . . . . . . . . . . . . . . . . . . . .

44
44
45

5 Results
5.1 Sub-periods within the full evaluation period . . . . . . . . . . . . . . . . . . . .
5.2 Portfolio out-of-sample performance . . . . . . . . . . . . . . . . . . . . . . . . .
5.2.1 SI versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.2.2 SM versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.2.3 RMT versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.2.4 GS versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.2.5 Summary of out-of-sample performance . . . . . . . . . . . . . . . . . . .
5.3 Covariance prediction analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3.1 SI versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3.2 SM versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3.3 RMT versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3.4 GS versus HC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3.5 Summary of covariance prediction analysis . . . . . . . . . . . . . . . . . .

47
47
48
48
51
54
57
59
62
62
63
64
65
66

6 Discussion
6.1 Portfolio out-of-sample performance . . . . . . . . . . . . . . . . . . . . . . . . .
6.2 Additional validity concerns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

69
69
71

7 Conclusion

72

Bibliography

73

4.3

iii

List of Figures
2.1
2.2

2.3

The MPT investment process. Illustration inspired by Fabozzi, Gupta, and Markowitz
(2002). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5
Density functions for a random variable R with E[R] = 1.1 and Var(R) = 0.32 .
For the left plot, R is N(1.1,0.32 )-distributed. For the right plot, the distribution
comes from a two point mixture of normal distributions. . . . . . . . . . . . . .
7
A bimodal density function symmetric around the mean, with E[R] = 1.1 and
Var(R) = 0.32 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

3.1

A feasible set spanned by three constraints. The shadowed area illustrates the
feasible set. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.2 An illustration of a convex function. Source: Sasane and Svanberg (2014, p.79). .
3.3 Q and all points in the interior of the segment AB are local minimizers, whereas
P is a global minimizer. Source: Sasane and Svanberg (2014, p. 105). . . . . . .
3.4 The blue frontier illustrates the efficient frontier. According to Markowitz, an
investor should only consider portfolios lying on the efficient frontier when selecting
a portfolio. The choice of portfolio on the efficient frontier depends on the riskaversion coefficient of the investor. . . . . . . . . . . . . . . . . . . . . . . . . . .
3.5 Illustration of the bias-variance tradeoff. The blue dots correspond to estimates.
Estimates within the red circle are considered accurate. Source: Fortmann-Roe
(2012, p. 1). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.6 Geometric interpretation of Theorem 1. Source: Ledoit and Wolf (2003b). . . . .
3.7 Simulated asset returns in a 24-month window. . . . . . . . . . . . . . . . . . . .
3.8 Thresholds imposed on both assets, which determines the sensitivity to variable
movement. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.9 An illustration of the significant pairwise returns. . . . . . . . . . . . . . . . . . .
3.10 Significant pairwise returns. Points lying in upper right area and the lower left
area correspond to significant pairwise returns in the same direction (marked blue).
Points lying in the upper left and lower right area correspond to significant pairwise
returns in the opposing direction (marked red). . . . . . . . . . . . . . . . . . . .
5.1
5.2

Ex-ante determination of sub-periods within the full evaluation period, conditional
on monetary policy signals as well as stock market signals. . . . . . . . . . . . . .
An illustration of realized performance in terms of annualized return and annualized volatility of portfolios corresponding to different target levels of risk. The
blue frontier corresponds to ex-post performance of SI-based portfolios, whereas
the orange one represents HC-based portfolios. . . . . . . . . . . . . . . . . . . .

iv

21
21
22

26

30
34
38
38
39

39

47

49

List of Figures

5.3

An illustration of realized performance in terms of annualized return and annualized volatility of portfolios corresponding to different target levels of risk. The
blue frontier corresponds to ex-post performance of SM-based portfolios, whereas
the orange one represents HC-based portfolios. . . . . . . . . . . . . . . . . . . .
5.4 An illustration of realized performance in terms of annualized return and annualized volatility of portfolios corresponding to different target levels of risk. The blue
frontier corresponds to ex-post performance of RMT-based portfolios, whereas the
orange one represents HC-based portfolios. . . . . . . . . . . . . . . . . . . . . .
5.5 An illustration of realized performance in terms of annualized return and annualized volatility of portfolios corresponding to different target levels of risk. The
blue frontier corresponds to ex-post performance of GS-based portfolios, whereas
the orange one represents HC-based portfolios. . . . . . . . . . . . . . . . . . . .
5.6 Realized risk-adjusted returns of portfolios over the full evaluation period, allocated through employing various covariance matrix estimators in MVO. . . . . .
5.7 Realized risk-adjusted returns of portfolios over expansionary periods, for five
different risk profiles. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.8 Realized risk-adjusted returns of portfolios over the full evaluation period, for five
different risk profiles. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.9 Maximum monthly drawdown of portfolios over the full evaluation period, for five
different risk profiles. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.10 Annualized portfolio weight turnover over the full evaluation period, for five different risk profiles. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.11 Time series for the mean of ratios of realized to predicted volatility for HC- and
SI-based estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.12 Histograms for the ratios of realized to predicted volatility for HC- and SI-based
estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.13 Time series for the mean of ratios of realized to predicted volatility for HC- and
SM-based estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.14 Histograms for the ratios of realized to predicted volatility for HC- and SM-based
estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.15 Time series for the mean of ratios of realized to predicted volatility for HC- and
SM-based estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.16 Histograms for the ratios of realized to predicted volatility for HC- and RMT-based
estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.17 Time series for the mean of ratios of realized to predicted volatility for HC- and
GS-based estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.18 Histograms for the ratios of realized to predicted volatility for HC- and GS-based
estimators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.19 Difference between realized and predicted risk for five different investor profiles,
with different target portfolio risks. Note that the y-axis is the difference in this
case, and not the ratio between realized and predicted portfolio risk. The percentage shown in the parentheses on the x-axis correspond to the target portfolio risk
for the particular investor profile. The closer the bar is to zero on the y-axis, the
better. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

v

52

54

57
59
60
60
61
61
63
63
64
64
65
65
66
66

67

List of Tables
4.1

Asset descriptive statistics for the nine assets to be considered in the empirical analysis. The series from which the descriptive statistics relate to constitute
monthly data from the period January 1994 to December 2013. Data from the
period January 1992 to January 1994 is excluded as the optimizer will require two
years worth of monthly data to initialize the first portfolio. I.e., January 1994
to December 2013 is the evaluation period. TR denotes total return data, which
accommodates for dividends and splits. . . . . . . . . . . . . . . . . . . . . . . .

Ex ante defined sub-periods within the full evaluation sample ranging from 1st
January 1994 to 31st December 2013. . . . . . . . . . . . . . . . . . . . . . . . .
5.2 Descriptive statistics for SI- and HC-based portfolios, at five different risk target
levels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.3 Relative return performance for the full evaluation period. The returns of SIbased portfolios are converted to risk-adjusted returns by deleveraging SI portfolio
volatility to HC portfolio volatility. ∆ denotes the difference in return between
SI-based portfolios and HC-based portfolios. . . . . . . . . . . . . . . . . . . . . .
5.4 Descriptive statistics for the four sub-periods from 1994 to 2013 for SI- and HCbased portfolios, at five different risk target levels. . . . . . . . . . . . . . . . . .
5.5 Relative return performance for the sub-periods that constitute the full evaluation
period. The returns of SI-based portfolios are converted to risk-adjusted returns
(RAR) by deleveraging SI portfolio volatility to HC portfolio volatility. ∆ denotes
the difference in return between SI-based portfolios and HC-based portfolios. . .
5.6 Full period descriptive statistics for SM- and HC-based portfolios, at five different
risk target levels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.7 Relative return performance for the full evaluation period. The returns of SMbased portfolios are converted to risk-adjusted returns (RAR) by deleveraging SM
portfolio volatility to HC portfolio volatility. ∆ denotes the difference in return
between SM-based portfolios and HC-based portfolios. . . . . . . . . . . . . . . .
5.8 Descriptive statistics for the four sub-periods from 1994 to 2013 for SM- and HCbased portfolios, at five different risk target levels. . . . . . . . . . . . . . . . . .
5.9 Relative return performance for the sub-periods that constitute the full evaluation
period. The returns of SM-based portfolios are converted to risk-adjusted returns
(RAR) by deleveraging SM portfolio volatility to HC portfolio volatility. ∆ denotes
the difference in return between SM-based portfolios and HC-based portfolios. . .
5.10 Full period descriptive statistics for RMT- and HC-based portfolios, at five different risk target levels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

43

5.1

vi

48
49

50
50

51
52

53
53

54
55

List of Tables

5.11 Relative return performance for the full evaluation period. The returns of RMTbased portfolios are converted to risk-adjusted returns (RAR) by deleveraging
RMT portfolio volatility to HC portfolio volatility. ∆ denotes the difference in
return between RMT-based portfolios and HC-based portfolios. . . . . . . . . . .
5.12 Descriptive statistics for the four sub-periods from 1994 to 2013 for RMT- and
HC-based portfolios, at five different risk target levels. . . . . . . . . . . . . . . .
5.13 Relative return performance for the sub-periods that constitute the full evaluation
period. The returns of RMT-based portfolios are converted to risk-adjusted returns (RAR) by deleveraging RMT portfolio volatility to HC portfolio volatility.
∆ denotes the difference in return between RMT-based portfolios and HC-based
portfolios. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.14 Full period descriptive statistics for GS- and HC-based portfolios, at five different
risk target levels. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5.15 Relative return performance for the full evaluation period. The returns of GSbased portfolios are converted to risk-adjusted returns (RAR) by deleveraging GS
portfolio volatility to HC portfolio volatility. ∆ denotes the difference in return
between GS-based portfolios and HC-based portfolios. . . . . . . . . . . . . . . .
5.16 Descriptive statistics for the four sub-periods from 1994 to 2013 for GS- and HCbased portfolios, at five different risk target levels. . . . . . . . . . . . . . . . . .
5.17 Relative return performance for the sub-periods that constitute the full evaluation
period. The returns of GS-based portfolios are converted to risk-adjusted returns
(RAR) by deleveraging GS portfolio volatility to HC portfolio volatility. ∆ denotes
the difference in return between GS-based portfolios and HC-based portfolios. . .
5.18 Portfolio risk prediction accuracy for the five covariance matrix estimators considered in this thesis. A ratio of 1 implies that the realized portfolio risk coincides
with the predicted portfolio risk. . . . . . . . . . . . . . . . . . . . . . . . . . . .

vii

55
56

56
57

58
58

59

67

Chapter 1

Introduction
This chapter aims to provide a background of the phenomenon under study as well as to introduce
the reader to the problem that constitutes the focal point of this thesis.

1.1

Background

Financial researchers and practitioners have long been interested in ways of allocating various
assets in an efficient manner. In this setting, an efficient portfolio refers to a portfolio that
yields the highest possible return given a certain level of risk that the investor is willing to take.
Naturally, such portfolios are appealing to portfolio managers around the world and the existing
body of knowledge includes a significant amount of research on this matter, which to a large
extent is dominated by quantitative models.
Out of these quantitative models, one particular approach is protruding and prominent:
Mean-Variance Optimization (MVO), introduced in a groundbreaking article published in 1952
by Markowitz (1952) for which he later was awarded the Sveriges Riksbank Prize in Economic
Sciences in Memory of Alfred Nobel (KVA 1990). The publication proved to become a cornerstone
in Modern Portfolio Theory (MPT) and an important stepping stone towards the creation of
further financial models such as the Capital Asset Pricing Model (CAPM), developed by Sharpe
(1964). Some purists of modern financial economics go as far as claiming that the publication
by Markowitz was “the moment of birth of modern financial economics”, as exemplified by
Rubinstein (2003).
The article by Markowitz (1952) can be seen as a reaction to previous and existing research
at that time, which to a large extent employed the law of large numbers theorem by Bernoulli,
leading to conclusions that all risk could be diversified away. Markowitz, however, did not share
this assertion. Instead, he claimed that the law of large numbers was not applicable to a portfolio
of securities, partly due to the prevalent interdependency and complexity in financial markets.
In other words, the inter-correlation between financial securities implies that diversification can
not entirely eliminate risk according to Markowitz. It is from this assertion that the essence of
Markowitz’s revolutionary theory stems, i.e. the existence of a trade-off relationship between
return and variance where variance is perceived as a measure of risk.
An underlying assumption of MVO is that investors are risk averse and rational. Thus,
an investor will always select a portfolio associated with less variance, ceteris paribus, and the
choice of the portfolio is solely based on the relationship between expected return and variance.
According to Markowitz (1952), the portfolio selection process is divided into two stages:

1

Chapter 1. Introduction

1. Parameter estimation: from historical observations and beliefs, one forms estimations of
future performance (in terms of return and variance) of the specified universe of securities.
2. Portfolio selection: employing the estimated parameters in the first stage, one choose an
efficient portfolio of securities. The security weights of the portfolio is obtained by solving
an optimization problem that is in line with the investor’s preferences.
In Markowitz’s pivotal publication in 1952, he was primarily interested in the portfolio selection
stage. In greater detail, this stage is related to his main proposal that an investor solely should
consider efficient portfolios. Recall that an efficient portfolio refers to a portfolio with a maximum
expected return for a given variance or less, or conversely, a portfolio with minimum variance
for a given expected return or more. In order to obtain efficient portfolios, Markowitz (1952)
presents a corresponding optimization problem, with the following mean-variance (MV) objective
function:
δ
(1.1.1)
w> µ − w> Σw
2
where w is a vector of portfolio weights, µ is a vector of expected returns for a set of assets and
Σ is a corresponding covariance matrix of asset returns. Furthermore, δ denotes the coefficient
of risk aversion, which controls the extent of how much additional risk is penalized. In short,
the objective function represents a trade-off between the expected return of a portfolio and its
expected variance. Markowitz (1952) showed that by solving this problem for different values of
the risk aversion coefficient, a set of efficient portfolios is obtained. These portfolios constitute
what Markowitz’s refers to as the efficient frontier and are the only portfolios that an investor
should consider in the context of portfolio selection.
During the years following the publication of Markowitz’s original work in 1952, his contribution to the modern portfolio theory has been eminent. Despite decades of research and debate,
Markowitz’s mean-variance MPT has remained the cornerstone of portfolio selection methods. It
is up to this state of time still topical and widely employed by financial practitioners, in particular
by portfolio managers (Tu and Zhou 2011). However, since the emergence of the MPT in 1952
and as economic and mathematical theory have progressed, several critics have questioned the
original work of Markowitz. While many agree that Markowitz’s mean-variance MPT is an important theoretical advance, questions have been raised regarding the first stage of Markowitz’s
portfolio selection process, i.e. the stage where one estimates the expected return vector, µ, and
the covariance matrix, Σ, both serving as essential inputs for the MVO problem. This was, and
still is, a universally encountered problem regarding the application of traditional MPT.
There is, however, an extensive literature devoted to the above mentioned critique. In greater
detail, within the realm of modern portfolio theory, the importance of the covariance matrix
has historically been overshadowed by the expected return parameter. This has led to several
proposals put forward regarding how one should estimate the expected return vector, some of
which are illuminated below:
• In 1992, Black and Litterman (1992) published their newly developed model: the BlackLitterman model. This model stems from an equilibrium assumption that the global market
portfolio is well diversified and efficient, serving as the neutral initial stage of the approach.
Thereafter, through the use of a reverse optimization process, the model derives returns of
assets implied by the market portfolio. In addition, an investor may convey agreement or
disagreement with the returns implied by the market, as they may claim expertise in value
investment that differs from the market consensus. Following the procedure of Black and
Litterman (1992), a vector consisting of the forecasted expected returns for a considered
asset universe can then be obtained.
2

Chapter 1. Introduction

• In addition to the Black-Litterman model, Sharpe (1964) and Fama and French (1993)
developed the CAPM and the Fama-French three-factor model respectively. In short, these
are models that attempt to describe asset returns and aid in estimating expected returns.
The above models merely capture a glimpse of the existing body of literature on this matter.
In recent times, however, research that point to the importance of the covariance matrix estimation in MVO has emerged. Thus, a demand for the relative performance between different
methods in estimating the covariance matrix has surged among financial practitioners. While
the existing body of knowledge regarding the estimation of expected returns is extensive, the
covariance matrix (Σ) estimation, which is an essential input in MVO, has arguably been overshadowed. Thus, in comparison to the return parameter, less is known regarding the performance
of MVO when the way of estimating the covariance matrix is varied, ceteris paribus. In a sense,
there exists a gap in the existing literature regarding this phenomenon, a gap which this thesis
aims to reduce.

1.2

Problematization

In the search for optimal portfolio allocations, investment managers have traditionally relied on
MPT as introduced by Markowitz in 1952. However, finding optimal allocations in MPT requires
estimation of covariances as well as expected returns of the assets included in the portfolio
optimization. Additionally, since portfolio optimization is dependent on expectations about
stochastic phenomena, the process of optimizing portfolios is prone to estimation error (Elton,
Gruber, Brown, et al. 2014).
It is well known that estimation errors of the input parameters, the expected return vector and
the covariance matrix, has a vital impact on the output of MV optimization. In other words, the
resulting portfolios obtained from the solution of the MVO problem are sensitive to the choice of
inputs. Hence, the out-of-sample performance of these optimized portfolios is strongly dependent
on the estimation accuracy of the input parameters. Michaud (1989) coined this phenomenon as
“Markowitz optimization enigma”, where he claims that MV optimized portfolios are unfeasible
in practice due to their susceptibility to estimation errors.
Despite these drawbacks, MV optimization continues to be a prominent method for portfolio
selection among investment managers as more accurate estimation techniques regarding expected
returns have emerged since the assertion by Michaud (1989). These include the Fama and French
(1993) response to CAPM and the Black and Litterman (1992) procedure that combines CAPM
with unique investor views.
However, research regarding estimation techiques for the covariance matrix input in MVO
has not been subject to the same level of attention until recently where Gerber, Markowitz, and
Pujara (2015) introduced a, supposedly, more robust co-movement measure as a replacement
for historical correlation in the estimation of the covariance matrix. Gerber et al. (2015) coin
this measure as the Gerber Statistic (GS) and their main findings suggest that MV optimized
portfolios using GS as a substitute to historical correlation in the covariance matrix estimation
consistently outperformed portfolios using the traditional estimation technique. More specifically,
their results indicate that the entire efficient frontier can be raised upward, thus resulting in higher
Sharpe ratio portfolios being attainable.
The findings by Gerber et al. (2015) are indeed interesting and can result in significant
monetary benefits for investment managers if deemed valid and reliable. More importantly, the
research poses a problem (or an opportunity) regarding whether financial practitioners that today
employ the traditional covariance matrix estimation technique should alter their quantitative
models. Clearly, the choice of estimation technique for the covariance matrix affect the MV
3

Chapter 1. Introduction

optimized portfolios. The problem is that there is an inadequate body of research regarding
the effect of out-of-sample performance when solely varying the estimation technique for the
covariance matrix. In other words, there is not a strong foundation regarding this field and
investment managers are most likely not willing to risk capital based on frontier research that
has not yet proven to stand up under close scrutiny.

1.3

Purpose, research questions and expected contribution

This thesis addresses investors that apply Markowitz’s mean-variance modern portfolio theory
in their quantiative models.
The purpose of this thesis is to investigate if portfolio managers that employ traditional
mean-variance optimization in practice can improve their asset-picking abilities by altering the
estimation technique regarding the covariance matrix. In addition, this thesis attempts to examine whether the relative forecasting performance of various covariance matrix estimators can
be tied to the prevailing market regime.
The study will attempt to achieve the purpose by answering the following research questions:
RQ 1: How does the choice of technique on how to estimate the covariance matrix
affect the out-of-sample performance of mean-variance optimization?
RQ 2: How does the prediction accuracy differ between the traditional sample covariance matrix and alternative estimators, and can the relative performance be tied
to the prevailing market dynamics?
In addition to providing insight for financial practitioners that employ MVO in their quantitative models, I expect this thesis to add to the current knowledge regarding the robustness of
different covariance estimation techniques. This knowledge can be applied to fields that are not
necessarily strictly related to modern portfolio theory. Examples of such fields are the insurance
industry and machine learning.

1.4

Disposition of the thesis

• Chapter 2, Literature Review, provides a literature review of relevant research within the
field of portfolio optimization. This chapter further motivates the theoretical framework
and methodology employed in this thesis.
• Chapter 3, Theoretical Framework, introduces the underlying theory that underlie this
thesis. It starts off by presenting some preliminaries, followed by a derivation of the portfolio
optimization problem considered in this thesis. Lastly, theory regarding the estimation
process of the expected return vector and the covariance matrix is presented.
• Chapter 4, Methodology, presents the general methodology employed for investigating the
research questions in this thesis.
• Chapter 5, Results, provides the empirical findings in this study.
• Chapter 6, Discussion, consists of a discussion regarding the empirical findings. In addition,
the validity of the thesis is challenged here.
• Chapter 7, Conclusion, concludes the thesis and presents some suggestions for further
research.

4

Chapter 2

Literature Review
This chapter will provide an extensive review of literature associated with portfolio optimization.
Essentially, this chapter motivates the theoretical framework and methodology employed in this
thesis.

2.1

Portfolio allocation

Almost 65 years have passed since Markowitz (1952) pioneered the use of mean-variance optimization in the context of portfolio management, where he quantified the concept of diversification
by employing the notions of return, volatility and covariance. Markowitz’s (1952) seminal work
has played a prominent role in modern portfolio theory and has been widely debated in the literature. Despite the fact that Markowitz became a Nobel laureate for the aforementioned work,
the framework has stirred controversy in scientific circles and has been subject to criticism which
challenges the validity of his proposed portfolio allocation model. Prior to illuminating the criticism in greater detail, it is important to understand the concept of mean-variance optimization
and how it is composed of several tethered components. In doing so, the literature review can
be divided into several parts, addressing different aspects of the mean-variance framework. The
following Figure 2.1 provides an overview of the portfolio selection process:
Consciousness
regarding the model’s
applicability

Constraints on
Portfolio Choice

Parameter Estimation

Mean-Variance
Optimization

Investor Objectives

- Expected Return Model
- Volatility & Correlation Estimates

Optimal Portfolio

Figure 2.1: The MPT investment process. Illustration inspired by Fabozzi, Gupta, and Markowitz
(2002).

5

Chapter 2. Literature Review

The node regarding model applicability is an extension of the original representation of the
MPT investment process described by Fabozzi, Gupta, and Markowitz (2002). Its appearance
in this thesis may be well motivated by tracing back to events where ignorance towards model
applicability resulted in catastrophic outcomes. Such an event can be found by recollecting the
work by Li (1999) who pioneered the use of Gaussian copulas for predicting the performance of
collateralized debt obligations (CDOs). In the years following his work, Li’s model became deeply
entrenched within the financial industry - in fact, it became so deeply entrenched that warnings
about its limitations and applicability were ignored by the practitioners. Moving forward to
the crisis of 2008, when the financial system’s foundation was severely ruptured, the financial
environment had altered in such magnitude that Li’s model could not anticipate. Hence, the
model became a recipe for disaster and has been partly credited to blame for laying the global
banking system in serious peril (Jones 2009). Already in 2005, Li warned the practitioners that
employed the model without being aware of the underlying assumptions of the formula. In
Whitehouse (2005, p.2), Li stated that:
‘The most dangerous part is when people believe everything coming out of it.’
In this context, Derman and Wilmott (2009) discuss the concept of model awareness within
the field of finance. They appreciate the simplicity in financial models, but hasten to assert that
while models are simple, reality is not. In other words, in the essence of models lies that they
do not perfectly mirror the reality. Confusing an illusion, which is what a model essentially is,
with reality can be a recipe for disaster. Moreover, Derman and Wilmott (2009, p.2) claim:
‘The most important question about any financial model is how wrong it is likely to
be, and how useful it is despite its assumptions.’
They further argue that, in stark contrast to true laws that may be found in the field of
physics, financial models are more fragile systems. The motivation behind this assertion is that
the world of finance is profoundly connected to human behavior which is too complex to entirely
capture in a simplified model. Thus, in contrast to true laws such as Newton’s law of gravity
found in the field of physics, there are no fundamental laws of finance - and even if there were,
it would not be possible to verify them through repeated experiments. Hence, Derman and
Wilmott (2009) argue that it is of utmost importance to be aware of the subtleties associated
with a quantitative model and not to confuse its illusion with reality. Knowing what is assumed
and what is swept out of view in a model is crucial.
Having illuminated the importance of being aware of the assumptions underpinning a model,
the focus of the literature review is now shifted towards the cornerstone model in this thesis: the
mean-variance framework. Fabozzi, Kolm, et al. (2007) note that a common misunderstanding
that is prevalent in the literature is that Markowitz’s mean-variance framework relies on the
assumption that security returns are jointly normally distributed. Fabozzi, Kolm, et al. (2007)
argue, however, that the mean-variance approach is consistent with the assumption of joint
normality and that the misconception may stem from this relationship. The basic assumption
and principle of the mean-variance framework may be found in a wide range of textbooks and
articles, such as inter alia Markowitz (1959) and Fabozzi, Kolm, et al. (2007) and follows as:
• The underlying assumption for the MPT mean-variance model is that an investor’s preferences can be captured by a utility function of the following two moments of portfolio
returns: the expected return and the variance of the portfolio.
• The principle underpinning mean-variance optimization is that for a given level of expected
return, a rational investor will choose a portfolio associated with a minimum amount of
variance amongst the feasible set of portfolios.
6

Chapter 2. Literature Review

Markowitz (1959) argue that these aspects underpin the mean-variance framework under the
umbrella of portfolio allocation. Following the introduction of MPT mean-variance optimization,
decades of debate and research on the subject have lead to an ambiguous academic support for the
framework. While the approach has found a widespread acceptance among financial practitioners
(Tu and Zhou 2011), the framework has been subject to a large extent of criticism in the academic
circles for not matching the real world in many ways (Feldstein (1969); Rockafellar and Uryasev
(2000)).
The first strand of criticism that will be illuminated relates to the concept of employing
variance as a proxy for risk. Hult et al. (2012) provide a lucid example that shed light on
one of the shortcomings with employing variance as a risk measure. Noting that the following
description may be considered as a parsimonious version of the mentioned example, Hult et
al. (2012) remark that location (mean) and dispersion (variance) are reasonable measures of
probable reward and risk, respectively, under the condition that the return is approximately
normally distributed. They further assert that since variance is a full domain measure that
quantifies a range of likely deviations from the mean, it may inaccurately describe the riskiness
of a position if the return is e.g. asymmetrically distributed. This is illustrated in Figure 2.2:
2

2
Mean = 1.1

Mean = 1.1

1.5

1.5

1

1

0.5

0.5

0

0
0

0.5

1

1.5

2

0

0.5

1

1.5

2

Figure 2.2: Density functions for a random variable R with E[R] = 1.1 and Var(R) = 0.32 . For the left
plot, R is N(1.1,0.32 )-distributed. For the right plot, the distribution comes from a two point mixture
of normal distributions.

Having the same mean and variance, both profiles in Figure 2.2 are identical from a meanvariance perspective. However, from a downside risk perspective, the riskiness of the positions
is inherently nonequivalent due to the asymmetric display of the density function to the right.
As the variance symmetrically accounts for deviations from the mean, it fails to adequately
discriminate between return distributions (Grootveld and Hallerbach 1999).
Research by, inter alia, Post and Van Vliet (2005) and Ang, Chen, and Xing (2006) suggest
that investors assign greater importance to downside risk as opposed to upside risk which they
view favorably. To cope with the shortcomings of employing variance as a risk measure in this
regard, several downside risk measures have emerged throughout the literature, leading to several
offsprings of the MPT mean-variance framework. Arguably, Value-at-Risk (VaR), popularized by
J.P. Morgan’s RiskMetrics in 1996, is the downside measure that has gained most traction over
the recent years within the field of finance (Glasserman, Heidelberger, and Shahabuddin 2002).
VaR allows the investor to measure the maximum predicted loss at a certain confidence level
(typically 95%). Consequently, empirical findings suggest that VaR enables one to better account
for downside risk in comparison to variance (Litterman (1996); Hendricks (1996)). However, a
general consensus prevails in the body of literature that VaR has its own drawbacks. Artzner
and Delbaen (1997) display that VaR has undesirable mathematical properties such as a lack

7

Chapter 2. Literature Review

of sub-additivity and convexity. Thus, VaR does not necessarily reward diversification and can
exhibit multiple local extrema, making it computationally intractable to find the global optimal
point in the optimization process for portfolio allocation (Beder (1995); Mausser and Rosen
(1999)). In addition, in the presence of fat left tails, VaR, like variance, has been criticized for
not accomodating for the magnitude of losses beyond the VaR value (Fabozzi, Kolm, et al. 2007).
Conditional Value-at-Risk (CVaR) is a risk measure that has surfaced the academic field of
finance in recent years. It attempts to rectify the mentioned shortcomings of VaR (Artzner,
Delbaen, et al. 1999). In fact, Pflug (2000) proved that CVaR is a coherent risk measure, which
was later supported by Krokhmal, Palmquist, and Uryasev (2002) who concluded that CVaR is
indeed a more reliable risk measure than VaR as it is sub-additive and convex.
While the above risk measures have been hailed in scientific circles and led to extensions of
the mean-variance framework (see e.g. mean-VaR, mean-CVaR; Artzner, Delbaen, et al. (1999)),
it comes at the cost of simplicity and computational tractability. Fantazzini (2004) argue that
this is why the vast majority of applied professionals prefer to rely on more traditional models
such as the mean-variance framework. He supports this assertion by drawing a parallel to the
financial field of option pricing where Black & Scholes still is the most practised model despite
the fact that more sophisticated models have emerged within the academia.
Following the aftermath of the financial tumult in 2008, downside risk measures have regained
an increased attention due to its ability to better consider for black swan events1 . However,
despite their theoretical appeal, Fabozzi, Kolm, et al. (2007) argue that downside risk measures
are difficult to implement in a portfolio allocation setting. They remark that this is partly due
to the fact that downside risk measures often entails computational intractability.
In light of the above review, it should be apparent that the definition of risk is ambiguous
in the literature. While some prefer the simplicity and interpretability of the mean-variance
approach, others argue for the use of downside risk measures. As the purpose of this thesis does
not lie in evaluating different risk measures, but rather how they can be estimated, it should
exist no ambiguity in how risk is measured. Hence, with the motivation that the mean-variance
framework is seemingly most employed in practice and in line with Markowitz (1952), this thesis
will restrain to variance as a proxy for risk, bearing its limitations in mind.

2.2

Parameter estimation

More importantly for this thesis is the strand of criticism that refers to the phenomenon that
mean-variance optimization suffers from the problem of estimation error. In the pivotal work by
Markowitz (1952), the focus lied on the theoretical soundness of the suggested portfolio selection
approach. Consequently, less emphasis was placed on how to implement MVO in practice.
In order to practically implement MVO, one needs to estimate the means and covariances of
asset returns as these moments are not known. These estimates are then employed to obtain a
solution for the investor’s optimization problem (Elton, Gruber, Brown, et al. 2014). There is
an abundance of literature that conclude that this leads to one of the most important drawbacks
of the mean-variance approach: i.e. the estimation error of the plug-in moments (see, inter
alia, Michaud (1989); Chopra and Ziemba (1993)). The drawback arises from the fact that the
optimizer is not aware that the inputs are statistical estimates and not known with certainty.
When estimating asset means and covariances of returns, the classical statistical procedure
has been to gather a history of past returns and compute their respective sample estimates.
This procedure relies on the assumption that historical data has some predictive power for
1 Coined by Taleb (2007), a black swan event an event that significantly deviates from the expected normal
case but has a major impact, i.e. an outlier that lives in the utmost point of a heavy tail.

8

Chapter 2. Literature Review

future performance. However, throughout the literature, several deficiencies of employing sample
estimates in a portfolio setting have been well documented. For instance, Frankfurter, Phillips,
and Seagle (1971) assert that MV optimized portfolios obtained by using sample estimates as
plug-in parameters do not necessarily outperform an equally weighted portfolio (also known
as the naı̈ve portfolio). This was later supported by Jobson and Korkie (1980) who obtained
similar results. Moreover, Best and Grauer (1991) show that the estimation error of the parameter
estimates is transferred to the obtained portfolio weights from the portfolio optimization. Hence,
the estimated optimal weights will almost certainly always deviate from the true optimal weights.
In this context, Ledoit and Wolf (2003a) contribute to the discussion by explaining why
sample estimates may come with severe problems in a portfolio setting. They argue that the
poor performance of MV optimized portfolios stem from sample estimates that contain a lot of
error: in this case, the most extreme sample coefficients tend to take extreme values not because
it is the truth, but rather due to the associated error. Consequently, they argue that the MV
optimizer will, consistently, latch onto these extreme coefficients and place the biggest portfolio
weights accordingly. It is from this phenomenon that the critique by Michaud (1989) stems,
claiming that the portfolio optimizers introduced by Markowitz (1952) are in fact “estimation
error maximizers”. Michaud (1989) further coined this puzzle as “Markowitz’s optimization
enigma”.
The critique by Michaud (1989) implies that in the presence of inaccurate plug-in estimates in
MVO, asset managers will underrepresent their true asset-picking abilities, leading to suboptimal
portfolios. To cope with this prevalent problem, a proliferation of studies regarding the subject
have emerged. However, as pointed out by Ledoit and Wolf (2014), while a substantial amount
of effort has been devoted to estimating the expected return vector, much less has been written
about the covariance matrix. To exemplify this assertion, they refer to Green, Hand, and Zhang
(2013) who list over 300 articles that relate to methods of estimating expected returns.
The disparity in attention devoted to the expected return vector versus the covariance matrix input may potentially be explained by tracing back to the findings by Chopra and Ziemba
(1993), suggesting that error in the expected return has more impact on the optimized portfolio’s out-of-sample performance than error contained in the covariance matrix. In addition, the
concept of forecasting returns has always been subject to a great deal of attention ever since
the introduction of financial markets, beyond the field of portfolio allocation. However, recently
Michaud, Esch, and Michaud (2012) challenged the aforementioned claim by Chopra and Ziemba
(1993). Michaud et al. (2012) remark that there is a persistent widespread error in the literature
regarding the relative importance of estimation error in return relative to the covariance matrix.
They argue that the paper by Chopra and Ziemba (1993) is highly flawed and unreliable: their
underlying argument for this assertion is that the analysis in Chopra and Ziemba (1993) is based
on an in-sample specific study, thus having no bearing regarding the impact of estimation error
in rigorous out-of-sample MV optimization. Furthermore, in stark contrast, they claim that the
estimation error in the covariance matrix may in fact overwhelm the optimization process as the
size of the asset universe increases.
In light of the above review, it should be apparent that estimation error is one of the most
important aspects of MV optimization. Reducing such error has the potential to increase out-ofsample performance of MV optimized portfolios. In other words, there is a potential to obtain
more efficient portfolios, leading to significant monetary benefits for asset managers that employ
MVO. Historically, the general consensus has been to focus on forecasting returns. Consequently,
the body of literature regarding this matter is extensive. However, more recently, research that
point to the importance of the covariance matrix estimation in MVO has emerged. Thus, a
demand for the relative performance between different methods in estimating the covariance
matrix has surged among financial practitioners.

9

Chapter 2. Literature Review

The following two subsections will (i) provide an overview of different approaches in forecasting expected returns, and (ii) review the literature that include research on how estimation error
in the covariance matrix can be reduced, in line with the research questions in this thesis.

2.2.1

Expected returns

Explaining future return characteristics may be considered as the grail of financial economics.
It can be considered as the quest for financial practitioners that seek efficient asset allocation
through portfolio optimization.
The classical approach when estimating expected returns in MVO is to rely on the sample
means as predictors. Under the hypothesis of normality, the sample mean is the maximum likelihood estimate (MLE) and thus the best linear unbiased estimator for the assumed distribution.
In this case, the sample mean exhibits the property that an increase in sample size, leads to an
improved performance of the estimate. Fabozzi, Kolm, et al. (2007), however, note that under
distributions that are heavy tailed or significantly deviate from a symmetrical unimodal distribution, the above properties of the sample mean are no longer valid. They support this claim by
referring to the work of Ibragimov (2005). An example where the sample mean is a poor forecast
is depicted in Figure 2.3.
2
Mean = 1.1

1.5

1

0.5

0
0

0.5

1

1.5

2

Figure 2.3: A bimodal density function symmetric around the mean, with E[R] = 1.1 and Var(R) = 0.32 .

As one can observe, the outcomes of the random variable (in this case the return of an
arbitrary asset) are not necessarily likely to take values described by the overall mean. In this
case, the sample mean is a poor predictor for future returns.
Furthermore, Fabozzi, Kolm, et al. (2007) argue that the return generating process of financial
time series usually do not exhibit stationarity, but varies over time. This implies that historical
data from a long past may have little explanatory power for future behavior. Hence, in such
a setting, the mean is not a good forecast of expected returns - consequently leading to large
estimation error in the optimization process.
To cope with the shortcomings regarding sample moments, two prevalent approaches can
be found in the literature. The first approach is to impose structure on the estimator, usually
by relying on some factor model to forecast expected returns. The other approach is to use
Bayesian estimators such as the Black-Litterman model. These different perspectives will be
further presented throughout this section.
Factor models
The capital asset pricing model (CAPM) is a single-factor model that marked the birth of asset
pricing theory. Building on the foundation of MPT set forth by Markowitz (1952), CAPM
was initially introduced by Sharpe (1964). Following the footsteps of Markowitz, Sharpe was
awarded the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel for his
10

Chapter 2. Literature Review

pioneering contribution to financial economics in 1990 (KVA 1990). In the context of CAPM,
Lintner (1965) and Mossin (1966) are worthy mentions as they independently proposed similar
theories as Sharpe (1964). The CAPM can be seen as an abstraction of the real-world capital
markets and is based on a strict set of equilibrium assumptions. Sharpe (1964) conclude that
the following three main assumptions underlie CAPM:
1. Investors are rational and choose mean-variance efficient portfolios according to Markowitz
(1952).
2. Investors are in complete agreement: i.e. investors have homogeneous expectations regarding the volatilities, correlations and expected returns of the assets.
3. Investors can borrow and lend at the risk-free rate, which is the same for all investors.
Under these assumptions, CAPM states that the aggregated market portfolio is efficient.
Furthermore, as investors can eliminate firm-specific risk by diversifying their portfolios, no
investor should price it. It is from this assertion that the key idea of CAPM stems: i.e. that all
risk originates from a single factor, the market. This is commonly referred to as the systematic
risk that cannot be diversified away. The major implication of CAPM is then that the expected
return of any asset is determined by its covariation to the market. In other words, the CAPM is
a single-factor linear model that relates expected returns of an asset and a market portfolio.
Within the field of finance, the CAPM has found a widespread application amongst practitioners ever since its introduction. Five decades later, the model is still regarded as a centerpiece
in finance literature (Da, Guo, and Jagannathan 2012) and widely employed in a portfolio allocation setting. Fama and French (2004) argue that the attraction of the CAPM lies in its
simplicity: it offers a quick quantitative insight into risk-reward interplay and an intuitive tool
for predicting expected returns. In addition, the model is supported by a strong theoretical background from economic theory. However, in academic circles, the empirical validity of CAPM has
been widely debated and there is a prevalent consensus that, owing to its idealized assumptions,
the empirical record regarding the model’s validity is poor. Fama and French (2004) claim that
it is in fact poor enough to invalidate the way CAPM is used in applications altogether. Levy
and Roll (2010) remark that the widespread belief regarding the invalidity of CAPM stems from
research that conclude that various, commonly used market proxies are inefficient (see, for example, Jobson and Korkie (1982); Shanken (1985); Gibbons, Ross, and Shanken (1989)). These
findings do not coincide with the CAPM theory, consequently casting doubt on CAPM. However,
Levy and Roll (2010) show that slight adjustments, well within estimation error bounds, on the
sample parameters employed in evaluating the market portfolio suffice to make the market proxy
efficient. Hence, in stark contrast to the belief that beta is dead, their findings suggest that market proxies may be consistent with the CAPM theory after all, thus strengthening the usefulness
of employing CAPM when estimating expected returns. It is important to note, however, that
Levy and Roll (2010) hasten to add that their findings do not constitute a proof of the empirical
validity of CAPM, but acts as a response to the rejection of the model. They further note that
the validity of the global CAPM is not empirically testable as the true market portfolio is in fact
not observable (hence why proxies are employed), in accordance to the critique by Roll (1977).
Following the conception of the CAPM, several offsprings have emerged throughout the literature. These offsprings are often referred to as multi-factor models, whose birth may be
motivated by the growing number of studies that found that the market alone is not sufficient
in explaining asset returns (see Banz (1981); Basu (1983)). Often viewed as a counterpart to
CAPM, Ross (1976) derived an asset pricing model solely based on arbitrage arguments such as
the no-arbitrage condition. This model is commonly referred to as the arbitrage pricing theory
(APT) model. Contrary to the CAPM, the APT postulates that the expected return of an asset
11

Chapter 2. Literature Review

is influenced by a variety of risk factors, allowing for a more accurate explanation of asset returns
(Ross 1976). In addition, supporters of the APT argue that an advantage over the CAPM is that
it relies on less restrictive assumptions, e.g. by not relying on the assumption that all investors
are mean-variance optimizers. However, the strength of allowing for more explanatory variables
in explaining returns is also its weakness, as the APT provides no specification on which factors
to include in the model. Consequently, there is no consensus on the identity of explanatory variables and no consensus on the number of factors to include, resulting in a less tractable approach
than the CAPM.
Moreover, the Fama and French three-factor model is a prominent extension of the CAPM
that attempts to rectify the strand of criticism related to the aspect of omitted explanatory
variables in the CAPM (Fama and French 1995). Based on findings implying higher average
returns on small stocks and high book-to-market stocks, Fama and French (1993) argue that
there are unidentified variables that produce undiversifiable risks in returns, not captured by
the market return. They support this claim by illuminating the phenomena that the returns
of small firms covary more with one another compared to large firms, and that returns on high
book-to-market stocks covary more with one another in relation to the covariation between low
book-to-market stocks. Based on this evidence, Fama and French (1995) proposed a three-factor
model that extends the CAPM with the addition of two factors.
In the academic world, the Fama and French three-factor model has been widely accepted
as a CAPM empirical successor (Zabarankin, Pavlikov, and Uryasev 2014). Less so by financial
practitioners, which Bartholdy and Peare (2005) argue can be explained by their findings that
the additional cost in complexity associated with Fama and French is not justified by its relative
performance over the CAPM. Fama and French (2004) argue that the main shortcoming of
their three-factor model lies in its empirical motivation: the added explanatory variables are not
intuitive from a theoretical perspective, they are rather mere ’brute force’ constructs meant to
capture return patterns illuminated by previous research.
Noting that the above review merely captures a glimpse of the literature on factor models
in a forecasting context, these models stem from the assertion that sample moments based on
historical data are likely to contain random noise and errors. Factor models attempt to smooth
historic data and focus on the underlying relationships, while ignoring deviations from perceived
statistical relationships that are inferred by random noise. In the literature, one can indeed find
that factor models tend to outperform the plug-in sample mean in a portfolio allocation setting.
For instance, Chan, Karceski, and Lakonishok (1999) show in their study that estimates based
on factor models lead to improved out-of-sample performance of optimized portfolios, compared
to when sample plug-in estimates are employed. However, no favorite specification emerges
regarding the choice of factor model.
Ait-Sahalia and Hansen (2009) note that moving from theoretical factors such as the market
portfolio, to empirical factors (see Fama and French (1995)) and potentially to statistical factors,
we may by construction capture more underlying relationships. However, in exchange, the factors
become more difficult to interpret, which Ait-Sahalia and Hansen (2009) argue raises concerns
regarding data mining. Ultimately, choosing between factor models involves a trade-off between
estimation error, bias and interpretability. In this context, invoking the following quote by Ledoit
and Wolf (2003b, p.2) is appropriate:
‘The art of choosing a factor model adapted to a given data set without seeing its
out-of-sample fit is just that: an art’
Essentially, Ledoit and Wolf (2003b) convey the message that as there is no general consensus
on the identity of factors (except for the market) to employ in factor models, choosing a specific
factor model is very ad hoc (we do not know how well it works a priori). The simplicity of the
12

Chapter 2. Literature Review

CAPM does not entail that it necessarily performs worse than more complex models. In fact,
a common finding in forecasting literature is that simple, parsimonious models that may suffer
from severe misspecification often provide stronger forecasts than more complicated models (see
inter alia Swanson and White (1997); Stock and Watson (1999)). Ultimately, when it comes
to factor models, the Fama-French factor model is regarded as the front figure in the academic
world. However, there is a presence of disparity between the academic world and the real world
setting as financial practitioners find the theoretical motivation and the mathematical simplicity
of the CAPM appealing. It provides a quick quantitative insight on the risk-reward interplay of
assets. Hence, as in the working paper by Gerber et al. (2015) who analyze portfolio performance
in a real-world setting, this thesis will rely on the CAPM when estimating expected returns, while
still acknowledging that its application have been heavily debated within scientific circles (see
e.g. Galagedera and Galagedera (2007) for a detailed capital asset pricing review).
The Black-Litterman model
Within the realm of asset allocation, the Black-Litterman model is regarded as a prominent
extension of traditional MVO. Developed in the original paper by Black and Litterman (1992),
the Black-Litterman model stems from an equilibrium assumption that the global portfolio is well
diversified and efficient, serving as the neutral initial stage of the approach. Thereafter, through
the use of a reverse optimization process, the model derives returns of assets implied by the
market equilibrium. At this stage, a natural question that arises is how the model differentiates
from the CAPM. The answer lies in the model’s flexibility to combine the market equilibrium
with additional market views of an investor. More precisely, the Black-Litterman model permits
analysts to convey agreement or disagreement with the returns implied by the market (Black
and Litterman 1992). The intuition behind the model is that an analyst may claim expertise in
value investment that differs from the market consensus: so why not let the analyst incorporate
these views when deriving the vector of expected returns? In their original paper, Black and
Litterman (1992) conclude the intuition of their proposed model in the following manner:
‘...our approach allows us to generate optimal portfolios that start at a set of neutral
weights and then tilt in the direction of the investor’s views.’
If carefully used, Nikbakhtt (2011) summarize some of the advantages that the Black-Litterman
model may endow on the final portfolio, in comparison to a portfolio obtained through traditional
MVO. These advantages include:
• Estimation error is usually reduced.
• Portfolio weights are often, by construction, more intuitive with respect to the expressed
views.
• The recommended portfolio obtained through the optimization should be more efficient
and less concentrated on individual assets.
However, the Black-Litterman model has been critcized for its ambiguity as Black and Litterman (1992) did not discuss the precise nature of how one practically applies the model. Nikbakhtt
(2011) argue that the incorporation of views is in fact a major limitation when putting the BlackLitterman model into practice. He claims that only the most naive analysts are confident in their
additional market views and in the presence of casually expressed views, the model may become
dangerous. Hence, it is of utmost importance that analysts integrate their views with the greatest
of care. How one estimates the parameter of confidence on views in the Black-Litterman model

13

Chapter 2. Literature Review

is not clear as Black and Litterman (1992) did not discuss the precise nature of this phenomenon
in their original article. The criticism regarding the great deal of vagueness associated with the
Black-Litterman model stems from the lack of properly described parameters underpinning the
model. In particular, the most severe problem of the model concerns the vagueness of how one
determines the confidence parameter often referred to as the weight-on-views or tau. Articles
such as: “A demystification of the Black-Litterman model” (Satchell and Scowcroft 2000) and
“A step-by-step guide to the Black-Litterman model” (Idzorek 2002) serve as strong examples
regarding the difficulties of interpreting the original work by Black and Litterman (1992).
In addition, Nikbakhtt (2011) illuminates the question regarding the impact of legal risk
when using the Black-Litterman model: he argues that in the absence of a reliable algorithm
that incorporates investor views, clients may use the “prudent expert” principle for portfolio
management in court. On the other hand, if the views are well documented, objectively defined
and well justified, legal risk may decline.
In this thesis, the neutral initial stage of the Black-Litterman model will be used (i.e. the
CAPM model) to estimate expected returns, without incorporating any additional market views.
The reasoning behind this is simply that I deem it inappropriate to dilute the portfolio outof-sample analysis with subjective opinions and leave this additional flexibility open for asset
managers to integrate, if sought.

2.2.2

The covariance matrix

Within the realm of modern portfolio theory, the importance of the covariance matrix has arguably been overshadowed by the expected return parameter. As previously mentioned, this
can partly be credited to the findings of Chopra and Ziemba (1993), suggesting that the relative
influence of errors in the expected return vector is higher. In recent times, this claim has been
challenged by the likes of Michaud et al. (2012) who take an opposing stance. They argue that
errors in the covariance matrix may in fact overwhelm the optimization process when the asset
universe grows large.
Nevertheless, the disparity in attention devoted to the different areas is not to be confused
by an absence of literature regarding the covariance matrix estimator. In fact, following the
advancements of mathematical theory and computational power in recent times, a fair amount of
consideration has been put in developing alternative methods of estimating the covariance matrix
(see e.g. Laloux et al. (2000); Ledoit and Wolf (2003b); Gerber et al. (2015)). Other reasons
for the gained interest regarding this phenomenon can be found in Bengtsson and Holst (2002)
who argue that the notorious difficulty of estimating expected returns compared to estimating
the covariance matrix implies that most of the improvement that can be made on MVO lies in
the covariance matrix estimation.
The classical method of estimating the covariance matrix in the context of MV optimization
is to employ the sample covariance matrix. During the years following the work by Markowitz
(1952), numerous studies have shown that the sample covariance matrix may suffer from drawbacks which undermine its forecasting power of future covariances (Elton, Gruber, Brown, et al.
2014). This may come as a surprise as the sample covariance matrix has the appealing property
of being the maximum likelihood estimate under normality. However, this is to forget what
maximum likelihood actually means. Following Ledoit and Wolf (2003b), it means that all the
trust is put in the data which clearly is a sound principle if there is enough data to trust it. Not
enough data is thus a problem and while increasing the amount of data is a potential solution, it
may come at the expense of employing outdated noisy data with no explanatory power regarding
the future. Consequently, as exemplified in Bengtsson and Holst (2002), an important drawback
of the sample covariance matrix is that it may follow noise too closely and suffer from overfitting

14

Chapter 2. Literature Review

which will undermine the out-of-sample fit, despite being the best estimate in-sample.
In other words, the sample covariance matrix has been shown to require a lot of data. This
is exemplified in Bengtsson and Holst (2002) who show that the sample covariance matrix of 100
(N ) assets implies that 5050 (N (N + 1)/2) parameters have to be estimated. Therefore, small
sample problems may occur when the considered asset universe is large.
In the literature, the cure for the feasible drawbacks associated with the sample covariance
matrix is to impose some form of structure on the estimator. The vast majority of challengers
stem from the notion that there exists a bias-variance tradeoff. While imposing structure may
reduce the instability (variance) of the estimator, it may come at the expense of specification
error (bias). The idea is to find an estimator that prevails at the optimum balance between
bias and variance. Strong challengers found in the body of research regarding this area include
estimators based on factor models, shrinkage models, random matrix theory and threshold theory
(Elton, Gruber, Brown, et al. 2014).
Factor models
Ledoit and Wolf (2003b) claim that a natural way to impose structure on the covariance matrix
estimator is to use a low-dimensional factor structure. Within the world of finance, factor models
have found a widespread traction. However, as in the discussion regarding factor models when
estimating expected returns, Ledoit and Wolf (2003b) argue that two questions arise in this
context: how many factors should one use and what factors should be considered in the model?
According to Elton, Gruber, Brown, et al. (2014), there is no general consensus regarding the
answers to these questions, except for the common understanding that a market factor should
be included. The use of a market factor to explain the return generating process of asset is
motivated by economic theory and was first introduced in such a setting by Sharpe (1964).
Not surprisingly, the single-index model by Sharpe (1964) is one of the most prominent
structural model found in the literature. The key idea of the single-index model is to impose
structure by assuming that the only reason that two securities move together is due to their
common response to market changes. Effectively, all other factors (such as industry factors)
beyond the market are assumed not to account for any comovement between securities. This
rather strong assumption is the core of the single-index model and the validity of the model is
thus strongly dependent on how well the assumption holds.
In some embodiments, empirical studies have found that the estimated covariance matrix
implied by the single-index outperforms the full historical covariance matrix. More specifically,
Elton, Gruber, and Urich (1978) investigated the relative ability in forecasting the correlation
structure between securities for various correlation estimation techniques. Some striking results
of their study was that the sample correlation matrix underperformed the correlation matrix
implied by the single-index model when comparing predicted and realized correlation between
financial securities. Furthermore, they showed that these results were statistically significant.
This suggests that a part of the correlation structure for the full historical model represents
random noise, which is not captured by the structured single-index model.
Ledoit and Wolf (2003b) argue that the sample covariance matrix and the estimated covariance matrix implied by the single-index model can be viewed as two extremes. The first one
can be regarded as a full factor model that puts all the trust in the data, whereas the second
one is a single factor model that makes a rather restrictive assumption regarding what data is
relevant. In the presence of effects beyond the market factor that account for security comovement, the single-index model may thus come at the expense of introducing specification error.
In this context, a similar discussion as in the CAPM model for expected returns can be carried out regarding the introduction of additional factors. Recall that Ait-Sahalia and Hansen

15

Chapter 2. Literature Review

(2009) argue that one may, by construction, capture more underlying relationships by moving
to a multi-factor model that accounts for e.g. industry factors. However, it does not necessarily
entail a better out-of-sample fit. In addition, the lack of a general consensus regarding factors
apart from the market factor remains a problem as it raises concerns about data mining. Ledoit
and Wolf (2003b) add that choosing between factor models thus becomes very ad hoc and calls
it an art.
In this context, Chan et al. (1999) study the performance of different factor specifications
in a realistic portfolio allocation setting. Their findings suggest no dominating factor specification emerges. In fact, the parsimonious single-index model performed only marginally worse
than more complex specifications based on a weaker theoretical foundation. However, all considered factor models outperformed the sample covariance matrix, which once again suggests that
improvements can be made on this estimate.
With respect to prediction, Elton, Gruber, Brown, et al. (2014) further conclude that parsimonious models tend to outperform more complex models in many tests. Their explanation for
this is that complex models with multiple factors tend to contain more noise than real information.
Shrinking
Thus far, two extreme estimators have been reviewed: the sample covariance matrix and the
covariance matrix estimator implied by Sharpe’s (1964) single-index model. In addition, the
drawbacks of these two estimators illuminated in the literature have been presented. To reiterate,
it is well known that the sample covariance matrix may suffer from overfitting as it puts all the
trust in the data which in turn may render a poor out-of-sample fit. On the other hand, the
strong structure imposed by the single-index model comes at the price of potentially introducing
specification error. Thus, there exists a trade-off between specification error (bias) and estimation
error (variance) within the realm of estimation.
With this in mind, a recent proposal by Ledoit and Wolf (2003b) is to take a different approach
in imposing structure. They suggest to take a weighted average of the sample covariance matrix
and the covariance matrix estimator implied by the single-index model. This way, they let the
weight assigned to the single-index estimator control how much structure that is imposed. The
approach is inspired by the concept of shrinking, dating back to the work by Stein (1956) where
the weight assigned to the single-index estimator is the shrinkage intensity and the shrinkage
target is the estimated covariance matrix implied by the single-index model. Ledoit and Wolf
(2003b) argue that this approach has the advantage of being able to account for effects beyond
the market factor without the need of specifying an arbitrary multi-factor structure. This is very
convenient, given that there is no general consensus regarding the identity of factors except for
the market factor. The method is commonly called shrinkage to market.
The central idea is to find an optimal compromise between estimation error commonly associated with the sample covariance matrix and specification error introduced by the structured
single-index model. In order to achieve this, Ledoit and Wolf (2003b) derives a formula for the
optimal linear shrinkage intensity that controls the amount of structure that is imposed. The
derivation is done by working under large-dimensional asymptotics.
Applying their shrinkage estimator in mean-variance optimization, they further show that for
NYSE and AMEX stock returns ranging from 1972 to 1995, lower out-of-sample variance of MV
optimized portfolios can be obtained compared to solely using, inter alia, the sample covariance
matrix or the covariance matrix estimator implied by the single-index model. Here, Ledoit and
Wolf (2003b) use an equally weighted index of the asset returns as a market proxy. In this
context, Bengtsson and Holst (2002) found similar results for Swedish asset returns. Both these

16

Chapter 2. Literature Review

studies only use the minimum variance portfolio in their evaluations. Consequently, their results
provides limited information regarding the performance of the covariance matrix estimators in a
portfolio selection context (only one special portfolio is considered). In practice, investors have
various risk profiles and may be interested in portfolios beyond the minimum variance portfolio.
However, a valid response to this critique can be found in Bengtsson and Holst (2002) where they
motivate this choice by not wanting to dilute the covariance matrix performance with potential
errors in the expected return vector.
It is important to note that the single-index estimator is not an exclusive shrinkage target in
this setting. This is illuminated in Ledoit and Wolf (2003a) where they develop a new shrinkage
estimator using the constant correlation model as the shrinkage target. However, they hasten to
add that for an asset universe consisting of different asset classes, the constant correlation model
is not appropriate.
Lastly, Ledoit and Wolf (2004) are very clear that the improvement that the linear shrinkage
introduced in Ledoit and Wolf (2003b) has over the sample covariance matrix is dependent on
the situation at hand. For a relatively large asset universe (N ) compared to the number of
observations per asset (T ), the improvement is expected to be significant. On the contrary, if
the data per estimated parameter is high (i.e. when N/T is small), the improvement may be
minuscule.
Random matrix theory
Although developed in the 1950s by quantum physicists, random matrix theory (RMT) is a
fairly recent area within the realm of portfolio optimization. In this context, Laloux et al. (2000)
conducted a pivotal study where they seek to identify measurement noise often associated with
the sample covariance matrix. Their approach stems from the idea that if one can devise a method
to distinguish measurement noise that devoid useful information from signal (useful information)
contained in the estimated covariance matrix, the estimate can be enhanced. Employing known
results from random matrix theory, they show that approximately 94% of the eigenvalues that
constitute the sample correlation matrix for S&P500 stock returns (daily data ranging from 19911996) agree with the theoretical prediction of RMT. This suggests that the sample correlation
matrix may be considered random, to a certain extent. In other words, merely a few eigenvalues
were found to contain useful information in the construction which is a remarkable finding in
a quantitative portfolio selection context. Under the assumption that these results are valid,
Laloux et al. (2000) as well as Plerou et al. (2002) that if one filters the noisy eigenvalues and
reconstructs a cleaned correlation matrix, the forecast of realized risk is improved. These findings
suggests that employing random matrix theory in estimating the covariance matrix in portfolio
optimization can be beneficial in portfolio optimization. However, Laloux et al. (2000) mention
that the noise filtering will in particular improve the least risky portfolios, as these diversified
portfolios seem to mostly be influenced by noise.
However, no general consensus exists in how one should filter the perceived noisy eigenvalues.
The only consensus lies in the fact that regardless of filtering method, the trace of the correlation
matrix should be preserved to ensure that the variance of the system is preserved. Nevertheless,
prominent filtering methods may be found in Laloux et al. (2000), Plerou et al. (2002) and Sharifi
et al. (2004).
Robust statistics
More recently, Gerber et al. (2015) contributed to the research of enhancing the covariance matrix
estimation. In their working paper, a new measure of comovement is introduced altogether, based
on the field of robust statistics. They coin this measure as the Gerber Statistic which aims to be
17

Chapter 2. Literature Review

more robust than the conventional historical correlation by accommodating for noise and outliers
in the data. The claim is supported by a following evaluation test, where Gerber et al. (2015)
compare out-of-sample performance of MV optimized portfolios in a realistic investment setting.
The asset universe is a multi-asset universe, consisting of various asset classes such as equity
indices, bonds and commodities. Furthermore, the results of the evaluation test indicated that
the entire realized efficient frontier could be raised upwards by replacing the sample covariance
matrix with an estimated covariance matrix based on the Gerber Statistic, implying larger Sharpe
ratios. Solely by changing the estimation technique for the covariance matrix, ceteris paribus,
they showed that the MV optimized portfolios obtained via their newly introduced measure
consistently outperformed obtained portfolios via the sample covariance matrix over a range of
different investor profiles.
Clearly, the findings by Gerber et al. (2015) are of great interest for portfolio managers that
employ MVO, if deemed valid. However, the study gives rise to the question whether previously
developed estimators would yield an even stronger performance (i.e. weak competition). In
addition, in the article by Gerber et al. (2015), it is recognized that in computing the Gerber
Statistic, a non-positive semidefinite matrix may be obtained in theory. Clearly, this poses a
problem in the optimization process since a solution to the problem cannot be guaranteed to be
a globally optimal solution. The authors, however, claim that this problem has not been found
to occur, neither in real nor in simulated practice.
Comparing covariance matrix estimators in MVO
The most common approach found in the literature to compare various covariance matrix estimators is to analyze obtained MV optimized portfolios in backtesting procedures (see e.g. Bengtsson
and Holst (2002); Ledoit and Wolf (2003b)). This enables one to study the out-of-sample performance of the obtained portfolios in a real world setting, as data from the testing period is left out
in the estimation phase. However, many studies only evaluate the minimum variance portfolio.
This portfolio merely captures a glimpse of the relative performance in practice, as investors
have varying risk profiles. A common response to this shortcoming is that by leaving out the
expected return parameter, more emphasis is made on the covariance matrix and the relative
performance is not diluted by potential errors in the estimated vector of expected returns. While
this is a valid response, it will nonetheless render the results to be less applicable in practice.
Furthermore, many studies exclude asset classes beyond stocks while it has been shown that a
portfolio may experience significant benefits if asset classes such as commodity is included in the
asset universe due to the recent increase in equity volatility (Conover et al. 2010).
In light of some of the above issues regarding matching the real-world environment associated
with the vast majority of studies on this matter, this thesis will attempt to complement the
existing body of research by studying out-of-sample performance over a range of investor profiles.
In addition, the asset universe will include different asset classes. The covariance estimators that
will be investigated are the ones reviewed in this section. There are indeed more estimators to
be found in the literature, but it is deemed that these are not as prominent as the ones that have
been reviewed.

18

Chapter 3

Theoretical Framework
This chapter introduces the foundation that the thesis operates within. The framework is divided
into three tethered components. The first part serves as an introduction in the form of presenting
some crucial preliminaries. The following two parts will provide extensive theory on meanvariance optimization and estimation techniques regarding the expected return and the covariance
matrix, with the focal point lying in the latter of these elements.

3.1

Basic preliminaries

This section serves as a brief introduction to fundamental notions and concepts with regard to
mean-variance optimization.

3.1.1

Return

The return of a financial security is is the gain or loss over a certain time horizon. The return
of a security in this thesis is denoted as R where the return for a particular security i is Ri .
In addition, the price of a security at a particular time, t, is denoted as St . The return for a
non-dividend security over a time period [t, T ], where T > t is then calculated in the following
manner:
ST − St
(3.1.1)
St
As for a security that pays dividends over the time period [t, T ], the return is calculated as:
Ri =

Ri =

Divt,T + (ST − St )
=
St

Divt,T
S
| {zt }

Yield of dividends

+

ST − St
S
| {zt }

(3.1.2)

Capital gain yield

At time t, the outcome of ST and Divt,T is not known, hence the return of a security over the
time period is a random variable. The expected outcome of this random variable will be denoted
as:
µi = E[Ri ]
(3.1.3)
In the setting of portfolio allocation, a portfolio can be composed of multiple securities where
each portfolio weight, wi , of security i is calculated accordingly:
Value of the investment in security i
(3.1.4)
wi =
Total value of the portfolio
19

Chapter 3. Theoretical Framework

Thus, for an asset universe of size n, we have that:
n
X

wi = 1

(3.1.5)

i=1

The expected return of such a portfolio, E[RP ] is then given by:
" n
#
n
n
X
X
X
E[RP ] = E
wi Ri =
wi E[Ri ] =
wi µi = w> µ
i=1

i=1

(3.1.6)

i=1

Here, w> = [w1 , . . . , wn ] and µ = [µ1 , . . . , µn ]> . I.e., these are vector notations where w, µ ∈
Rn×1 .

3.1.2

Variance

In the setting of mean-variance optimization introduced by Markowitz (1952), the notion of
variance is also fundamental. For a portfolio of assets, where the asset universe consists of n
assets, its variance can be derived in the following manner:

V ar(RP ) = E 

n
X

!2 
wi Ri − E[RP ]



=E

n
X

i=1


=E

=

!2 
wi (Ri − E[Ri ])



i=1

n
X


! n
X
wi (Ri − E[Ri ]) 
wj (Rj − E[Rj ])

i=1

j=1

n X
n
X
wi wj σij = w> Σw
wi wj E[(Ri − E[Ri ])(Rj − E[Rj ])] =
|
{z
}
i=1 j=1
i=1 j=1

n X
n
X

:=σij

To summarize, the variance of the portfolio is given by:
V ar(RP ) = w> Σw

(3.1.7)

where Σ denotes the covariance matrix of the asset returns, composed of all covariances between
the returns defined as σij (σii ∀ i = {1, . . . , n} simply is the variance of asset i’s return, these
constitute the diagonal of the covariance matrix).

3.1.3

Optimization

In mathematics, optimization refers to the selection of a best element, with regard to certain
conditions, from a set of possible alternatives. A mathematical representation of an optimization
problem is presented below:
(
min
f (x)
x
(3.1.8)
s.t.
gi (x) ≤ bi , i = 1, . . . , m
The solution to this general form must lie within the feasible set which is given by F = {x ∈
Rn : gi (x) ≤ bi , i = 1, . . . , m}. A graphical example of a feasible set with the form as in (3.1.8),
spanned by three constraints, is illustrated in Figure 3.1. In this example, it is assumed that x1
and x2 (two dimensional case) cannot take negative values.
20

Chapter 3. Theoretical Framework

x2
8
7

C g3
b
1

6
≤

(x
)

g
1(
x

)

b
3

≤

4

≤

)
(x
g2

5

b2

B

3

D

2
1

E

0

A

0

1

2

3

4

5

6

7

8

9

10

x1

Figure 3.1: A feasible set spanned by three constraints. The shadowed area illustrates the feasible set.

Convexity
With regard to opimization, convex functions plays an important role. They have the appealing
property that a local minimum is also a global minimum. For concave functions, this means
that a local maximum is also a global maximum. Reconnecting to the feasible set spanned by a
certain number of constraints in the general optimization problem 3.1.8, a set F ⊂ Rn is referred
to as convex if for all x, y ∈ F and t ∈ [0, 1], it holds that:
(1 − t)x + ty ∈ F

(3.1.9)

Now, a function f : F → R is called convex if for all x, y ∈ F, the following holds:
f ((1 − t)x + ty) ≤ (1 − t)f (x) + tf (y) ∀t ∈ (0, 1)

(3.1.10)

If the above holds when the equality sign is flipped, f is said to be concave. A graphical
illustration of a convex function is shown in Figure 3.2.

Figure 3.2: An illustration of a convex function. Source: Sasane and Svanberg (2014, p.79).

21

Chapter 3. Theoretical Framework

Optimality
Moreover, optimality is an important concept in mean-variance optimization. Given a real-valued
function of a random variable x, a point x̂ ∈ F is said to be a local minimizer if it holds that:
f (x̂) ≤ f (x) ∀x ∈ F such that |x − x̂| < δ, δ ∈ R+

(3.1.11)

and a global minimizer of f if the following holds true:
f (x̂) ≤ f (x) ∀x ∈ F

(3.1.12)

An illustration for these two events is depicted in Figure 3.3

Figure 3.3: Q and all points in the interior of the segment AB are local minimizers, whereas P is a global
minimizer. Source: Sasane and Svanberg (2014, p. 105).

Reconnecting to convexity, the special property of a convex (concave) problem entails that a
local minimizer (maximizer) is also a global minimizer (maximizer). Clearly, this is an appealing
property in the optimization process as a solution to the problem can be guaranteed to be a
globally optimal solution. This renders the optimizer to produce stable solutions.
Quadratic programming
Quadratic programming (QP) refers to the problem of minimizing or maximizing a quadratic
function subject to linear equality and inequality constraints. Let f : Rn → R be a quadratic
function with the following representation:
1
f (x) = x> Hx + c> x + c0
(3.1.13)
2
where x ∈ Rn , H ∈ Rn×n is a symmetric matrix, c ∈ Rn and c0 ∈ R. A general mathematical
representation of minimizing the above function is as follows:

1 >
>

2 x Hx + c x + c0
min
x
(QP) s.t. Ax = b
(3.1.14)


x≥0
where A ∈ Rm×n and b ∈ Rm . Note that all of these are given, including H, c and c0 , with an
exception of x, sought to be solved. The above problem (QP) is a general form of the optimization
problem presented by Markowitz (1952) for portfolio allocation. Reconnecting to convexity, the
quadratic function f (x) is convex if and only if H is positive semi-definite. Consequently, the
optimization problem is perceived to be nice if this holds. This implication is important to
remember throughout this thesis.
22

Chapter 3. Theoretical Framework

3.2

Portfolio optimization

This section introduces the fundamental concepts of Modern Portfolio Theory (MPT). Serving
as the cornerstone of MPT, the mean-variance optimization framework for portfolio selection
is presented. Moreover, this section derives the MV formulation for allocating asset portfolios
in a quantitative manner. The structure of this section is largely inspired by Lundström and
Svensson (2014).
In its most basic form, the problem of portfolio selection may be summarized by the following
four aspects (Steuer, Qi, and Hirschberger (2008); Lundström and Svensson (2014)):
• A fixed amount of money to be invested
• An asset universe of size n, constituted of possible security investments
• A predetermined holding period for the portfolio
• A portfolio rebalance frequency, determining the length of possible sub-periods within the
holding period
Recalling the notations outlined in Section 3.1.1, the portfolio weight for the i:th asset is
denoted as wi . As shown in equation 3.1.4, the portfolio weights are defined as proportions of
the fixed sum to be invested. Therefore, the portfolio weights must sum to 1. This relation serves
as the first constraint for the optimization problem and is often referred to as the condition of
a fully invested portfolio. Furthermore, note that future returns of assets are unknown at the
beginning of the holding period, and thus to be considered as random variables. However, in
MPT, the optimizer assumes that all future asset characteristics (µi , σii and σij ) are known
when the optimization is initialized. As this does not hold true in reality, these parameters have
to be estimated. The real performance of the optimized portfolio is thus heavily reliant on the
estimation accuracy of these parameters. How these are estimated in practice will be further
described in the subsequent sections, Section 3.3 and Section 3.4.
Mathematically, the random return for a portfolio is defined below:
RP =

n
X

wi Ri = w> R

(3.2.1)

i=1

where R = (R1 , R2 , . . . , Rn )> . Now, assuming that an investor is solely interested in maximizing the uncertain portfolio return, the portfolio selection problem has the following stochastic
programming representation:
(
maximize RP = w> R
w
(SP)
subject to w ∈ F
Here, F defines the feasible region that is spanned by the portfolio constraints:
F = {w ∈ Rn |

n
X

wi = 1, αi ≤ wi ≤ βi }

(3.2.2)

i=1

The second constraint bounds the weights for each asset, where αi and βi is the lower and
upper bound, respectively. Two notable cases are referred to as the unconstrained and constrained case. The unconstrained case allows for wi to take any value (i.e. αi → −∞ and βi
→ ∞), which is equivalent to removing the weight constraint altogether. On the other hand, a

23

Chapter 3. Theoretical Framework

common constrained case is to bound the portfolio weights by imposing αi = 0. This is equal to
not allowing short positions in the portfolio. Moreover, an important aspect to note is that (SP)
is a stochastic programming problem. This follows from the fact that the future returns of the
securities are random variables and the composition of the portfolio (i.e. the vector of portfolio
weights, w) must be determined at the beginning of the holding period (Steuer et al. 2008).
Hence, in similar fashion as in Steuer et al. (2008), (SP) will be referred to as the investor’s
initial stochastic programming problem. At this stage, (SP) is not a tractable problem to solve
due to the presence of stochastic variables in the objective function.

3.2.1

MVO - an equivalent deterministic formulation

The difficulty with a stochastic programming problem is that its solution is not well defined.
This renders the problem to become intractable in a mathematical sense as it is not solvable
through standard optimization methods.
To cope with this difficulty and to solve (SP), one requires an interpretation and a decision
(Steuer et al. 2008). A common approach taken in the literature is to formulate the stochastic
problem as an equivalent deterministic problem (Lundström and Svensson 2014). Typically,
these formulations involve the utilization of some statistical characteristic or characteristics of
the random variables. In short, (SP) has to be transformed into a simplified deterministic
problem in order to be solvable. In this context, Steuer et al. (2008) argue that it is illuminating
to delve into the rationale that leads from (SP) to an equivalent deterministic formulation and
reviews some historical findings as follows.
In the early 17th century, mathematicians assumed that a gambler would be indifferent in
receiving the uncertain outcome of a gamble and receiving its expected outcome in cash. In the
context of portfolio selection, this assumption translates to the situation where an investor is
indifferent in holding a portfolio of stocks or receiving its certainly equivalent (CE), defined as
follows:
CE = E[RP ]

(3.2.3)

Under the assumption that an investor seek to maximize the amount of cash received for
certain, one arrives at the following deterministic representation:
(
maximize E[RP ] = w> µ
w
(3.2.4)
subject to w ∈ F
Recall here that the column vector of random future returns over a time period of n assets is
denoted as: R = (R1 , R2 , . . . , Rn )> . The column vector of expected values is further defined as
µ = E[R].
However, in 1738, Bernoulli discovered the famously known St. Petersburg paradox (a translated version of his work can be found in Bernoulli (1954)). The paradox provides an example of
a gamble with an infinite expected value but where a gambler in reality would be willing to forgo
the gamble in exchange for a finite/less amount of money (Steuer et al. 2008). This example
contradicts the classical theory that an investor is solely interested in maximizing the expected
cash outcome without taking its volatility into account. Hence, in order to better match the real
world setting and in contrast to the previous theory, Bernoulli suggested not to directly compare
cash outcomes, but rather to compare the utilities of cash outcomes. Now, if the utility of a
cash outcome is given by a function U : R → R, the utility of the certain equivalent can be
represented as follows:

24

Chapter 3. Theoretical Framework

U (CE) = E[U (RP )]

(3.2.5)

More specifically, the utility of the certainty equivalent equals the expected utility of the
uncertain future portfolio return. Given that an investor seeks to maximize the utility of the
certainty equivalent, we arrive at Bernoulli’s principle of maximizing expected utility as follows:
(
maximize E[U (RP )]
w
(UP)
subject to w ∈ F
This moves us one step further towards the mean-variance optimization problem suggested
by Markowitz (1952). However, in its present form, (UP) cannot be solved as the utility function
and its parameters are unknown. Therefore, the next step is to find a suitable utility function
that properly mirrors the utility function of an investor.
In the literature, two schools of thought have evolved for dealing with the undetermined
nature of the utility function. The first one involves attempting to incorporate an investor’s
preference structure into (UP) and obtaining an optimal portfolio, as suggested by Roy (1952).
The other one, in the spirit of Markowitz, has arguably found most traction within the field of
portfolio optimization. It involves a parameterization of the utility function U : thereafter, (UP)
is solved for all possible values of its unknown parameters. In this context, Markowitz (1952)
considered the following parameterized quadratic and concave utility function:
U (x) = x − (δ/2)x2

(3.2.6)

Using this quadratic utility function, Markowitz (1952) showed that an optimal portfolio for
an investor with a risk-aversion coefficient δ can be obtained by solving the following deterministic
problem:
(
maximize E[RP ] − 2δ V ar(RP )
w
(DP)
subject to w ∈ F
The set of all optimal solutions of w ∈ Rn is called the efficient set, which constitutes the
efficient frontier. This frontier refers to efficient portfolios that has the highest possible expected
return given a specified level of variance, or conversely, the lowest possible variance given a
specified level of return. The interpretation of Markowitz’s suggested utility function is that
an investor is solely interested in the relationship between expected return and variance when
choosing between portfolios, where variance is a proxy for risk. Here, it is assumed that all
investors are risk-averse. This means that any additional risk lowers the perceived utility. The
risk-aversion coefficient determines the perceived cost of risk.
In greater detail, mean-variance optimization (MVO) seeks to find optimal asset allocations
when both expected risk and return is considered. Indeed, this can be formulated mathematically
in a number of ways. In this study the utility-maximization problem is considered, which implies
that maximal utility is the objective for the investor. This is obtained by including both the
expected risk and the expected return in the objective function, where the trade-off is determined
by a risk aversion parameter. Extending (DP), we finally arrive at the following MVO problem:

25

Chapter 3. Theoretical Framework


δ

maximize w> µ − w> Σw


w

2
n
(MVO) subject to P wi = 1


i=1


αi ≤ wi ≤ βi , ∀i = {1, 2, . . . , n}.
In this form, the problem is solvable by employing tractable optimization methods. Furthermore, recall that w> µ denotes the expected portfolio return and that w> Σw is the portfolio
variance, where Σ is the covariance matrix of all asset returns. Solving (MVO) for different
values of the risk-aversion coefficient, one obtains an efficient frontier as displayed in Figure 3.4:

Expected Return

Eﬃcient frontier

All feasible portfolios that can be
formed within the asset universe.
Transition point

Volatility (I.e. Standard Deviation)

Figure 3.4: The blue frontier illustrates the efficient frontier. According to Markowitz, an investor should
only consider portfolios lying on the efficient frontier when selecting a portfolio. The choice of portfolio
on the efficient frontier depends on the risk-aversion coefficient of the investor.

The unconstrained optimal solution to the above MVO, i.e. when the portfolio weight constraints are lifted, has the following representation:



Σ−1 1 1
1a
?
−1
w =
+
Σ
µ−
(3.2.7)
c
δ
c
where
a = 1> Σ−1 µ
c = 1> Σ−1 1
where 1 is a column vector of ones. For future references, note that the only portfolio that
does not require an estimate of expected returns, thus only depndent on the covariance matrix,
is the minimum variance portfolio (MVP). This is given by:
Σ−1 1
(3.2.8)
c
The proof of these solutions can be found in Hult et al. (2012, p.88) for the interested reader.
wM V P =

26

Chapter 3. Theoretical Framework

3.2.2

Matching the real world - short selling and transaction costs

In the setting of portfolio optimization, short-selling constraints are commonly imposed. The
mathematical implication of this is to require the portfolio weights wi :s to be greater or equal to
zero. In practice, this constraint is frequently imposed as many funds and institutional investors
are prohibited from selling short. A natural question that one may ask is whether enforcing
such a constraint leads to suboptimal solutions as the optimizer is not allowed to freely allocate
portfolio weights within the set of real numbers. Assuming that the optimizer is accurate and
not influenced by estimation error, this is true in a global context. However, if the investor is
prohibited from taking short positions, the constraint is not viewed as a limitation, but rather an
adjustment so that the optimizer properly reflects the environment where the investor prevails.
Interestingly, it has also been shown in the literature that enforcing short-selling constraints
often improves out-of-sample performance of the optimized portfolios. This was shown in Jagannathan and Ma (2003) who proved that mean-variance optimizers are implicitly applying
some form of shrinkage on the sample covariance matrix when short positions are not allowed,
consequently leading to more stable portfolio weights. Their findings can be explained by the
prevalent perception that the optimal portfolio tends to amplify large estimation errors in certain
directions. This stems from the inherent behavior of the mean-variance optimizer, which will
assign large weights to assets that appear to have a small variance due to a significant underestimation. Similarly, if the expected return of an asset is significantly overestimated and appears to
be large, a large weight will be assigned to the corresponding asset. Thus, the portfolio risk of the
optimal portfolio is typically underpredicted and the return overpredicted (Karoui 2013). However, imposing short-selling constraints does not allow the optimizer to assign extreme weights
as they are then bounded between zero and one. As a result, the problem of error amplification is reduced. This finding further motivates investors that apply MVO to rule out short sale
positions.
The long-only MVO problem thus becomes:

δ

maximize w> µ − w> Σw


w

2
n
P
(3.2.9)
subject
to
wi = 1


i=1


0 ≤ wi ≤ 1, ∀i = {1, 2, . . . , n}.
Furthermore, the inclusion of transaction costs in the portfolio selection problem is important
to consider in order to better reflect the real world capital markets. In the initial MVO problem
presented by Markowitz (1952), transaction costs were ignored. However, when portfolios are
frequently rebalanced, the effect of transaction costs are far from insignificant. Recall that MV
optimizers are by nature prone to estimation error due to the stochastic nature of future returns
and covariances. In addition, small changes in these estimates (the vector of expected returns and
the covariance matrix) can result in reallocations that would not necessarily occur if transaction
costs were incorporated in the model. As a result, considering the inclusion of transaction costs
and incorporating it into the model is expected to reduce the amount of trading and rebalancing.
This is appealing for investors, as a low portfolio turnover is preferable.
Moreover, disregarding transaction costs may render inefficient portfolios as the cost of rebalancing may overwhelm the expected monetary gain of reallocating the portfolio. Thus, the
investor’s ability to allocate optimal portfolios is undermined.
There are numerous ways to implement transaction costs into the problem of portfolio selection. Some of these involve complicated nonlinear functions that emulate the cost penalty
function of transaction costs. However, while these functions might potentially capture the actual incurred effects of transaction costs, they come with the cost of computational intractability.
27

Chapter 3. Theoretical Framework

Thus, common approaches in practice involve a simplification to the transaction cost function
which assumes the penalty function to only be dependent on the proportional cost of portfolio
weight changes. Mathematically, the net expected portfolio return can thus be represented as
follows:
E[RP ] = w> µ − (b> max{0, w − w0 } + s> max{0, w0 − w})
{z
}
|

(3.2.10)

penalty function of transaction costs

n

Here, b ∈ R is the proportional cost to purchase assets and s ∈ Rn the proportional cost
to sell assets. Furthermore, w0 contain the weights of the current portfolio, which is held at the
time as the optimizer is initialized. In practice, when incorporating transaction costs, b and s
are often set to be equal. Furthermore, one can make the assumption that the proportional cost
of all individual assets are homogeneous. Clearly, this is a simplified case. It has the advantage
of allowing the model to accommodate for transaction costs while still being easy to implement.
Under the assumption that the cost of purchasing and selling assets is the same, in addition to
assuming that the proportional transaction cost is the same for every asset, we can define the
term accommodating for transaction costs in the following manner (using equation 3.2.10):
ψ1> Λ = b> max{0, w − w0 } + s> max{0, w0 − w}

(3.2.11)

Here, Λ is a vector of absolute values of portfolio weight changes, ψ a fixed proportional transaction cost, and 1 is a column vector of ones. In practice, ψ typically ranges between 10 to 50
basis points (Bessler, Opfer, and Wolff (2014);Gerber et al. (2015)).
Now, combining the long-only MVO problem in 3.2.9 with equation 3.2.11, we arrive at the
following long-only MVO representation, where the effect of transaction costs is incorporated in
the asset allocation model:

δ


max w> µ − ψ1> Λ − w> Σw


w
2


n
X
(3.2.12)
(MVO)
s.t.
wi = 1




i=1


0 ≤ wi ≤ 1, ∀i = {1, 2, . . . , n}.
To clarify once again, Λ is a vector of absolute values of portfolio weight changes, ψ a fixed
proportional transaction cost, and 1 is a column vector of ones. δ is a risk aversion parameter
that determines the trade-off between risk and return in the problem. The first constraint forces
the portfolio to be fully invested in the included assets. The second constraint bounds the weights
for each asset, where 0 and 1 is the lower and upper bound, respectively (short-selling is thus
prohibited). Indeed, (3.2.12) requires both µ and Σ to be estimated before the optimization
is attempted. This leaves us the important task of constructing viable estimates, which is the
purpose of the subsequent sections.
Furthermore, in (3.2.12) the matrix Σ needs to be positive semidefinite for the problem to be
concave and consequently have the property that a local optimum is also a global optimum (see
Section 3.1.3). If the matrix is negative definite, the problem is non-concave, making it difficult
to find the global optimum. In order for the matrix to be a valid covariance matrix, it should by
definition be positive semidefinite.

28

Chapter 3. Theoretical Framework

3.3

Estimating expected returns

The expected returns need to be estimated before problem 3.2.12 can be solved. The Capital
Asset Pricing Model (CAPM) proposes a method for such estimation. Under rational expectations of investors, the CAPM model offers a quick quantitative insight of risk-reward interplay
of assets. Three main assumptions underlie CAPM (Berk and DeMarzo 2014):
1. Investors are rational and choose mean-variance efficient portfolios according to Markowitz
(1952).
2. Investors are in complete agreement: i.e. investors have homogeneous expectations regarding the volatilities, correlations and expected returns of the assets.
3. Investors can borrow and lend at the risk-free rate, which is the same for all investors.
Under these assumptions, the CAPM implies that the market portfolio of all risky securities
is an efficient portfolio. Furthermore, the capital market line (CML) is spanned by the set of
portfolios with the highest possible expected return for a given level of volatility (risk) (Berk and
DeMarzo 2014).
At its core, CAPM is a single-factor linear model (commonly referred to as the security market
line) that relates the expected return of an asset and the market portfolio. The mathematical
representation of the security market line is as follows:
µi = E[Ri ] = rf + βi,mkt (E[Rmkt ] − rf )
{z
}
|

(3.3.1)

Risk premium for security i

Here, βi,mkt serves as a measure of non-diversifiable (systematic) risk. It measures the amount
of risk associated with an asset, that is common to the market risk. Ri is the return of some
asset, Rmkt the total return of the market, and rf the risk-free interest rate available to all
investors in the market. The estimate of βi,mkt can be obtained by using the sample covariance
and variance:
Cov(Ri , Rmkt )
βi,mkt =
(3.3.2)
V ar(Rmkt )
An estimate of the expected market return E[µmkt ], denoted as µ̂mkt , is obtained by taking
the mean return over a lookback period for each asset, where the mean returns are denoted
as λ̂i , and then weighing the means by their respective market weight mi . In other words,
µ̂mkt = m> λ̂ where m = (m1 , . . . , mN )> and λ̂ = (λ̂1 , . . . , λ̂N )> . This is referred to as a
value-weighted market index and serves merely as a proxy for the market portfolio. In reality,
the global market portfolio is unknown which is why proxies are employed. Now, the CAPM
estimated return vector is denoted:
µ̂ = (µ̂1 , . . . , µ̂N )>

(3.3.3)

where each individual return estimate is obtained by the estimated CAPM returns, i.e.:
µ̂i = rf + β̂i,mkt (µ̂mkt − rf )

(3.3.4)

In Section 3.4.2, the idea behind single-index models such as CAPM is explained in greater
detail, hopefully providing a better understanding of the model.

29

Chapter 3. Theoretical Framework

3.4

Estimating the covariance matrix

In order to estimate the risk of a portfolio, one needs to know to which degree the assets in the
portfolio face common risks and how their returns move together. For this intended purpose, the
covariance is a widely employed measure.
In the context of covariance matrix estimation, the conventional sample covariance matrix
has the appealing property of being the maximum likelihood estimator under the assumption
of normality. This means that it is the best unbiased estimator. However, it also comes with
drawbacks. Being the maximum likelihood estimator, all the trust is put in the data. This is a
sound principle, provided that there is enough data (Ledoit and Wolf 2003b). In small samples,
however, the estimator is subject to the risk of overfitting the data (follows noise too closely). This
means that the sample covariance matrix (also referred to as the historical covariance matrix)
may perform poorly out-of-sample, despite the fact that it performs best in-sample. Intuitively,
one might think that increasing the lookback window of the sample solves this, but in reality,
it may come at the cost of trusting outdated data with little explanatory power for the future.
In a sense, the estimator is a double edged sword as the low bias often comes at the cost of
high variance (see top right circle in Figure 3.5). In addition, the sample covariance matrix runs
the risk of becoming ill-conditioned if the number of assets under consideration is large relative
to the number of historical observations. More specifically, if the number of assets exceeds the
number of observations for every asset, the sample covariance matrix will not be invertible which
is very alarming in a portfolio optimization context (Bengtsson and Holst 2002).

Figure 3.5: Illustration of the bias-variance tradeoff. The blue dots correspond to estimates. Estimates
within the red circle are considered accurate. Source: Fortmann-Roe (2012, p. 1).

3.4.1

The sample covariance matrix

Let ri,t denote the historical return for asset i at time period t. Then, the average historical
return over the time span [1, T ] with step increment of size one for each asset i is given by:
T

r̄i =

1X
ri,t
T t=1

(3.4.1)

The sample covariance between two assets can then be estimated in the following manner:
Cov(ri , rj ) =

T
1 X
(ri,t − r̄i )(rj,t − r̄j ) := σ̂ij
T − 1 t=1

30

(3.4.2)

Chapter 3. Theoretical Framework

Performing equation 3.4.2 for all combinations i, j of assets, the historical covariance matrix
for N assets is then obtained by:


σ̂11 σ̂12 · · · σ̂1N
 σ̂21 σ̂22 · · · σ̂2N 

b HC = 
Σ
 ..
..
..  .
..
 .
.
.
. 
σ̂N 1

σ̂N 2

···

σ̂N N

Note that the covariance matrix Σ can be constructed according to:
Σ = diag(σ) C diag(σ)

(3.4.3)

where σ is a column vector of standard deviations. diag(σ) denotes a matrix with the elements of
σ on the main diagonal. C is a correlation matrix. Thus, by changing the method of calculating
correlation and in turn the correlation matrix C, different covariance matrices can be obtained.
Here, the estimate of σ, denoted as σ̂, is obtained as the sample standard deviation of the
historical asset returns. Note that σ̂ will not depend on which correlation estimation method is
used. Therefore, the estimated covariance matrices obtained by historical correlation are:
b HC = diag(σ̂)C
bHC diag(σ̂)
Σ
An expression of CHC is given by:


ρ1,1
 ρ2,1

CHC =  .
 ..

ρ1,2
ρ2,2
..
.

···
···
..
.


ρ1,N
ρ2,N 

.. 
. 

ρN,1

ρN,2

···

ρN,N

(3.4.4)

where ρi,j is the correlation between asset returns ri and rj . The estimate of CHC , denoted as
bHC , is computed by using the pairwise sample correlations of the historical asset returns.
C

3.4.2

The single-index market model

A prominent competitor to the sample covariance matrix is the single-index model introduced by
Sharpe (1964). The single-index model is a one factor model that attempts to cure the problem of
overfitting associated with the sample covariance matrix by imposing structure on the estimator.
By observing stock prices, one can see that individual stock prices tend to move together with
the aggregated market. Not surprisingly, the market factor thus usually proves to be the most
important factor when explaining the return generating process of stock returns. In addition,
this suggests that the comovement between stocks to some extent stems from a common response
to market changes.
In Sharpe’s (1964) si

[The evaluation harness truncated this reference: showing the first 120000 of 247834 characters.]
</reference>

<statements>
1. Risk is measured by variance of portfolio returns, given a vector of expected returns and a covariance matrix; optimization trades off expected return versus variance subject to constraints.
2. Variance (or standard deviation) of portfolio returns is the canonical risk measure; risk is determined by the covariance matrix \(\Sigma\) and weights \(w\) through \(w^\top \Sigma w\).
3. Markowitz’s framework is consistent with either (a) quadratic utility over mean and variance or (b) joint normality of returns; in practice, it assumes the mean vector and covariance matrix of future returns are known or well‑estimated, which is rarely true.
4. Optimization problem: Choose weights \(w\) to minimize variance for a target expected return, or maximize expected return for a given variance, or maximize mean-variance utility \(w^\top \mu - \gamma w^\top \Sigma w\), subject to budget and possibly other constraints.
5. Mean–Variance: Allocation behavior: Single‑period optimization; static weights chosen to trade off mean vs variance.
6. Mean–Variance: Main strengths: Simple, tractable; foundational; clear geometry of efficient frontier.
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.