























In various economic applications, people want to compare $n$ units with respect to certain quantities $Y_1, Y_2, \ldots, Y_n$ measuring their performance. The latter, however, is often influenced by certain factors which are beyond control of the units, and one would like to extract an adjusted performance from the data. Specifically, let $X_i \in \mathcal{X}$ summarize the factors of the $i$-th unit. Then one could think of a model equation $Y_i = f_o(X_i) + ε_i$ with a regression function $f_o : \mathcal{X} \to \mathbb{R}$ describing the unavoidable influence of the factors $X_i$ and $ε_i$ being the adjusted performance of the $i$-th unit. Now a common proposal is to estimate $f_o$ via regression methods by a function $\hat{f}$ depending on the current data $(X_i,Y_i)$, possibly augmented by additional past data, and to use the residuals $\hatε_i := Y_i - \hat{f}(X_i)$ as surrogates for the adjusted performances $ε_i$. In the present report we discuss this approach, its potential pitfalls and (mis)interpretation. In particular, an unavoidable property of the residuals $\hatε_i$ is that they measure only parts of the adjusted performance while the remaining parts get hidden in the estimated function $\hat{f}$. Possible alternatives are mentioned briefly.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。