























Let $(X,Y)$ be a random couple in $S\times T$ with unknown distribution $P$ and $(X_1,Y_1),...,(X_n,Y_n)$ be i.i.d. copies of $(X,Y).$ Denote $P_n$ the empirical distribution of $(X_1,Y_1),...,(X_n,Y_n).$ Let $h_1,...,h_N:S\mapsto [-1,1]$ be a dictionary that consists of $N$ functions. For $λ\in {\mathbb{R}}^N,$ denote $f_λ:=\sum_{j=1}^Nλ_jh_j.$ Let $\ell:T\times {\mathbb{R}}\mapsto {\mathbb{R}}$ be a given loss function and suppose it is convex with respect to the second variable. Let $(\ell \bullet f)(x,y):=\ell(y;f(x)).$ Finally, let $Λ\subset {\mathbb{R}}^N$ be the simplex of all probability distributions on $\{1,...,N\}.$ Consider the following penalized empirical risk minimization problem \begin{eqnarray*}\hatλ^{\varepsilon}:={\mathop {argmin}_{λ\in Λ}}\Biggl[P_n(\ell \bullet f_λ)+\varepsilon \sum_{j=1}^Nλ_j\log λ_j\Biggr]\end{eqnarray*} along with its distribution dependent version \begin{eqnarray*}λ^{\varepsilon}:={\mathop {argmin}_{λ\in Λ}}\Biggl[P(\ell \bullet f_λ)+\varepsilon \sum_{j=1}^Nλ_j\log λ_j\Biggr],\end{eqnarray*} where $\varepsilon\geq 0$ is a regularization parameter. It is proved that the ``approximate sparsity'' of $λ^{\varepsilon}$ implies the ``approximate sparsity'' of $\hatλ^{\varepsilon}$ and the impact of ``sparsity'' on bounding the excess risk of the empirical solution is explored. Similar results are also discussed in the case of entropy penalized density estimation.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。