An Input Skip Can Reparameterize a VAE Without Improving the Generator
Kevin Lee
Abstract
A constant input skip in a Gaussian variational autoencoder (VAE) decoder is an affine reparameterization of the generator. Rescaling the residual by $b=1-\alpha$ and its variance by $b^2$ preserves the joint distribution of observations and latents and scales reconstruction mean squared error (MSE) by $b^2$. Each proper score of $p(x)$ is invariant to this map. Skip and $\alpha=0$ nets trained independently still use that decoder family. Since Adam and gradient clipping affect residual parameters, the fitted generators can differ. Changing skip and observation variance together on a Gaussian factor model with known density reduces MSE but degrades energy score and importance weighted log density. On seven annual \mbox{Deep-ocean} Assessment and Reporting of Tsunamis (DART) records, energy score at the same observation variance is mixed on four previously chosen records and better for the skip on three 2018 records. Fitted probabilistic principal component analysis (PPCA) has a lower energy score than every neural DART model. Minimizing the training reconstruction term over a scalar observation variance can raise the energy score of the implied $p(x)$.
Chat is not available.
Successful Page Load