Loss As Model Inconsistency.pdf

richardson22b.pdf
Preview of Loss as Model Inconsistency
🔗 Source: proceedings.mlr.press
📊 Size: 717 KB
📄 Pages: 30 pages
⬇️ Downloads: 2,277

Summary

Many tasks in artificial intelligence have been cast as optimization problems, but the choice of objective is not unique. A key component of a machine learning system is a loss function which the system must minimize, and a wide variety of losses are used in practice. Each implicitly represents different values and results in different behavior, so the choice between them can be quite important. Yet, because it's unclear how to choose a "good" loss function, the choice is usually made by empirics, tradition, and an instinctive calculus acquired through practice—not by explicitly laying out beliefs. Furthermore, there is something to be gained by fiddling with these loss functions: one can add regularization terms, to (dis)incentivize (un)desirable behavior. However, the process of tinkering with the objective until it works is often unsatisfying. It can be a tedious game without clear rules or meaning, while results so obtained are arguably overfitted and difficult to motivate.

A choice of model admits more principled discussion, in part because models are testable; it makes sense to ask if a model is accurate. This observation motivates the proposal: instead of specifying a loss function directly, one articulates a situation that gives rise to it, in the (more interpretable) language of probabilistic beliefs and certainties. Concretely, we use the machinery of Probabilistic Dependency Graphs (PDGs), a particularly expressive class of graphical models that can incorporate arbitrary (even inconsistent) probabilistic information in a natural way, and comes equipped with a well-motivated measure of inconsistency.

A primary goal of this paper is to show that PDGs and their associated inconsistency measure can provide a "universal" model-based loss function. Towards this end, we show that many standard objective functions—cross entropy, square error, many statistical distances, the ELBO, regularizers, and the log partition function—arise naturally by measuring the inconsistency of the appropriate underlying PDG. This is somewhat surprising, since PDGs were not designed with the goal of capturing loss functions at all. Specifying a loss function indirectly like this is in some ways more restrictive, but it is also more intuitive (it no technical familiarity with losses, for instance), and admits more grounded defense and criticism.

For a particularly powerful demonstration, consider the variational autoencoder (VAE), an enormously successful class of generative model that has enabled breakthroughs in image generation, semantic interpolation, and unsupervised feature learning. Structurally, a VAE for a space X consists of a (smaller) latent space Z, a prior distribution p(Z), a decoder d(X|Z), and an encoder e(Z|X). A VAE is not considered a "graphical model" for two reasons. The first is that the encoder e(Z|X) has the same target variable as p(Z), so something like a Bayesian Network cannot simultaneously incorporate them both (besides, they could be inconsistent with one another). The second reason: it is not a VAE's structure, but rather its loss function that makes it tick. A VAE is typically trained by maximizing the "ELBO", a somewhat difficult-to-motivate function of a sample x, originating in variational calculus.

We show that −ELBO(x) is also precisely the inconsistency of a PDG containing x and the probabilistic information of the autoencoder (p, d, and e). We can form such a PDG precisely because PDGs allow for inconsistency. Thus, PDG semantics simultaneously legitimize the strange structure of the VAE, and also justify its loss function, which can be thought of as a property of the model itself (its inconsistency), rather than some mysterious construction borrowed from physics.

Representing objectives as model inconsistencies, in addition to providing a principled way of selecting an objective, also has beneficial pedagogical side effects, because of the structural relationships between the underlying models. For instance, these relationships will allow us to derive simple and intuitive visual proofs of technical results, such as the variational inequalities that traditionally motivate the ELBO, and the monotonicity of R´enyi divergence.

In the coming sections, we show in more detail how this concept of inconsistency, beyond simply providing a permissive and intuitive modeling framework, reduces exactly to many standard objectives used in machine learning and to measures of statistical distance. We demonstrate that this framework clarifies the relationships between them, by providing clear derivations of otherwise opaque inequalities.

Description

Richardson argues that loss function choice in probabilistic dependency graphs (PDGs) should align with the model, not personal preference. He shows that many standard loss functions and statistical divergences can be derived from PDG inconsistency, offering intuitive visual language and clarifying the evidence lower bound (ELBO) in variational inference.

Technical Information

  • File Format: PDF
  • File Size: 717 KB
  • Pages: 30
  • Language: EN
  • Total Downloads: 2,277
  • Last Updated: 5 hours ago

Document Overview

This PDF document about Loss as Model Inconsistency provides comprehensive information and guidance. Whether you're a beginner or advanced user, this resource offers valuable insights into Loss as Model Inconsistency.

Related Topics

If you're interested in Loss as Model Inconsistency, you might also want to explore:

Download Loss as Model Inconsistency eBooks for free and learn more about Loss as Model Inconsistency. These books contain exercises and tutorials to improve your practical skills, at all levels!

Not satisfied with this document? We have related documents to Loss as Model Inconsistency, try searching with similar keywords: Loss as Model Inconsistency, Inconsistency Asymmetry And Non Locality A Philoso, time inconsistency, Inconsistency, Inconsistency Book, Inconsistency Reference, Inconsistency Pdf, Inconsistency Guide

You can download PDF versions of the user's guide, manuals and ebooks about Loss as Model Inconsistency, you can also find and download for free A free online manual (notices) with beginner and intermediate, Downloads Documentation, You can download PDF files (or DOC and PPT) about Loss as Model Inconsistency for free, but please respect copyrighted ebooks.