How to train your MAML

Antreas Antoniou, Harrison Edwards, Amos Storkey

Edinburgh Research Explorer (University of Edinburgh) · 2018 · 83 citations · 24 references

DOIFull text

Open access

Concepts

TL;DR

Few‑shot learning has advanced via meta‑learning, with MAML being a leading but sensitive and computationally heavy approach that suffers from instability and hyperparameter tuning challenges. This work introduces MAML++, a set of modifications that stabilize training and enhance generalization, convergence speed, and computational efficiency. MAML++ achieves these gains through a suite of architectural and algorithmic changes that reduce sensitivity, accelerate convergence, and lower computational cost.

Abstract

The field of few-shot learning has recently seen substantial advancements. Most of these advancements came from casting few-shot learning as a meta-learning problem. Model Agnostic Meta Learning or MAML is currently one of the best approaches for few-shot learning via meta-learning. MAML is simple, elegant and very powerful, however, it has a variety of issues, such as being very sensitive to neural network architectures, often leading to instability during training, requiring arduous hyperparameter searches to stabilize training and achieve high generalization and being very computationally expensive at both training and inference times. In this paper, we propose various modifications to MAML that not only stabilize the system, but also substantially improve the generalization performance, convergence speed and computational overhead of MAML, which we call MAML++.

References

24