Monotonic Chunkwise Attention

Abstract

Sequence-to-sequence models with soft attention have been successfullyapplied to a wide variety of problems, but their decoding process incurs aquadratic time and space cost and is inapplicable to real-time sequencetransduction. To address these issues, we propose Monotonic Chunkwise Attention(MoChA), which adaptively splits the input sequence into small chunks overwhich soft attention is computed. We show that models utilizing MoChA can betrained efficiently with standard backpropagation while allowing online andlinear-time decoding at test time. When applied to online speech recognition,we obtain state-of-the-art results and match the performance of a model usingan offline soft attention mechanism. In document summarization experimentswhere we do not expect monotonic alignments, we show significantly improvedperformance compared to a baseline monotonic attention-based model.

Quick Read (beta)

loading the full paper ...