We tackle the problem of action-conditioned generation of realistic anddiverse human motion sequences. In contrast to methods that complete, orextend, motion sequences, this task does not require an initial pose orsequence. Here we learn an action-aware latent representation for human motionsby training a generative variational autoencoder (VAE). By sampling from thislatent space and querying a certain duration through a series of positionalencodings, we synthesize variable-length motion sequences conditioned on acategorical action. Specifically, we design a Transformer-based architecture,ACTOR, for encoding and decoding a sequence of parametric SMPL human bodymodels estimated from action recognition datasets. We evaluate our approach onthe NTU RGB+D, HumanAct12 and UESTC datasets and show improvements over thestate of the art. Furthermore, we present two use cases: improving actionrecognition through adding our synthesized data to training, and motiondenoising. Our code and models will be made available.