Research in image captioning has mostly focused on English because of theavailability of image-caption paired datasets in this language. However,building vision-language systems only for English deprives a large part of theworld population of AI technologies' benefit. On the other hand, creatingimage-caption paired datasets for every target language is expensive. In thiswork, we present a novel unsupervised cross-lingual method to generate imagecaptions in a target language without using any image-caption corpus in thesource or target languages. Our method relies on (i) a cross-lingual scenegraph to sentence translation process, which learns to decode sentences in thetarget language from a cross-lingual encoding space of scene graphs using asentence parallel (bitext) corpus, and (ii) an unsupervised cross-modal featuremapping which seeks to map an encoded scene graph features from image modalityto language modality. We verify the effectiveness of our proposed method on theChinese image caption generation task. The comparisons against several existingmethods demonstrate the effectiveness of our approach.