Abstract
As pre-trained language models become more resource-demanding, the inequalitybetween resource-rich languages such as English and resource-scarce languagesis worsening. This can be attributed to the fact that the amount of availabletraining data in each language follows the power-law distribution, and most ofthe languages belong to the long tail of the distribution. Some research areasattempt to mitigate this problem. For example, in cross-lingual transferlearning and multilingual training, the goal is to benefit long-tail languagesvia the knowledge acquired from resource-rich languages. Although beingsuccessful, existing work has mainly focused on experimenting on as manylanguages as possible. As a result, targeted in-depth analysis is mostlyabsent. In this study, we focus on a single low-resource language and performextensive evaluation and probing experiments using cross-lingual post-training(XPT). To make the transfer scenario challenging, we choose Korean as thetarget language, as it is a language isolate and thus shares almost no typologywith English. Results show that XPT not only outperforms or performs on parwith monolingual models trained with orders of magnitudes more data but also ishighly efficient in the transfer process.