Domain Adaptation of Foundation LLMs for e-Commerce

  • 2025-01-16 17:58:32
  • Christian Herold, Michael Kozielski, Tala Bazazo, Pavel Petrushkov, Hadi Hashemi, Patrycja Cieplicka, Dominika Basaj, Shahram Khadivi
  • 0

Abstract

We present the e-Llama models: 8 billion and 70 billion parameter largelanguage models that are adapted towards the e-commerce domain. These modelsare meant as foundation models with deep knowledge about e-commerce, that forma base for instruction- and fine-tuning. The e-Llama models are obtained bycontinuously pretraining the Llama 3.1 base models on 1 trillion tokens ofdomain-specific data. We discuss our approach and motivate our choice of hyperparameters with aseries of ablation studies. To quantify how well the models have been adaptedto the e-commerce domain, we define and implement a set of multilingual,e-commerce specific evaluation tasks. We show that, when carefully choosing the training setup, the Llama 3.1models can be adapted towards the new domain without sacrificing significantperformance on general domain tasks. We also explore the possibility of mergingthe adapted model and the base model for a better control of the performancetrade-off between domains.

 

Quick Read (beta)

loading the full paper ...