### Abstract

Since reinforcement learning algorithms are notoriously data-intensive, thetask of sampling observations from the environment is usually split acrossmultiple agents. However, transferring these observations from the agents to acentral location can be prohibitively expensive in terms of communication cost,and it can also compromise the privacy of each agent's local behavior policy.Federated reinforcement learning is a framework in which $N$ agentscollaboratively learn a global model, without sharing their individual data andpolicies. This global model is the unique fixed point of the average of $N$local operators, corresponding to the $N$ agents. Each agent maintains a localcopy of the global model and updates it using locally sampled data. In thispaper, we show that by careful collaboration of the agents in solving thisjoint fixed point problem, we can find the global model $N$ times faster, alsoknown as linear speedup. We first propose a general framework for federatedstochastic approximation with Markovian noise and heterogeneity, showing linearspeedup in convergence. We then apply this framework to federated reinforcementlearning algorithms, examining the convergence of federated on-policy TD,off-policy TD, and $Q$-learning.